Senior/Staff FDE - CUA
full-time
lead
Posted 2 days ago
Before you apply
Build my evidence-backed draft — free Apply on company site →Paste your relevant resume section or 2–4 true bullets. See supported requirements and honest gaps. No account and no application sent.
About this role
About Snorkel
At Snorkel, we believe meaningful AI doesn’t start with the model, it starts with the data.
We’re on a mission to help enterprises transform expert knowledge into specialized AI at scale. The AI landscape has gone through incredible changes since 2015, when Snorkel started as a research project in the Stanford AI Lab, to the generative AI breakthroughs of today. But one thing has remained constant: the data you use to build AI is the key to achieving differentiation, high performance, and production-ready systems. We work with some of the world’s largest organizations to empower scientists, engineers, financial experts, product creators, journalists, and more to build custom AI with their data faster than ever before. Excited to help us redefine how AI is built? Apply to be the newest Snorkeler!
About the Role
Snorkel AI is hiring a Forward Deployed Engineer focused on Computer Use Agents to partner with leading AI labs and enterprises on their most critical agentic-AI initiatives.
In this role, you will lead the technical execution of complex customer engagements involving agents that operate computers, browsers, and software environments to complete realistic, multi-step tasks. You will translate ambiguous product and model challenges into robust task environments, datasets, evaluators, and delivery plans that improve agent reliability and downstream performance.
You will work across the full delivery lifecycle—from technical discovery and solution design through implementation, evaluation, and production delivery. You will also identify patterns across engagements and turn successful approaches into reusable capabilities, technical standards, and product improvements.
Main Responsibilities
Computer Use Agents, Data, and Evaluation
Design and build task environments, datasets, and evaluation workflows for computer-using agents operating across browsers, desktop applications, terminals, and other software interfaces
Translate customer goals, agent failure modes, and real-world workflows into representative, multi-step tasks with clear success criteria
Develop data-generation, validation, and quality-assurance pipelines for multimodal and agentic training and evaluation data
Build automated evaluators, checks, and measurement frameworks to assess task completion, correctness, robustness, efficiency, and adherence to requirements
Diagnose agent failures across planning, tool use, perception, state management, and interaction with user interfaces; turn findings into improved tasks, data, and evaluations
Design and run experiments to measure how data, task design, and evaluation changes affect downstream agent performance
Deliver reusable, production-grade task suites, datasets, and evaluation assets that help customers train, benchmark, and improve computer-use agents
Forward Deployed Engineering & Customer Partnership
Lead technical workstreams from initial solution design through production delivery, navigating ambiguity and making sound technical decisions
Build, refine, and iterate on solutions that address customer needs, incorporating feedback to ensure the delivered work provides tangible value
Rapidly prototype and productionize solutions across models, agent frameworks, APIs, browser or desktop environments, and custom applications
Communicate technical tradeoffs, experimental results, and recommendations clearly to technical and cross-functional stakeholders
Serve as a trusted technical partner to customers and internal delivery teams, resolving complex blockers and driving alignment
Technical Leadership & Scale
Identify recurring patterns across customer engagements and turn successful solutions into reusable task frameworks, evaluators, tooling, and best practices
Define and improve technical standards for agent task design, environment reliability, evaluation, and delivery
Partner with DaaS Engineering, Research, and Product teams to influence platform and product capabilities based on real-world customer needs
Lead technical design reviews, share expertise, and provide guidance to other engineers
Stay current with emerging agentic-AI, computer-use, evaluation, and data-curation techniques and assess their applicability to customer problems
What We're Looking For
5+ years of experience in machine learning engineering, software engineering, applied AI, forward deployed engineering, solutions engineering, or a similar technical role
Strong Python skills and experience building reliable production software, data, or ML systems
Hands-on experience building, evaluating, or deploying LLM-based or agentic systems, including computer-use agents (CUA)
Strong understanding of experimentation and evaluation, including LLM-as-a-judge / model-based evaluation, defining metrics, and using empirical results to guide technical decisions
Experience designing task environments, datasets, and verifiers for agents, including reward & verifier desi
Similar Jobs
Related searches:
On-site Jobs
Lead Jobs
On-site Lead Jobs
Lead AI Agents & RAGLead NLP & Language AILead Data EngineeringLead Generative AILead Machine LearningLead Robotics & Autonomy
AI Jobs in New York
AI Agents & RAG in New YorkNLP & Language AI in New YorkData Engineering in New YorkGenerative AI in New YorkMachine Learning in New YorkRobotics & Autonomy in New York
generative-aifine-tuningagentsdata-pipelinellmreinforcement-learning
Get jobs like this delivered weekly
Free AI jobs newsletter. No spam.