Senior/Staff FDE - CUA

Snorkel AI · New York, NY · $180k - $320k
full-time lead Posted 1 day ago

Before you apply

Build my evidence-backed draft — free Apply on company site →

Paste your relevant resume section or 2–4 true bullets. See supported requirements and honest gaps. No account and no application sent.

Get weekly job alerts like this →

About this role

About Snorkel At Snorkel, we believe meaningful AI doesn’t start with the model, it starts with the data. We’re on a mission to help enterprises transform expert knowledge into specialized AI at scale. The AI landscape has gone through incredible changes since 2015, when Snorkel started as a research project in the Stanford AI Lab, to the generative AI breakthroughs of today. But one thing has remained constant: the data you use to build AI is the key to achieving differentiation, high performance, and production-ready systems. We work with some of the world’s largest organizations to empower scientists, engineers, financial experts, product creators, journalists, and more to build custom AI with their data faster than ever before. Excited to help us redefine how AI is built? Apply to be the newest Snorkeler! About the Role Snorkel AI is hiring a Forward Deployed Engineer focused on Computer Use Agents to partner with leading AI labs and enterprises on their most critical agentic-AI initiatives. In this role, you will lead the technical execution of complex customer engagements involving agents that operate computers, browsers, and software environments to complete realistic, multi-step tasks. You will translate ambiguous product and model challenges into robust task environments, datasets, evaluators, and delivery plans that improve agent reliability and downstream performance. You will work across the full delivery lifecycle—from technical discovery and solution design through implementation, evaluation, and production delivery. You will also identify patterns across engagements and turn successful approaches into reusable capabilities, technical standards, and product improvements. Main Responsibilities Computer Use Agents, Data, and Evaluation Design and build task environments, datasets, and evaluation workflows for computer-using agents operating across browsers, desktop applications, terminals, and other software interfaces Translate customer goals, agent failure modes, and real-world workflows into representative, multi-step tasks with clear success criteria Develop data-generation, validation, and quality-assurance pipelines for multimodal and agentic training and evaluation data Build automated evaluators, checks, and measurement frameworks to assess task completion, correctness, robustness, efficiency, and adherence to requirements Diagnose agent failures across planning, tool use, perception, state management, and interaction with user interfaces; turn findings into improved tasks, data, and evaluations Design and run experiments to measure how data, task design, and evaluation changes affect downstream agent performance Deliver reusable, production-grade task suites, datasets, and evaluation assets that help customers train, benchmark, and improve computer-use agents Forward Deployed Engineering & Customer Partnership Lead technical workstreams from initial solution design through production delivery, navigating ambiguity and making sound technical decisions Build, refine, and iterate on solutions that address customer needs, incorporating feedback to ensure the delivered work provides tangible value Rapidly prototype and productionize solutions across models, agent frameworks, APIs, browser or desktop environments, and custom applications Communicate technical tradeoffs, experimental results, and recommendations clearly to technical and cross-functional stakeholders Serve as a trusted technical partner to customers and internal delivery teams, resolving complex blockers and driving alignment Technical Leadership & Scale Identify recurring patterns across customer engagements and turn successful solutions into reusable task frameworks, evaluators, tooling, and best practices Define and improve technical standards for agent task design, environment reliability, evaluation, and delivery Partner with DaaS Engineering, Research, and Product teams to influence platform and product capabilities based on real-world customer needs Lead technical design reviews, share expertise, and provide guidance to other engineers Stay current with emerging agentic-AI, computer-use, evaluation, and data-curation techniques and assess their applicability to customer problems What We're Looking For 5+ years of experience in machine learning engineering, software engineering, applied AI, forward deployed engineering, solutions engineering, or a similar technical role Strong Python skills and experience building reliable production software, data, or ML systems Hands-on experience building, evaluating, or deploying LLM-based or agentic systems, including computer-use agents (CUA) Strong understanding of experimentation and evaluation, including LLM-as-a-judge / model-based evaluation, defining metrics, and using empirical results to guide technical decisions Experience designing task environments, datasets, and verifiers for agents, including reward & verifier desi

Similar Jobs

Related searches:

On-site Jobs Lead Jobs On-site Lead Jobs Lead NLP & Language AILead Data EngineeringLead Generative AILead Machine LearningLead Robotics & AutonomyLead AI Agents & RAG AI Jobs in New York NLP & Language AI in New YorkData Engineering in New YorkGenerative AI in New YorkMachine Learning in New YorkRobotics & Autonomy in New YorkAI Agents & RAG in New York generative-aifine-tuningagentsdata-pipelinellmreinforcement-learning

Get jobs like this delivered weekly

Free AI jobs newsletter. No spam.