Staff ML Engineer, Frontier AI
full-time
lead
Posted 5 months ago
Before you apply
Build my evidence-backed draft — free Apply on company site →Paste your relevant resume section or 2–4 true bullets. See supported requirements and honest gaps. No account and no application sent.
About this role
About Us:
Here at Ambience, we never set out to be just another scribe. We’re building the AI intelligence platform that restores humanity to healthcare and drives meaningful ROI for health systems across the country.
Our technology helps providers focus on delivering great care by removing the administrative burden that pulls them away from patients and away from their most impactful work. Ambience delivers real-time coding-aware documentation and clinical workflow support across ambulatory, emergency and inpatient settings at the top health systems in North America.
Our teams operate relentlessly with extreme ownership to build the best solutions for our health system partners. We value candor, positivity and deep thought — and we expect a lot from each other because we know the problems we’re solving truly matter.
Ambience was ranked #1 for Improving the Clinician Experience in the KLAS Research Emerging Solutions Top 20 Report, recognized by Fast Company as one of the Next Big Things in Tech, named one of the best AI companies in healthcare by Inc., and selected as a LinkedIn Top Startup in 2024 and 2025. We’re backed by Oak HC/FT, Andreessen Horowitz (a16z), OpenAI Startup Fund, and Kleiner Perkins — and we’re just getting started.
THE ROLE:
As a Staff Machine Learning Engineer at Ambience, you will help set the technical direction for the AI systems that power our clinical products. You’ll identify the highest-impact opportunities to improve model behavior, evaluation, post-training, and agentic systems, and lead the design and execution of cross-cutting initiatives.
This is a highly hands-on role with broad technical influence. You’ll work closely with clinicians, product managers, researchers, and engineers to translate cutting-edge research into reliable, production-grade AI systems.
Our engineering roles are hybrid — working onsite at our San Francisco office three days per week.
WHAT YOU’LL OWN:
- Define the AI Quality and Evaluation Strategy: Establish how we measure production AI quality across LLM and agentic systems, including automated graders, regression testing, human evaluation, failure taxonomies, and offline-to-online validation.
- Drive Production Model Improvement: Identify high-impact failure modes across products and lead improvements across prompting, retrieval, context, routing, data, and post-training. Design experiments that determine the right intervention and prove that improvements hold in production.
- Advance Post-Training Capabilities: Shape how Ambience uses techniques such as supervised fine-tuning, preference optimization, reinforcement learning, distillation, and synthetic data to improve model behavior.
- Architect Agentic AI Systems: Set technical direction for production systems involving tool use, retrieval, context and state management, routing, orchestration, tracing, and failure recovery.
- Build Self-Improving AI Loops: Design the systems that turn production failures and user feedback into better datasets, evaluations, and model behavior, creating scalable improvement flywheels across products.
- Shape the AI Roadmap: Co-develop the 12-month technical roadmap, balancing near-term product wins with longer-term research and platform investments.
- Raise the Technical Bar: Lead technical reviews, mentor engineers, establish best practices, and create reusable systems and frameworks that improve the effectiveness of the broader AI team.
- Stay at the Cutting Edge: Distill insights from recent research in LLMs, agents, post-training, NLP, speech, and multimodal AI, and drive experiments that keep Ambience at the forefront of clinical AI.
WHO YOU ARE:
- Expert in Production AI Systems
7+ years in production ML, research engineering, or applied AI.
Deep expertise with modern LLMs, transformers, and complex production AI systems.
Have led or materially shaped consequential AI systems with real production impact.
Able to reason across models, data, orchestration, evaluation, serving, and product behavior.
- Deep Evaluation Rigor
Significant experience designing evaluations for LLMs, agents, or other complex AI systems.
Can turn ambiguous product-quality problems into robust metrics, datasets, and experiments.
Experienced with grader bias, leakage, contamination, misleading aggregate metrics, regression detection, and offline-online mismatch.
Knows how to establish whether a model or system improvement actually translates into better user outcomes.
- Agentic Systems Expertise
Have built or architected production systems involving multiple models, tools, retrieval, context and state management, routing, orchestration, tracing, and failure recovery.
Understand how to evaluate and debug complex agent behavior across model, tool, and system boundaries.
Comfortable making architectural decisions that span multiple products or teams.
- Post-Training and Model Improvement Expertise
Similar Jobs
Related searches:
On-site Jobs
Lead Jobs
On-site Lead Jobs
Lead AI InfrastructureLead Robotics & AutonomyLead NLP & Language AILead Healthcare AILead Data EngineeringLead Generative AILead Machine Learning
AI Jobs in San Francisco
AI Infrastructure in San FranciscoRobotics & Autonomy in San FranciscoNLP & Language AI in San FranciscoHealthcare AI in San FranciscoData Engineering in San FranciscoGenerative AI in San FranciscoMachine Learning in San Francisco
generative-aimlopssearchpytorchreinforcement-learninghealthcarenlpfine-tuning
Get jobs like this delivered weekly
Free AI jobs newsletter. No spam.