Senior Machine Learning Engineer
full-time
senior
Posted 2 weeks ago
Before you apply
Build my evidence-backed draft — free Apply on company site →Paste your relevant resume section or 2–4 true bullets. See supported requirements and honest gaps. No account and no application sent.
About this role
About Us:
Here at Ambience, we never set out to be just another scribe. We’re building the AI intelligence platform that restores humanity to healthcare and drives meaningful ROI for health systems across the country.
Our technology helps providers focus on delivering great care by removing the administrative burden that pulls them away from patients and away from their most impactful work. Ambience delivers real-time coding-aware documentation and clinical workflow support across ambulatory, emergency and inpatient settings at the top health systems in North America.
Our teams operate relentlessly with extreme ownership to build the best solutions for our health system partners. We value candor, positivity and deep thought — and we expect a lot from each other because we know the problems we’re solving truly matter.
Ambience was ranked #1 for Improving the Clinician Experience in the KLAS Research Emerging Solutions Top 20 Report, recognized by Fast Company as one of the Next Big Things in Tech, named one of the best AI companies in healthcare by Inc., and selected as a LinkedIn Top Startup in 2024 and 2025. We’re backed by Oak HC/FT, Andreessen Horowitz (a16z), OpenAI Startup Fund, and Kleiner Perkins — and we’re just getting started.
THE ROLE:
As a Senior Machine Learning Engineer at Ambience, you will build and improve the AI systems that power our clinical products. You’ll own complex projects end-to-end, from diagnosing production failures and designing evaluations to building, deploying, and iterating on model and agentic systems.
This is a highly hands-on role with significant technical ownership. You’ll work closely with clinicians, product managers, and fellow engineers to translate cutting-edge research into reliable, production-grade AI systems.
Our engineering roles are hybrid — working onsite at our San Francisco office three days per week.
WHAT YOU’LL DO:
- Build Trustworthy AI Evaluation Systems: Design and own evaluation pipelines for LLM and agentic systems, combining automated graders, regression testing, production feedback, and human evaluation to measure real product quality.
- Improve Production Model Behavior: Diagnose high-impact failure modes and test improvements across prompting, retrieval, context, routing, data, fine-tuning, or other model and system interventions.
- Build Agentic AI Systems: Develop production systems involving tool use, retrieval, context and state management, routing, orchestration, tracing, and failure recovery.
- Build Data and Improvement Flywheels: Turn production failures and user feedback into better datasets, evaluations, and model behavior through active learning and systematic iteration.
- Stay at the Cutting Edge: Distill insights from recent research in LLMs, agents, NLP, speech, and multimodal AI and translate promising ideas into practical experiments.
- Own AI Systems End-to-End: Work across models, data, evaluation, orchestration, serving, and observability, while remaining deeply hands-on in code and production debugging.
WHO YOU ARE:
- Strong Production AI Experience
5+ years in production ML, research engineering, or applied AI.
Have built a consequential production AI system or materially improved model behavior in production.
Strong understanding of modern LLMs, transformers, and production AI systems.
- Deep Evaluation Experience
Experienced designing evaluations for LLMs, agents, or other complex AI systems.
Can turn ambiguous quality problems into measurable dimensions, datasets, and experiments.
Familiar with challenges such as grader bias, leakage, misleading aggregate metrics, regression detection, and offline-online mismatch.
- Agentic Systems Experience
Experience building production systems involving multiple models, tools, retrieval, context, state, routing, or orchestration.
Understands reliability and failure modes in complex AI workflows, not just individual model calls.
- Production-Grade Software Engineer
Proficient in Python and modern ML frameworks; PyTorch preferred.
Comfortable with deployment, observability, CI/CD, and containerized systems.
Still highly hands-on: writes code, inspects traces, analyzes failures, and debugs production systems.
- Data-Centric AI Developer
Skilled at building high-quality datasets and feedback loops.
Experienced using production failures, user feedback, and active learning to improve model and system quality.
- Effective Interdisciplinary Collaborator
Able to work closely with clinicians, product managers, and fellow engineers.
Strong communicator who can simplify complex AI concepts for diverse audiences.
Comfortable owning ambiguous technical problems and driving them to measurable outcomes.
Nice-to-Haves
- Experience with realtime voice, conversational AI, or multimodal systems.
- Experience with fine-tuning, post-training, or model adaptation.
- Prior work i
Similar Jobs
Related searches:
On-site Jobs
Senior Jobs
On-site Senior Jobs
Senior Fintech & Payments AISenior NLP & Language AISenior Healthcare AISenior Data EngineeringSenior Generative AISenior Machine Learning
AI Jobs in San Francisco
Fintech & Payments AI in San FranciscoNLP & Language AI in San FranciscoHealthcare AI in San FranciscoData Engineering in San FranciscoGenerative AI in San FranciscoMachine Learning in San Francisco
llmpaymentspytorchfine-tuningsearchgenerative-ainlphealthcare
Get jobs like this delivered weekly
Free AI jobs newsletter. No spam.