Product Lead, AI/ML (Evals)
full-time
lead
Posted 3 weeks ago
Apply Now
Stand out: build a proof-of-work pitch →
Free GitHub-based preview. Direct apply stays one click away.
Get weekly job alerts like this →Hiring for this role?
AI Market Demand Pack · $29 one-time
Compare this role's skills with the full AI hiring market. Get ranked demand, salary bands, leading companies, public source URLs, and a decision brief.
About this role
ABOUT ABRIDGE
Abridge was founded in 2018 with the mission of powering deeper understanding in healthcare. Our AI-powered platform was purpose-built for medical conversations, improving clinical documentation efficiencies while enabling clinicians to focus on what matters most—their patients.
Our enterprise-grade technology transforms patient-clinician conversations into structured clinical notes in real-time, with deep EMR integrations. Powered by Linked Evidence and our purpose-built, auditable AI, we are the only company that maps AI-generated summaries to ground truth, helping providers quickly trust and verify the output. As pioneers in generative AI for healthcare, we are setting the industry standards for the responsible deployment of AI across health systems.
We are a growing team of practicing MDs, AI scientists, PhDs, creatives, technologists, and engineers working together to empower people and make care make more sense. We have offices located in the Mission District in San Francisco, the SoHo neighborhood of New York, and East Liberty in Pittsburgh.
The Role
Abridge's has multiple products - the core product is ambient clinical notes, but we’ve expanded surfaces: billing, clinical decision support, orders, nursing, etc. Measuring quality reliably and iterating fast without breaking clinician trust sit at the core of every single product’s success.
The Evals team owns the tooling, templates, and consultation that product teams use across the eval lifecycle: curate, build, iterate, and deploy. This builds the strategy, platform, and process that enables evals to be fast, trustworthy, consistent, and a true moat.
We are hiring a Product Manager to lead this platform. You will work with 10+ teams and partner closely with engineering, ML, data science, and clinician science. You will help set the standard for how Abridge measures AI quality, decide what the platform builds, and make evaluating a new model cheap and routine as frontier models ship frequently.
What You'll Do
- Drive product strategy and execution for the evals platform. Own the roadmap across the eval lifecycle and own outcomes against it.
- Build the shared measurement infrastructure. Help build the systems that let any pod define quality, run experiments, compare models, and watch production. Own the standards for LLM judges, rule-based evaluators, human annotation, and online monitoring, and be clear about where each belongs.
- Make model selection fast and routine. Give teams a repeatable way to evaluate a new frontier model within days of release, and a defensible framework for when a post-trained specialty model is worth it over a prompted frontier one. Keep the eval system model-agnostic so it stays a neutral referee.
- Own the operating model across pods. Define the eval gates from early build through GA and steady-state monitoring, which gates are hard versus advisory, and who owns non-negotiable floors like critical-error rates. Land this across pods you don't own, without formal authority.
- Cross functional execution. Work with Data Engineering on the de-identification pipeline, Data Science on bootstrapping judge quality with less human annotation, Clinical Science on flagged production cases, and the agent platform team as workflows go agentic.
- Operate with a high bar for quality, speed, and accountability.
What You Bring
- 5+ years of product management experience with significant ownership of ML powered products or platform systems.
- Deep understanding of how to measure and improve model quality, including evaluation frameworks, annotation pipelines, and benchmark design.
- Strong technical fluency across ML, data pipelines, and distributed systems.
- Experience working closely with ML researchers and engineers to drive impact in production.
- Ability to balance long term architectural investments with near term quality improvements.
- Strong communication skills and the ability to translate complex technical concepts into clear decisions and narratives.
- A track record of delivering high quality products in domains where accuracy, reliability, and trust are paramount.
Bonus Points If…
- You have experience building evaluation platforms, ML observability systems, or quality measurement pipelines.
- You have worked in clinical, healthcare, or regulated environments with a high bar for accuracy and compliance.
- You have worked on specialty specific or domain specific model adaptations.
- You have worked on personalization systems, context ingestion frameworks, or ambient intelligence products.
- You have experience shipping large scale ML products with human in the loop workflows.
WHY WORK AT ABRIDGE?
At Abridge, we’re transforming healthcare delivery experiences with generative AI, enabling clinicians and patients to connect in deeper, more meaningful ways. Our mission is clear: to power deeper understanding in healthcare. We’re driving real, lasting chan
Similar Jobs
Related searches:
Remote Jobs
Lead Jobs
Remote Lead Jobs
Lead AI Agents & RAGLead Machine LearningLead NLP & Language AILead Backend & SystemsLead Fintech & Payments AILead Data EngineeringLead Generative AILead Healthcare AILead AI InfrastructureLead AI Research
AI Jobs in San Francisco
AI Agents & RAG in San FranciscoMachine Learning in San FranciscoNLP & Language AI in San FranciscoBackend & Systems in San FranciscoFintech & Payments AI in San FranciscoData Engineering in San FranciscoGenerative AI in San FranciscoHealthcare AI in San FranciscoAI Infrastructure in San FranciscoAI Research in San Francisco
llmdata-pipelinedistributed-systemspaymentsagentsgenerative-aihealthcareevaluation
Get jobs like this delivered weekly
Free AI jobs newsletter. No spam.