LLM Inference Engineer
full-time
mid
Posted 6 days ago
Before you apply
Build my evidence-backed draft — free Apply on company site →Paste your relevant resume section or 2–4 true bullets. See supported requirements and honest gaps. No account and no application sent.
About this role
ROLE MISSION
As HAI's LLM Inference Engineer, you will own the serving infrastructure that determines whether our breakthrough healthcare AI reaches patients efficiently and reliably. You'll optimize the systems that translate raw model capability into sub-100ms responses—making the difference between conversational experiences that feel natural and those that feel broken. This role exists because inference optimization at scale is where research meets reality: your work directly determines latency, cost, and availability for millions of patient conversations across healthcare systems.
WHAT YOU WILL ACCOMPLISH
Own your first major outcome: By day 90, you will have shipped a measurable improvement to our inference serving stack (reduced latency, improved throughput, or optimized cost per inference), validated the gains across our production deployment scenarios, and established the performance optimization roadmap that will guide infrastructure investment.
Drive lasting impact: At 12 months, you will have designed and deployed advanced serving architectures (disaggregated inference, optimized caching, speculative decoding) that meaningfully improve patient experience and operational efficiency, contributed novel optimization techniques that become part of our core infrastructure, and made our serving stack a durable competitive advantage in healthcare AI deployment.
THE TEAM
You'll work alongside systems engineers, ML researchers, and infrastructure experts who are obsessed with making AI systems fast, reliable, and cost-effective. This is a team that values deep technical rigor, continuous benchmarking, and solving hard systems problems that have real impact on patient experience and business unit economics.
WHAT YOU'LL DO
- Design and implement multi-node serving architectures for distributed LLM inference
- Optimize multi-LoRA serving systems
- Apply advanced quantization techniques (FP4/FP6) to reduce model footprint while preserving quality
- Implement speculative decoding and other latency optimization strategies
- Develop disaggregated serving solutions with optimized caching strategies for prefill and decoding phases
- Continuously benchmark and improve system performance across various deployment scenarios and GPU types
LOCATION REQUIREMENT
We believe the best ideas happen together. This role is based in our Menlo Park, California office, expected to be five days a week. We're also exploring establishing a presence in the Bellevue area—if that develops, flexibility on location may be available for exceptional candidates.
COMPENSATION
Compensation is based on experience, expertise, and level of responsibility. We offer competitive packages that reflect the seniority and scope of the role, along with equity, health insurance, and other benefits.
WHAT YOU BRING
MUST-HAVE:
- Experience optimizing LLM inference systems at scale
- Proven expertise with distributed serving architectures for large language models
- Hands-on experience implementing quantization techniques for transformer models
- Strong understanding of modern inference optimization methods, including:
- Speculative decoding techniques with draft models
- Eagle speculative decoding approaches
- Proficiency in Python and C++
- Experience with CUDA programming and GPU optimization
NICE-TO-HAVE:
- Contributions to open-source inference frameworks such as vLLM, SGLang, or TensorRT-LLM
- Experience with custom CUDA kernels
- Track record of deploying inference systems in production environments
- Deep understanding of performance optimization systems
Show us what you've built: Tell us about an LLM inference or training project that makes you proud! Whether you've optimized inference pipelines to achieve breakthrough performance, designed innovative training techniques, or built systems that scale to billions of parameters - we want to hear your story.
Open source contributor? Even better! If you've contributed to projects like vllm, sglang, lmdeploy or similar LLM optimization frameworks, we'd love to see your PRs. Your contributions to these communities demonstrate exactly the kind of collaborative innovation we value.
Join a team where your expertise won't just be appreciated—it will be celebrated and amplified. Help us shape the future of AI deployment at scale!
References
1. Polaris: A Safety-focused LLM Constellation Architecture for Healthcare, https://arxiv.org/abs/2403.13313
2 https://arxiv.org/abs/2403.133132. Polaris 2: https://www.hippocraticai.com/polaris2
3 https://www.hippocraticai.com/polaris23. Personalized Interactions: https://www.hippocraticai.com/personalized-interactions
4 https://www.hippocraticai.com/personalized-interactions4. Human Touch in AI: https://www.hippocraticai.com/the-human-touch-in-ai
5 https://www.hippocraticai.com/the-human-touch-in-ai5. Empathetic Intelligence: https://www.hippocraticai.com/empathetic-intelligence https:/
Similar Jobs
Related searches:
On-site Jobs
Mid-Level Jobs
On-site Mid-Level Jobs
Mid-Level AI InfrastructureMid-Level Fintech & Payments AIMid-Level NLP & Language AIMid-Level AI ResearchMid-Level Healthcare AIMid-Level Generative AIMid-Level Machine Learning
AI Jobs in Menlo Park
AI Infrastructure in Menlo ParkFintech & Payments AI in Menlo ParkNLP & Language AI in Menlo ParkAI Research in Menlo ParkHealthcare AI in Menlo ParkGenerative AI in Menlo ParkMachine Learning in Menlo Park
llmgpuhealthcarepaymentsfine-tuninginferenceresearch
Get jobs like this delivered weekly
Free AI jobs newsletter. No spam.