ML Infrastructure Engineer
full-time
senior
Posted 22 hours ago
Before you apply
Build my evidence-backed draft — free Apply on company site →Paste your relevant resume section or 2–4 true bullets. See supported requirements and honest gaps. No account and no application sent.
About this role
About Zipline
Zipline is the world’s largest and most experienced drone delivery service. We are on a mission to serve all humans equally by ensuring access to food, medicine and essential goods anytime, anywhere. We design, build, and operate the world’s largest autonomous logistics system, delivering critical supplies quickly and reliably. Today, Zipline operates on four continents, makes a delivery somewhere in the world every 30 seconds, and has completed millions of deliveries to date, including blood, vaccines, medical supplies, food, and retail products.
Our customers include the world’s largest and most prominent healthcare systems, governments, retailers, restaurants and global businesses who rely on us to save lives, reduce emissions, increase economic opportunity, and provide delivery from point A to point B as fast as possible. The drone is only 15% of what we’ve built to enable seamless, reliable, global operations.
Our system strengthens supply chains, reduces congestion, and gives people time back. With more than 140 million commercial autonomous miles safely flown, Zipline is redefining access to healthcare, consumer products, and food across the globe.
We operate at a global scale and are looking for practical problem solvers who thrive on real-world challenges and rapid growth. Our team is motivated by building systems that have a direct, meaningful impact on people’s lives and by scaling the future of logistics. We are seeking people who sculpt from first principles, enjoy facing adversity, and can do the impossible at record breaking speeds.
About You and The Role
As an ML Training & Inference Infrastructure Engineer on the Data Platform team you will be building and scaling the systems powering our data flywheel. This person will work at the intersection of autonomy and the infrastructure, owning systems that make ML development faster, reproducible, observable, and safe.
This role is for a strong software engineer who enjoys the full ML development cycle: data ingestion, processing pipelines, dataset management, distributed training, continuous model integration, evaluation, and deployment. The ideal candidate has strong production engineering habits and is excited to build infrastructure that helps real autonomous systems improve over time.
What You'll Do
Build and operate software infrastructure that enables learning algorithms to leverage Zipline’s large-scale (quickly growing!) fleet data.
Design scalable, maintainable data and ML infrastructure for autonomy teams, including dataset creation, validation, training, evaluation, and deployment.
Own and improve data pipelines that feed into the ML development loop.
Identify and mitigate bottlenecks in the ML development cycle, especially around orchestration, performance, and reproducibility to increase the rate at which we can improve and scale the delivery experience.
What You'll Bring
3+ years of professional software engineering experience, ideally including ML infrastructure, data infrastructure, robotics, autonomy, aerospace, medical devices, or another safety-critical hardware/product environment.
Strong software engineering practices in Python in a production setting; comfort designing APIs, services, schemas, jobs, and operational workflows.
Experience building reproducible data pipelines and machine-learning pipelines.
Experience monitoring data statistics, system performance metrics, pipeline failures, and model/evaluation signals.
Working knowledge of ML concepts such as datasets, training, evaluation, optimization, statistics, and modern deep learning workflows.
Generalist mindset and willingness to work across cloud services, data platforms, developer tooling, and embedded/robotics-adjacent constraints.
Experience with PyTorch or similar ML frameworks.
Strong ownership, clear communication, and interest in building secure systems for mission-critical workflows.
Experience with Kubernetes or other container orchestration systems for production workloads.
Experience with cloud and on-premise production infrastructure, preferably AWS, and infrastructure-as-code tools such as Terraform or CloudFormation.
BONUS POINTS
Experience deploying or evaluating ML systems on real robots, autonomous vehicles, drones, or other hardware products.
Experience with large-scale training systems, feature stores, data/versioned artifact stores, model registries, or experiment tracking.
Experience with annotation systems, dataset inspection tooling, or active-learning workflows.
What Else You Need To Know
The starting cash ranges for this role is $160,000 - $250,000. Please note that this is a target, starting cash range for a candidate who meets the minimum qualifications for this role. The final cash pay for this role will depend on a variety of factors, including a specific candidate's experience, qualifications, skills, working location, and projected impact. The total compens
Similar Jobs
Related searches:
On-site Jobs
Senior Jobs
On-site Senior Jobs
Senior Computer VisionSenior Robotics & AutonomySenior Backend & SystemsSenior Healthcare AISenior AI InfrastructureSenior Data EngineeringSenior Machine Learning
AI Jobs in San Francisco
Computer Vision in San FranciscoRobotics & Autonomy in San FranciscoBackend & Systems in San FranciscoHealthcare AI in San FranciscoAI Infrastructure in San FranciscoData Engineering in San FranciscoMachine Learning in San Francisco
cloudpytorchhealthcareroboticsdata-pipelineautonomous-vehiclesdeep-learningdistributed-systems
Get jobs like this delivered weekly
Free AI jobs newsletter. No spam.