Staff/Principal DevOps Engineer, AI Inference
full-time
principal
Posted 1 day ago
Apply Now
Stand out: build a proof-of-work pitch →
Free GitHub-based preview. Direct apply stays one click away.
Get weekly job alerts like this →Hiring for this role?
AI Market Demand Pack · $29 one-time
Compare this role's skills with the full AI hiring market. Get ranked demand, salary bands, leading companies, public source URLs, and a decision brief.
About this role
Your Impact at LILA
The Staff/Principal DevOps Engineer - AI Inference will drive the design, implementation, and optimization of infrastructure purpose-built for serving machine learning models at scale. This role bridges platform engineering, site reliability, and ML infrastructure, building the systems that power low-latency, high-throughput inference across GPU clusters and cloud accelerators. You will collaborate with ML engineers, research scientists, and software engineers to build inference platforms that serve models reliably to production users while maximizing compute efficiency.
What You'll Be Building
GPU/accelerator infrastructure on Kubernetes: scheduling, resource isolation, multi-tenant GPU sharing, device plugins, and topology-aware placement for inference workloads
Model serving platforms using frameworks such as vLLM, Triton Inference Server, TGI, or custom serving stacks with optimized batching, caching, and request routing
Intelligent request routing and load balancing across heterogeneous accelerator fleets (NVIDIA GPUs, AWS Inferentia/Trainium) to maximize utilization and minimize latency
Autoscaling systems that dynamically match inference compute supply with demand across production, research, and experimental workloads
Production-grade deployment pipelines for ML models: canary rollouts, A/B testing, model versioning, and safe rollback across multi-region deployments
Infrastructure-as-code with Terraform and Helm for GPU-accelerated EKS clusters, including node pools, spot/on-demand strategies, and accelerator-specific networking
Observability and performance optimization: GPU utilization monitoring, inference latency profiling, token throughput dashboards, and SLO/SLI tracking for model endpoints
CI/CD pipelines for model artifacts: container image builds with CUDA/driver dependencies, model registry integration, and automated inference benchmarking in CI
AWS cloud infrastructure for ML: EKS with GPU node groups, EC2 accelerated instances (P4/P5, Inf2, Trn1), S3 model storage, EFA/high-bandwidth networking, and IAM least privilege
Cost optimization and capacity planning: right-sizing accelerator instances, spot instance strategies for inference, and fleet-wide efficiency reporting
What You'll Need to Succeed
Expertise in DevOps, SRE, or Platform Engineering with significant experience operating GPU/accelerator infrastructure at scale
Deep experience with Kubernetes for ML workloads: GPU scheduling, resource quotas, node affinity, and accelerator device management
Strong proficiency deploying to AWS using infrastructure-as-code (Terraform, Helm) with hands-on experience managing GPU-based compute (EKS, EC2 P-series/Inf/Trn instances)
Experience with model serving infrastructure: inference servers, request batching, KV-cache optimization, or LLM serving frameworks
Strong understanding of networking for distributed inference: high-bandwidth interconnects, NCCL, VPC/PrivateLink, and load balancing at L4/L7
Strong proficiency in Python for automation, tooling, and integration with ML frameworks
Bonus Points For
Experience with LLM inference optimization: continuous batching, speculative decoding, quantization (GPTQ, AWQ, FP8), tensor parallelism, and pipeline parallelism
Hands-on experience with multiple accelerator families (NVIDIA A100/H100, AWS Inferentia2, Trainium, AMD MI300X) and maintaining hardware-agnostic serving infrastructure
Multi-region deployment experience with geographic routing and failover for latency-sensitive inference endpoints
Proficiency in Rust or Go for performance-critical infrastructure components
SRE practices for ML systems: chaos engineering on GPU workloads, incident management, capacity modeling for bursty inference traffic
Experience with model registries, artifact versioning, and ML supply chain security
Observability platform expertise: building custom metrics for token-level throughput, time-to-first-token, and per-request GPU memory profiling
Prior startup/high-growth experience balancing velocity with reliability in rapidly scaling AI systems
Compensation
We offer competitive base compensation with bonus potential and generous early-stage equity. Your final offer will reflect your background, expertise, and expected impact.
U.S. Benefits. Full-time U.S. employees receive a comprehensive benefits program including medical, dental, and vision coverage; employer-paid life and disability insurance; flexible time off with generous company wide holidays; paid parental leave; an educational assistance program; commuter benefits, including bike share memberships for office based employees; and a company subsidized lunch program.
International Benefits. Full-time employees outside the U.S. receive a comprehensive benefits program tailored to their region. USD salary ranges apply only to U.S.-based positions; international salaries are set to local market.
Expected Base Salary Range
$192,000 — $272,000
Similar Jobs
Related searches:
Get jobs like this delivered weekly
Free AI jobs newsletter. No spam.