{"access":{"catalog_url":"https://aidevboard.com/api/v1/catalog","description":"Public read endpoints are open and free. API keys are optional for stable agent identity and keyed hourly throttling.","docs_url":"https://aidevboard.com/docs","employer_pilot_url":"https://aidevboard.com/verified-interview-pilot","mode":"open","register_url":"https://aidevboard.com/api/v1/register"},"candidate_resume_action":{"application_authorized":false,"candidate_charge":0,"endpoint":"https://aidevboard.com/api/v1/candidate/resume-preview","job_id_json_path":"jobs[].id","method":"POST","preview_requires_identity":false,"required_body_fields":["job_id","evidence_bullets"],"requires_explicit_human_review":true,"saved_artifact_protocol":"mcp","saved_artifact_requires_verified_human":true,"saved_artifact_tool":"compile_job_specific_resume","search_requires_identity":false,"status":"available_after_candidate_selects_job","submission_performed":false,"uses_candidate_verified_evidence":true},"degraded":false,"estimated":false,"has_next":true,"jobs":[{"id":"b040d81d-82bc-4198-a020-8cf6e2705157","company_id":"e8dfc4ee-9649-4fd0-9c16-90d38a1954e1","title":"Staff Machine Learning Scientist, Applied Causal Inference","slug":"staff-machine-learning-scientist-applied-causal-inference-4231759e","description":"About the Team \n DoorDash is building the next generation of causal decisioning systems for New Verticals: grocery, convenience, retail, alcohol, pets, flowers, and other emerging categories. These businesses operate in high-dimensional, messy marketplaces where every consumer, merchant, item, promotion, substitution, search result, and delivery promise creates a causal question.\n About the Role \n We are hiring a Causal Machine Learning Engineer to help build the causal ML foundation behind how DoorDash grows New Verticals. This is not a generic ML role with some experimentation work on the side. We are looking for someone who has built or deeply worked on production causal systems: uplift models, heterogeneous treatment effect models, surrogate metrics, experimentation platforms, counterfactual policy evaluation, promotion optimization, or marketplace decisioning systems.\n You will join a small, senior pod of causal ML and econometrics experts working across ML, Analytics, Product, and Engineering. The mandate is to build the causal spine for a large-scale consumer marketplace.\n You're excited about this opportunity because you will… \n \n Design, build, and productionize causal ML systems that influence real marketplace decisions across New Verticals.\n Build uplift / heterogeneous treatment effect models for consumer lifecycle value, promotions, retention, and reactivation.\n Develop counterfactual evaluation frameworks for ranking, recommendations, search, promotions, substitutions, and marketplace interventions.\n Build systems that connect experimentation, observational data, and ML decisioning so teams can make better tradeoffs when randomized experiments are slow, noisy, or incomplete.\n Design surrogate metrics and early indicators that help teams move faster while preserving long-term marketplace health.\n Partner with econometrics and analytics leaders to choose the right methods: doubly robust estimation, IV, diff-in-diff, synthetic controls, double ML, CUPED-style variance reduction, contextual bandits, off-policy evaluation, and related approaches.\n Translate causal models into production systems that can shape decisions in ranking, targeting, budget allocation, inventory-aware discovery, and consumer growth.\n Raise the bar for causal reasoning across ML teams: when to trust a model, when not to, and how to debug causal claims in a real marketplace.\n \n We're excited about you because you have… \n \n Deep practical experience with causal inference, econometrics, experimentation, or causal ML .\n Experience shipping models or decision systems in production, ideally in consumer marketplaces, ads, recommendations, search, pricing, promotions, logistics, fintech, or other high-scale settings.\n Strong judgment around the tradeoffs between randomized experiments, observational estimation, and model-based decisioning.\n Comfort debating and applying methods such as doubly robust estimation, double ML, IV, diff-in-diff, CUPED, uplift modeling, contextual bandits, and off-policy evaluation .\n Strong ML engineering ability: you can build reliable pipelines, train models, evaluate them rigorously, and partner with platform teams to put them into production.\n Strong product judgment: you can connect methods to business decisions, not just optimize offline metrics.\n The ability to operate across functions with ML engineers, economists, data scientists, product managers, and business leaders.\n Compensation \n The successful candidate’s starting pay will fall within the pay range listed below and is determined based on job-related factors including, but not limited to, skills, experience, qualifications, work location, and market conditions. Base salary is localized according to an employee’s work location. Ranges are market-dependent and may be modified in the future.\n In addition to base salary, the compensation for this role includes opportunities for equity grants. Talk to your recruiter for more information.\n DoorDash cares about you and your overall well-being. That’s why we offer a comprehensive benefits package to all regular employees, which includes a 401(k) plan with employer matching, 16 weeks of paid parental leave, wellness benefits, commuter benefits match, paid time off and paid sick leave in compliance with applicable laws (e.g. Colorado Healthy Families and Workplaces Act). DoorDash also offers medical, dental, and vision benefits, 11 paid holidays, disability and basic life insurance, family-forming assistance, and a mental health program, among others.\n To learn more about our benefits, visit our careers page here .\n See below for paid time off details:\n \n For salaried roles: flexible paid time off/vacation, plus 80 hours of paid sick time per year.\n For hourly roles: vacation accrued at about 1 hour for every 25.97 hours worked (e.g. about 6.7 hours/month if working 40 hours/week; about 3.4 hours/month if working 20 hours/week), and paid sick time accrued at 1 hour for every 30 hours worked","salary_min":203500,"salary_max":299300,"location":"San Francisco, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["fine-tuning","healthcare","cloud","payments","machine-learning","inference"],"apply_url":"https://job-boards.greenhouse.io/doordashusa/jobs/8140067","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-18T20:55:11Z","expires_at":"2026-09-29T13:49:21.704976Z","created_at":"2026-08-25T18:33:45.109934Z","updated_at":"2026-08-30T13:49:21.836357Z","company_name":"DoorDash","company_slug":"doordash","company_logo_url":"https://www.google.com/s2/favicons?domain=doordash.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/b040d81d-82bc-4198-a020-8cf6e2705157"},{"id":"e9ebce41-3ac8-47ba-a5ee-5e22e28ef4f1","company_id":"e8dfc4ee-9649-4fd0-9c16-90d38a1954e1","title":"Staff Machine Learning Engineer, Causal Inference","slug":"staff-machine-learning-engineer-causal-inference-da3f0671","description":"About the Team \n DoorDash is building the next generation of causal decisioning systems for New Verticals: grocery, convenience, retail, alcohol, pets, flowers, and other emerging categories. These businesses operate in high-dimensional, messy marketplaces where every consumer, merchant, item, promotion, substitution, search result, and delivery promise creates a causal question.\n About the Role \n We are hiring a Causal Machine Learning Engineer to help build the causal ML foundation behind how DoorDash grows New Verticals. We are looking for someone who has built or deeply worked on production causal systems: uplift models, heterogeneous treatment effect models, surrogate metrics, experimentation platforms, counterfactual policy evaluation, promotion optimization, or marketplace decisioning systems.\n You will join a small, senior pod of causal ML and econometrics experts working across ML, Analytics, Product, and Engineering. The mandate is to build the causal spine for a large-scale consumer marketplace.\n You're excited about this opportunity because you will… \n \n Design, build, and productionize causal ML systems that influence real marketplace decisions across New Verticals.\n Build uplift / heterogeneous treatment effect models for consumer lifecycle value, promotions, retention, and reactivation.\n Develop counterfactual evaluation frameworks for ranking, recommendations, search, promotions, substitutions, and marketplace interventions.\n Build systems that connect experimentation, observational data, and ML decisioning so teams can make better tradeoffs when randomized experiments are slow, noisy, or incomplete.\n Design surrogate metrics and early indicators that help teams move faster while preserving long-term marketplace health.\n Partner with econometrics and analytics leaders to choose the right methods: doubly robust estimation, IV, diff-in-diff, synthetic controls, double ML, CUPED-style variance reduction, contextual bandits, off-policy evaluation, and related approaches.\n Translate causal models into production systems that can shape decisions in ranking, targeting, budget allocation, inventory-aware discovery, and consumer growth.\n Raise the bar for causal reasoning across ML teams: when to trust a model, when not to, and how to debug causal claims in a real marketplace.\n \n We're excited about you because you have… \n \n Deep practical experience with causal inference, econometrics, experimentation, or causal ML .\n Experience shipping models or decision systems in production, ideally in consumer marketplaces, ads, recommendations, search, pricing, promotions, logistics, fintech, or other high-scale settings.\n Strong judgment around the tradeoffs between randomized experiments, observational estimation, and model-based decisioning.\n Comfort debating and applying methods such as doubly robust estimation, double ML, IV, diff-in-diff, CUPED, uplift modeling, contextual bandits, and off-policy evaluation .\n Strong ML engineering ability: you can build reliable pipelines, train models, evaluate them rigorously, and partner with platform teams to put them into production.\n Strong product judgment: you can connect methods to business decisions, not just optimize offline metrics.\n The ability to operate across functions with ML engineers, economists, data scientists, product managers, and business leaders.\n Compensation \n The successful candidate’s starting pay will fall within the pay range listed below and is determined based on job-related factors including, but not limited to, skills, experience, qualifications, work location, and market conditions. Base salary is localized according to an employee’s work location. Ranges are market-dependent and may be modified in the future.\n In addition to base salary, the compensation for this role includes opportunities for equity grants. Talk to your recruiter for more information.\n DoorDash cares about you and your overall well-being. That’s why we offer a comprehensive benefits package to all regular employees, which includes a 401(k) plan with employer matching, 16 weeks of paid parental leave, wellness benefits, commuter benefits match, paid time off and paid sick leave in compliance with applicable laws (e.g. Colorado Healthy Families and Workplaces Act). DoorDash also offers medical, dental, and vision benefits, 11 paid holidays, disability and basic life insurance, family-forming assistance, and a mental health program, among others.\n To learn more about our benefits, visit our careers page here .\n See below for paid time off details:\n \n For salaried roles: flexible paid time off/vacation, plus 80 hours of paid sick time per year.\n For hourly roles: vacation accrued at about 1 hour for every 25.97 hours worked (e.g. about 6.7 hours/month if working 40 hours/week; about 3.4 hours/month if working 20 hours/week), and paid sick time accrued at 1 hour for every 30 hours worked (e.g. about 5.8 hours/month if working 40 hours/week; about 2.9 hours/mon","salary_min":203500,"salary_max":299300,"location":"San Francisco, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["cloud","healthcare","payments","fine-tuning","machine-learning","inference"],"apply_url":"https://job-boards.greenhouse.io/doordashusa/jobs/8139942","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-18T20:55:10Z","expires_at":"2026-09-29T13:49:21.417077Z","created_at":"2026-08-25T18:33:45.096852Z","updated_at":"2026-08-30T13:49:21.548155Z","company_name":"DoorDash","company_slug":"doordash","company_logo_url":"https://www.google.com/s2/favicons?domain=doordash.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/e9ebce41-3ac8-47ba-a5ee-5e22e28ef4f1"},{"id":"e21299fc-4a21-4671-a661-ef416117aba6","company_id":"c93e0284-9c76-4a85-9905-494865ab9278","title":"Inference Systems Performance Architect","slug":"inference-systems-performance-architect-dcf7b6cc","description":"The era of pervasive AI has arrived. In this era, organizations will use generative AI to unlock hidden value in their data, accelerate processes, reduce costs, drive efficiency and innovation to fundamentally transform their businesses and operations at scale. \n SambaNova Suite™ is the first full-stack, generative AI platform, from chip to model, optimized for enterprise and government organizations. Powered by the intelligent SN40L chip, the SambaNova Suite is a fully integrated platform, delivered on-premises or in the cloud, combined with state-of-the-art open-source models that can be easily and securely fine-tuned using customer data for greater accuracy. Once adapted with customer data, customers retain model ownership in perpetuity, so they can turn generative AI into one of their most valuable assets. \n About the role \n As an Architect on the Inference Systems Performance team, you'll own the discipline of end-to-end performance for large-scale LLM inference at SambaNova, from how a request moves through tokenization, prefill, decode, and the fabric between them, to how an entire deployment is sized against customer SLOs. Inference-systems performance is a nascent field; the results of design choices are being discovered daily rather than inherited from a mature craft, and this role exists to bring rigor to that frontier.\n The work spans two coupled pillars.\n The first is reproducible workload capture and benchmarking -- building faithful, replayable representations of real and increasingly agentic traffic, so that what we measure reflects production rather than an artifact of a naive load script. \n The second is performance modeling and simulation - analytic and simulation models that turn measurement into a \"what-if\" capability, letting us reason about configurations and hardware that do not exist yet.\n Together these feed both today's serving optimization and the next generation of system planning.\n The technical frontier you'll help define is heterogeneous, disaggregated inference - GPU on prefill, the RDU on decode - which explores hard problems across networking, storage, prompt caching, and tail-latency-bound data movement.\n You will be the go-to person for inference-systems performance across SambaNova, and a resource the entire organization relies on to answer \"how fast can this go, and what will it take.\"\n Responsibilities \n \n Define and drive the technical strategy for inference-systems performance including workload capture, benchmarking, modeling, and simulation, while developing and architecture that enables many potential futures\n Build the workload-capture and agentic-benchmarking capability - capture representative production traffic and enforce the discipline of interrogating results, spotting artificial contention or misleadingly high cache-hit rates that never occur in real use\n Own the performance-modeling and simulation practice - models that predict how a configuration change moves the output, informing capacity planning against customer SLOs and next-generation system and hardware planning\n Attack the end-to-end profiling gap - drive tooling that produces accurate, actionable profiles of a distributed inference pipeline so bottlenecks can be localized across host, accelerator, and fabric\n Serve as the senior technical voice across model-optimization, systems, hardware, and product, tying together multiple engineering activities and teams, and weighing trade-offs of reliability, scalability, operational cost, and ease of adoption\n Act as a resource for the entire organization including representing SambaNova's performance story to customers and partners\n Mentor and multiply by raising the capability of principal and senior engineers, building the systems, tools, and patterns that make everyone more productive\n Drive the resolution of the most ambiguous, novel challenges that span organizational boundaries or have no established answer in the field yet\n \n Required qualifications \n \n 12+ years of experience in performance engineering, with a demonstrated record of technical leadership on large-scale, complex systems\n Deep expertise in end-to-end performance analysis of distributed systems with many moving parts and the ability to localize bottlenecks that others cannot\n Proven command of realistic workload generation and simulation and of performance modeling, including calibrating models against real, variable workloads\n Demonstrated ability to enter an unfamiliar domain and apply core performance methods with transferable discipline expertise \n Ability to lead cross-functional efforts, mentor senior engineers, and influence organizational direction\n Experience representing an organizations credibly to customers and partners\n Track record of independently scoping and delivering high-complexity, high-ambiguity work with significant impact on products or roadmap\n \n Preferred qualifications \n \n Direct experience with LLM inference serving - continuous batching, prompt/KV cach","salary_min":245000,"salary_max":325000,"location":"San Jose, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"principal","tags":["distributed-systems","generative-ai","agents","llm","inference"],"apply_url":"https://sambanova.ai/sambanova-available-positions/?gh_jid=6144835004","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-16T16:59:17Z","expires_at":"2026-09-29T13:34:49.128051Z","created_at":"2026-08-25T18:27:21.848015Z","updated_at":"2026-08-30T13:34:49.267513Z","company_name":"SambaNova Systems","company_slug":"sambanova","company_logo_url":"https://www.google.com/s2/favicons?domain=sambanova.ai\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/e21299fc-4a21-4671-a661-ef416117aba6"},{"id":"6753c0ee-7501-493c-89c5-8fdeebce90f8","company_id":"74257563-5513-4a8d-a0f7-01f00c59aed6","title":"Senior Data Scientist - Payments (Inference)","slug":"senior-data-scientist-payments-inference-a5421a89","description":"Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. \n The Community You Will Join: \n You will join the Payments Data Science organization, which sits at the intersection of Trust and Payments and powers the systems that move money safely and efficiently across Airbnb's global marketplace. The team spans payment optimization for guests and hosts, fraud and risk mitigation, complex measurement, and regulatory compliance. We partner directly with Payments product and engineering leadership, Finance, and Trust to ensure every transaction is fast, safe, and compliant at global scale. Our work directly shapes decisions made by senior leaders, including Payments executive leadership, and requires a rigorous, evidence-based approach to every recommendation we make. Our Data Science team enables this mission by providing reliable measurement frameworks to deliver robust data insights, build and enable state-of-the-art data products/models, and provide actionable and reliable business guidance.\n The Difference You Will Make: \n We are looking for a passionate data scientist to lead quantitative measurement efforts and bring novel scientific approaches to drive decision making across our platform’s payment experience. This data scientist will perform careful hypothesis generation, causal inference framework development, and model development/evaluation to ideate and drive payment strategies on our platform. This role will have a particular focus on payments fraud mitigation and loss optimization, with the goal of making our platform safer for our community. Our Data Scientists have a deep understanding of causal framework development, statistical analysis, machine learning model development and evaluation strategies, and the complications of running experiment/quasi-experimental methods in a two-sided marketplace. They have keen business sense and are able to develop novel solutions to fraud and risk problems that don't have an established playbook and utilize their findings to communicate across a wide range of partners to drive our data \u0026 product roadmaps. They are not only the trusted data expert on their team, but also a storyteller.\n Examples of projects you may work on include, development of novel metrics and frameworks that can efficiently measure outcomes (often balancing competing tradeoffs), generating deep root cause investigations and long term impact measurements, and building/evaluating ML and agentic models to optimize guest, host, and business outcomes.\n A Typical Day:  \n \n Inference: Develop and apply causal inference methods, including experimental, econometric regressions, and quasi-experimental methods to measure a wide-range of platform/product impacts.\n AI/ML: Build methods for robust evaluation of ML/AI model efficiency and performance. Ability to identify use-cases for and develop predictive models to classify, segment, and interpret our users’ behavior. Support evaluation and optimization of agentic and LLM-based systems. \n Optimization:  Develop methodologies to explore/simulate the impact of new interventions and develop data products to optimize product/operational strategies.\n Communication:  Deliver robust research reports and effective data visualizations. Collaborate with and present to stakeholders to identify opportunities and communicate findings, and drive impact.\n Empowerment:  Think strategically about opportunities to improve and scale our brand measurement and customer insights.\n \n Your Expertise: \n \n 5+ years of industry experience in a quantitative analysis role with a Master’s degree in a quantitative field (math / economics / statistics, and etc.), or 3+ years of experience with a Phd degree.\n Strong knowledge of causal inference, experimentation, applied statistical modeling, and end-to-end ML development.\n Skilled in statistical programming (Python or R) and database usage (SQL)\n Demonstrated track record of owning a business or technical domain end-to-end at a prior company: setting your own roadmap, being the accountable expert others escalate to, and driving a problem to resolution.\n Proven ability to communicate clearly and effectively to audiences of varying technical levels\n Ability to work independently, set your own roadmap, and drive cross-functional alignment\n Payments Fraud/Risk Domain expertise is a strong plus.\n Familiarity with evaluating agentic or LLM-based systems (e.g., decision-quality measurement, human-in-the-loop calibration) is a plus.\n \n Your Location: \n This position is US - Remote Eligible. The role may include occasional work at an Airbnb office or attendance at offsites, as agreed to with your manager. While the position is Remote ","salary_min":179000,"salary_max":210000,"location":"Remote (US)","workplace":"remote","remote_scope":"restricted","job_type":"full-time","experience_level":"senior","tags":["llm","agents","payments","data-science","inference"],"apply_url":"https://careers.airbnb.com/positions/8123037?gh_jid=8123037","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-14T17:25:31Z","expires_at":"2026-09-29T13:39:39.386068Z","created_at":"2026-08-25T18:29:15.498252Z","updated_at":"2026-08-30T13:39:39.528611Z","company_name":"Airbnb","company_slug":"airbnb","company_logo_url":"https://www.google.com/s2/favicons?domain=airbnb.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/6753c0ee-7501-493c-89c5-8fdeebce90f8"},{"id":"6109f5c9-3808-4bdf-a225-316edd299a6b","company_id":"6734f15a-40ed-4186-ae4a-d774c655ae58","title":"Staff/Principal DevOps Engineer, AI Inference","slug":"staffprincipal-devops-engineer-ai-inference-09cb28c0","description":"Your Impact at LILA \n The Staff/Principal DevOps Engineer - AI Inference will drive the design, implementation, and optimization of infrastructure purpose-built for serving machine learning models at scale. This role bridges platform engineering, site reliability, and ML infrastructure, building the systems that power low-latency, high-throughput inference across GPU clusters and cloud accelerators. You will collaborate with ML engineers, research scientists, and software engineers to build inference platforms that serve models reliably to production users while maximizing compute efficiency.\n What You'll Be Building \n \n GPU/accelerator infrastructure on Kubernetes: scheduling, resource isolation, multi-tenant GPU sharing, device plugins, and topology-aware placement for inference workloads\n Model serving platforms using frameworks such as vLLM, Triton Inference Server, TGI, or custom serving stacks with optimized batching, caching, and request routing\n Intelligent request routing and load balancing across heterogeneous accelerator fleets (NVIDIA GPUs, AWS Inferentia/Trainium) to maximize utilization and minimize latency\n Autoscaling systems that dynamically match inference compute supply with demand across production, research, and experimental workloads\n Production-grade deployment pipelines for ML models: canary rollouts, A/B testing, model versioning, and safe rollback across multi-region deployments\n Infrastructure-as-code with Terraform and Helm for GPU-accelerated EKS clusters, including node pools, spot/on-demand strategies, and accelerator-specific networking\n Observability and performance optimization: GPU utilization monitoring, inference latency profiling, token throughput dashboards, and SLO/SLI tracking for model endpoints\n CI/CD pipelines for model artifacts: container image builds with CUDA/driver dependencies, model registry integration, and automated inference benchmarking in CI\n AWS cloud infrastructure for ML: EKS with GPU node groups, EC2 accelerated instances (P4/P5, Inf2, Trn1), S3 model storage, EFA/high-bandwidth networking, and IAM least privilege\n Cost optimization and capacity planning: right-sizing accelerator instances, spot instance strategies for inference, and fleet-wide efficiency reporting\n \n What You'll Need to Succeed \n \n Expertise in DevOps, SRE, or Platform Engineering with significant experience operating GPU/accelerator infrastructure at scale\n Deep experience with Kubernetes for ML workloads: GPU scheduling, resource quotas, node affinity, and accelerator device management\n Strong proficiency deploying to AWS using infrastructure-as-code (Terraform, Helm) with hands-on experience managing GPU-based compute (EKS, EC2 P-series/Inf/Trn instances)\n Experience with model serving infrastructure: inference servers, request batching, KV-cache optimization, or LLM serving frameworks\n Strong understanding of networking for distributed inference: high-bandwidth interconnects, NCCL, VPC/PrivateLink, and load balancing at L4/L7\n Strong proficiency in Python for automation, tooling, and integration with ML frameworks\n \n Bonus Points For \n \n Experience with LLM inference optimization: continuous batching, speculative decoding, quantization (GPTQ, AWQ, FP8), tensor parallelism, and pipeline parallelism\n Hands-on experience with multiple accelerator families (NVIDIA A100/H100, AWS Inferentia2, Trainium, AMD MI300X) and maintaining hardware-agnostic serving infrastructure\n Multi-region deployment experience with geographic routing and failover for latency-sensitive inference endpoints\n Proficiency in Rust or Go for performance-critical infrastructure components\n SRE practices for ML systems: chaos engineering on GPU workloads, incident management, capacity modeling for bursty inference traffic\n Experience with model registries, artifact versioning, and ML supply chain security\n Observability platform expertise: building custom metrics for token-level throughput, time-to-first-token, and per-request GPU memory profiling\n Prior startup/high-growth experience balancing velocity with reliability in rapidly scaling AI systems\n \n \n Compensation \n We offer competitive base compensation with bonus potential and generous early-stage equity. Your final offer will reflect your background, expertise, and expected impact.\n U.S. Benefits. Full-time U.S. employees receive a comprehensive benefits program including medical, dental, and vision coverage; employer-paid life and disability insurance; flexible time off with generous company wide holidays; paid parental leave; an educational assistance program; commuter benefits, including bike share memberships for office based employees; and a company subsidized lunch program.\n International Benefits. Full-time employees outside the U.S. receive a comprehensive benefits program tailored to their region. USD salary ranges apply only to U.S.-based positions; international salaries are set to local market.\n Expected Base Salary Range\n $192,000 — $272,000","salary_min":192000,"salary_max":272000,"location":"Boston, MA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"principal","tags":["llm","mlops","gpu","cloud","devops","inference"],"apply_url":"https://job-boards.greenhouse.io/lilasciences/jobs/4248032009","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-07-28T15:38:59Z","expires_at":"2026-09-29T13:48:28.77755Z","created_at":"2026-07-29T14:18:39.098522Z","updated_at":"2026-08-30T13:48:28.909371Z","company_name":"Lila Sciences","company_slug":"lila-sciences","company_logo_url":"https://www.google.com/s2/favicons?domain=lila.ai\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/6109f5c9-3808-4bdf-a225-316edd299a6b"},{"id":"8c1a812e-e15b-4f41-adec-d56b9a321557","company_id":"332b7698-676b-4a3e-8b02-81b1195c5af6","title":"Staff Software Engineer- Foundation Model Inference","slug":"staff-software-engineer-foundation-model-inference-0efd5575","description":"P-1930\n At Databricks, we are passionate about enabling data and AI teams to solve the world's toughest problems — from making the next mode of transportation a reality to accelerating the development of medical breakthroughs. We do this by building and running the world's best data and AI infrastructure platform so our customers can use deep data insights to improve their business. Founded by engineers — and customer-obsessed — we leap at every opportunity to solve technical challenges, from designing next-gen UI/UX for interfacing with data to scaling our services and infrastructure across millions of virtual machines. And we're only getting started.\n As part of the AI team, you'll build the platforms and products that power everything from data apps, AI agents, model training, model serving, and Vector Search. You'll be joining a high-agency, high-visibility team operating at the frontier of AI infrastructure — with deep ties to research, product, and real-world enterprise use cases. Databricks Mosaic AI is one of our fastest-growing businesses, helping thousands of our customers democratize AI within their organizations. We're building the products and infrastructure that power the next generation of AI.\n The Foundation Model Inference team is the backbone of Databricks’ generative AI capabilities. We build the infrastructure that enables our customers to serve, scale, and optimize frontier models with enterprise-grade reliability and performance. Our Foundation Model APIs provide a unified platform that gives customers access to LLMs with the governance, flexibility, and scalability required for enterprise production workloads.\n We are looking for high-agency engineers who are excited to work on powering model inference at enterprise scale.\n The impact you will have: \n \n Build LLM infrastructure powering large-scale inference workloads for customers through partner models (OpenAI, Anthropic, Gemini) and self-hosted models (Qwen, GPT-OSS, Llama)\n Improve reliability, latency, and efficiency of distributed AI workloads\n Collaborate with platform, infra, and ML teams to deliver seamless end-to-end experiences\n Shape how developers and data scientists build and interact with AI on Databricks\n \n What we look for: \n \n 8+ years of experience in backend or infrastructure engineering\n Experience with distributed systems, scalable APIs, or cloud-native infrastructure\n Experience with real-time serving, ML infrastructure, or GPU orchestration\n Familiarity with service-oriented architecture, deployment pipelines, and system observability\n \n Bonus points for: \n \n Exposure to platforms like SageMaker, Vertex AI, or Azure ML\n Contributions to OSS projects like MLflow, PyTorch, Ray, vLLM, SGLang\n Built developer platforms or internal tools supporting AI workflows\n  \n Pay Range Transparency \n Databricks is committed to fair and equitable compensation practices. The pay range(s) for this role is listed below and represents the expected salary range for non-commissionable roles or on-target earnings for commissionable roles.  Actual compensation packages are based on several factors that are unique to each candidate, including but not limited to job-related skills, depth of experience, relevant certifications and training, and specific work location. Based on the factors above, Databricks anticipates utilizing the full width of the range. The total compensation package for this position may also include eligibility for annual performance bonus, equity, and the benefits listed above. For more information regarding which range your location is in visit our page here . \n  \n Local Pay Range\n $190,000 — $265,000 USD \n About Databricks \n Databricks is the Data and AI company. More than 20,000 organizations worldwide — including adidas, AT\u0026T, Bayer, Block, Mastercard, Rivian, Unilever, and 70% of the Fortune 500 — rely on the Databricks Data + AI Platform to build and scale data and AI apps, analytics and agents. Headquartered in San Francisco with 30+ offices around the globe, Databricks offers a unified platform that includes Genie, Lakebase, Agent Bricks, Lakeflow, Lakehouse, and Unity Catalog. To learn more, follow Databricks on LinkedIn , X , YouTube , and Instagram . Benefits At Databricks, we strive to provide comprehensive benefits and perks that meet the needs of all of our employees. For specific details on the benefits offered in your region click here . \n Our Commitment to Diversity and Inclusion \n At Databricks, we are committed to fostering a diverse and inclusive culture where everyone can excel. We take great care to ensure that our hiring practices are inclusive and meet equal employment opportunity standards. Individuals looking for employment at Databricks are considered without regard to age, color, disability, ethnicity, family or marital status, gender identity or expression, language, national origin, physical and mental ability, political affiliation, race, religion, sexual orientation","salary_min":190000,"salary_max":265000,"location":"San Francisco, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["llm","pytorch","mlops","cloud","distributed-systems","agents","generative-ai","inference"],"apply_url":"https://databricks.com/company/careers/open-positions/job?gh_jid=8649279002","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-07-24T16:49:01Z","expires_at":"2026-09-29T13:32:35.999517Z","created_at":"2026-07-25T14:02:34.398709Z","updated_at":"2026-08-30T13:32:36.14272Z","company_name":"Databricks","company_slug":"databricks","company_logo_url":"https://www.google.com/s2/favicons?domain=databricks.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/8c1a812e-e15b-4f41-adec-d56b9a321557"},{"id":"15d0b3ec-2358-4c4d-b3e1-146969722a5b","company_id":"2721f049-2cf2-4e3e-82d0-8d8df89c8f90","title":"Senior Machine Learning Engineer, LLM Inference Optimization","slug":"senior-machine-learning-engineer-llm-inference-optimization-5eb66d9e","description":"About Nebius: \n Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.\n Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.\n Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R\u0026D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R\u0026D.\n The role   \n Nebius Token Factory is building an AI training and model post-training capability for frontier model improvement. This role owns the infrastructure that makes large-scale training and RL experiments possible, reliable, reproducible, and efficient. The work sits at the intersection of distributed systems, GPU performance, model training frameworks, RL pipelines, and production engineering. \n A Senior MLE owns substantial model and endpoint optimization projects end to end. They are deeply hands-on, can debug difficult serving problems independently, and can deliver measurable improvements without needing heavy supervision. \n Your responsibilities :   \n \n \n Own optimization work for specific model families, customer endpoints, or serving backends.\n \n Run engine comparisons and recommend practical serving configurations for specific workloads.\n \n Debug model quality or performance regressions during production rollouts.\n \n Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, and cost per token.\n \n Deploy, configure, benchmark, and extend inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, or similar systems.\n \n Build and productionize model-compression workflows, including quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery.\n \n Implement or integrate speculative decoding, draft-model approaches, KV -cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving.\n \n Build reproducible benchmark harnesses for TTFT , TPOT , tokens per second per GPU, p95/p99 latency, GPU memory, reliability, and cost per token.\n \n Partner with GPU kernel engineers and platform engineers to diagnose bottlenecks across model code, kernels, runtime, scheduler, gateway, and cluster layers.\n \n Write clear design docs, performance reports, rollout plans, and customer-facing technical explanations.\n \n Must-haves :   \n \n \n Strong Python and PyTorch engineering skills.\n \n Hands-on experience deploying or optimizing LLM, VLM , or high-throughput transformer inference systems.\n \n Practical knowledge of at least one modern inference stack such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, or equivalent internal systems.\n \n Strong understanding of transformer inference bottlenecks, including KV cache, attention, memory bandwidth, batching, parallelism, and long-context serving.\n \n Ability to reason quantitatively about latency, throughput, quality, utilization, and cost tradeoffs.\n \n Strong communication skills and ability to collaborate with research, kernel, infrastructure, product, and customer teams.\n \n Nice - to - have s :   \n \n \n Experience with quantization-aware training, post-training quantization, FP8 , INT8 , INT4 , NVFP4 , MXFP4 , AWQ , GPTQ , SmoothQuant, or related techniques.\n \n Experience with distillation, speculative decoding, EAGLE, Medusa, multi-token prediction, or other inference acceleration methods.\n \n Experience with agentic workloads, including tool calling, structured outputs, streaming APIs, high concurrency, and multi-step orchestration.\n \n CUDA or Triton familiarity, even if the role is not primarily a kernel-engineering role.\n \n Open-source contributions to vLLM, SGLang, TensorRT-LLM, FlashInfer, LMCache, PyTorch, Triton, Ray, KServe, or related projects.\n \n Key employee benefits in the US: \n \n \n Health insurance:  100% company-paid medical, dental, and vision coverage for employees and families.\n \n 401(k) plan:  Up to 4% company match with immediate vesting.\n \n Parental leave:  20 weeks paid for primary caregivers, 12 weeks for secondary caregivers.\n \n Remote work reimbursement:  Up to $85/month for mobile and internet.\n \n Disability \u0026 life insurance : Company-paid short-term, long-term and life insurance coverage.\n \n  \n #LI-BH3 \n  \n Pay Transparency \n We offer competitive compensation and benefits packages. Actual compensation will be determined based on job-related factors, including experience, skills, qualifications, the level at which the candidate is hired, and geographic location, consistent with applicable law.","salary_min":195200,"salary_max":262200,"location":"Palo Alto, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"senior","tags":["agents","llm","cloud","gpu","distributed-systems","pytorch","machine-learning","inference"],"apply_url":"https://careers.nebius.com/?gh_jid=4921522101","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-07-22T22:54:04Z","expires_at":"2026-09-29T13:45:08.012855Z","created_at":"2026-07-24T14:15:06.696498Z","updated_at":"2026-08-30T13:45:08.14846Z","company_name":"Nebius","company_slug":"nebius","company_logo_url":"https://www.google.com/s2/favicons?domain=nebius.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/15d0b3ec-2358-4c4d-b3e1-146969722a5b"},{"id":"f2f0947c-9182-4d73-b1c8-ab97cc666183","company_id":"2721f049-2cf2-4e3e-82d0-8d8df89c8f90","title":"Senior Applied Scientist, Efficient LLM Inference \u0026 Model Optimization","slug":"senior-applied-scientist-efficient-llm-inference-model-optimization-0c77e3de","description":"About Nebius: \n Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.\n Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.\n Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R\u0026D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R\u0026D.\n The role   \n Nebius Token Factory needs scientists who can turn frontier inference bottlenecks into research problems, publish credible work, and then help ship the results into production. This is not a papers-only research role. The Applied Scientist is expected to design rigorous experiments, write strong code, collaborate with engineers, and convert research into deployed inference capabilities. \n A Senior Applied Scientist owns well-scoped research and production optimization projects. They can publish or prepare high-quality technical work while also producing code, experiments, and prototypes that engineers can use.\n Your responsibilities :   \n \n \n Own focused research projects from hypothesis through experiment, ablation, prototype, and production handoff.\n \n Prepare internal reports, technical blogs, or papers when the work is externally credible.\n \n Partner directly with MLEs to ensure research prototypes become usable production components.\n \n Define and execute research programs in efficient LLM and VLM inference with measurable production impact.\n \n Invent, evaluate, and productionize methods for quantization, QAT , distillation, speculative decoding, KV -cache reuse, KV -cache compression, long-context inference, MoE routing, and model/runtime co-optimization.\n \n Build high-quality prototypes in PyTorch, Triton, CUDA -adjacent tooling, or inference-serving frameworks, then work with MLEs and platform engineers to productionize them.\n \n Design rigorous evaluation methodology covering quality, latency, throughput, numerical stability, memory footprint, tail latency, and cost per token.\n \n Publish papers, technical reports, blog posts, and open-source artifacts that build external credibility for Nebius Token Factory.\n \n Collaborate with MLE, GPU kernel, backend infrastructure, product, and customer teams to choose high-leverage research bets.\n \n Mentor engineers and scientists on experimental design, scientific rigor, and model/system tradeoffs.\n \n Must-haves :   \n \n \n PhD in computer science, machine learning, ML systems, computer systems, computer architecture, electrical engineering, applied math, or a closely related field.\n \n Strong publication record or equivalent research artifacts in ML, ML systems, efficient inference, model compression, quantization, distillation, serving systems, or related areas.\n \n Strong hands-on coding ability in Python and PyTorch; ability to move from idea to experiment to prototype quickly.\n \n Deep understanding of LLMs, VLMs, transformer inference, decoding algorithms, model compression, quantization, and production-serving tradeoffs.\n \n Strong experimental design skills, including ablations, baselines, metrics, statistical reasoning, and failure analysis.\n \n Excellent written and verbal communication.\n \n Nice - to - have s :   \n \n \n First-author publications in NeurIPS, ICML , ICLR , MLSys, ACL , EMNLP , ASPLOS , OSDI , SOSP , ISCA , HPCA , or comparable venues.\n \n Experience deploying ML models or inference optimizations in production.\n \n Experience with vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo, FlashAttention, FlashInfer, Triton, CUDA , or PyTorch internals.\n \n Experience with post-training, SFT , DPO , RLHF , RLAIF , preference optimization, or synthetic data generation when connected to inference quality or efficiency.\n \n Open-source research artifacts, widely used benchmarks, high-quality technical blogs, or invited talks in efficient AI systems.\n \n Key employee benefits in the US: \n \n \n Health insurance:  100% company-paid medical, dental, and vision coverage for employees and families.\n \n 401(k) plan:  Up to 4% company match with immediate vesting.\n \n Parental leave:  20 weeks paid for primary caregivers, 12 weeks for secondary caregivers.\n \n Remote work reimbursement:  Up to $85/month for mobile and internet.\n \n Disability \u0026 life insurance : Company-paid short-term, long-term and life insurance coverage.\n \n  \n Pay Transparency \n We offer competitive compensation and benefits packages. Actual compensation will be determined based on job-related factors, including experience, skills, qualifications, the level at which the candidate is hired, and geographic location, consistent with applicable law.\n B","salary_min":195200,"salary_max":262200,"location":"Palo Alto, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"senior","tags":["gpu","llm","reinforcement-learning","pytorch","nlp","cloud","inference"],"apply_url":"https://careers.nebius.com/?gh_jid=4921549101","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-07-22T22:53:48Z","expires_at":"2026-09-29T13:45:06.39248Z","created_at":"2026-07-24T14:15:05.082048Z","updated_at":"2026-08-30T13:45:06.526116Z","company_name":"Nebius","company_slug":"nebius","company_logo_url":"https://www.google.com/s2/favicons?domain=nebius.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/f2f0947c-9182-4d73-b1c8-ab97cc666183"},{"id":"4d8ed5b4-93f4-4766-aa31-2fc37f443b6c","company_id":"332b7698-676b-4a3e-8b02-81b1195c5af6","title":"Engineering Manager, Foundation Model Inference (FMAPI)","slug":"engineering-manager-foundation-model-inference-fmapi-26cd54fd","description":"Engineering Manager, Foundation Model Inference (FMAPI) \n RDQ427R519\n At Databricks, we are driven by a passion to empower data teams in tackling the world's most challenging problems — from revolutionizing transportation to accelerating medical innovations. Our mission is to build and operate the premier data and AI infrastructure platform, enabling our customers to leverage deep data insights for transformative business improvements. If you are passionate about advancing data and AI technologies and are eager to lead a cutting-edge product to success, we invite you to join our dynamic team at Databricks.\n The Foundation Model Inference team is the backbone of Databricks’ generative AI capabilities. We build the infrastructure that enables our customers to serve, scale, and optimize frontier models with enterprise-grade reliability and performance. Our Foundation Model APIs provide a unified platform that gives customers access to LLMs with the governance, flexibility, and scalability required for enterprise production workloads.\n Databricks is looking for an Engineering leader to lead part of our Foundation Model Inference organization. We are seeking a leader who can build strong teams, uphold a high technical bar, and drive execution on large multi-quarter initiatives in a fast-moving AI infrastructure environment.\n The impact you will have \n \n Lead and grow a team of talented, product-minded infrastructure Engineers working on Databricks Foundation Model Inference and Foundation Model API; The teams owns the systems powering large-scale inference workloads for customers through partner models (OpenAI, Anthropic, Gemini) and self-hosted models (Qwen, GPT-OSS, Llama).\n Help shape the roadmap for products spanning real-time inference, provisioned throughput, and batch inference use cases.\n Partner closely with product and engineering leadership to deliver the right capabilities with strong reliability, quality, and service health.\n Build an inclusive, high-performing team that attracts, develops, and retains exceptional engineers.\n Maintain a deep understanding of your team’s technical area and uphold a strong bar for architecture, implementation quality, and operational excellence.\n \n What we look for \n \n Experience managing and growing high-performing software engineering teams.\n Strong technical judgment in distributed systems, platform infrastructure, AI/ML infrastructure, or large-scale backend services, with the ability to maintain quality even in areas you have not worked on personally.\n A track record of delivering complex, multi-quarter engineering initiatives with high quality and predictable execution.\n Experience partnering effectively with product management and peer engineering teams to translate customer and business needs into roadmaps and shipped outcomes.\n Operational rigor, including experience with service health, incident response, postmortems, and continuous improvement.\n Excellent communication, coaching, and hiring skills.\n BS in Computer Science or related field. \n  \n Pay Range Transparency \n Databricks is committed to fair and equitable compensation practices. The pay range(s) for this role is listed below and represents the expected salary range for non-commissionable roles or on-target earnings for commissionable roles.  Actual compensation packages are based on several factors that are unique to each candidate, including but not limited to job-related skills, depth of experience, relevant certifications and training, and specific work location. Based on the factors above, Databricks anticipates utilizing the full width of the range. The total compensation package for this position may also include eligibility for annual performance bonus, equity, and the benefits listed above. For more information regarding which range your location is in visit our page here . \n  \n Zone 1 Pay Range\n $190,000 — $261,250 USD \n About Databricks \n Databricks is the Data and AI company. More than 20,000 organizations worldwide — including adidas, AT\u0026T, Bayer, Block, Mastercard, Rivian, Unilever, and 70% of the Fortune 500 — rely on the Databricks Data + AI Platform to build and scale data and AI apps, analytics and agents. Headquartered in San Francisco with 30+ offices around the globe, Databricks offers a unified platform that includes Genie, Lakebase, Agent Bricks, Lakeflow, Lakehouse, and Unity Catalog. To learn more, follow Databricks on LinkedIn , X , YouTube , and Instagram . Benefits At Databricks, we strive to provide comprehensive benefits and perks that meet the needs of all of our employees. For specific details on the benefits offered in your region click here . \n Our Commitment to Diversity and Inclusion \n At Databricks, we are committed to fostering a diverse and inclusive culture where everyone can excel. We take great care to ensure that our hiring practices are inclusive and meet equal employment opportunity standards. Individuals looking for employment at Databric","salary_min":190000,"salary_max":261250,"location":"San Francisco, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["llm","distributed-systems","generative-ai","inference"],"apply_url":"https://databricks.com/company/careers/open-positions/job?gh_jid=8643979002","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-07-22T01:19:06Z","expires_at":"2026-09-29T13:32:29.735157Z","created_at":"2026-07-22T14:02:25.056967Z","updated_at":"2026-08-30T13:32:29.873458Z","company_name":"Databricks","company_slug":"databricks","company_logo_url":"https://www.google.com/s2/favicons?domain=databricks.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/4d8ed5b4-93f4-4766-aa31-2fc37f443b6c"},{"id":"3f34c43f-1307-4759-b68f-8073291a1c99","company_id":"332b7698-676b-4a3e-8b02-81b1195c5af6","title":"Staff Software Engineer, Foundation Model Inference ","slug":"staff-software-engineer-foundation-model-api-b7d81082","description":"P-1930 \n At Databricks, we are passionate about enabling data and AI teams to solve the world's toughest problems — from making the next mode of transportation a reality to accelerating the development of medical breakthroughs. We do this by building and running the world's best data and AI infrastructure platform so our customers can use deep data insights to improve their business. Founded by engineers — and customer-obsessed — we leap at every opportunity to solve technical challenges, from designing next-gen UI/UX for interfacing with data to scaling our services and infrastructure across millions of virtual machines. And we're only getting started.\n As part of the AI team, you'll build the platforms and products that power everything from data apps, AI agents, model training, model serving, and Vector Search. You'll be joining a high-agency, high-visibility team operating at the frontier of AI infrastructure — with deep ties to research, product, and real-world enterprise use cases. Databricks Mosaic AI is one of our fastest-growing businesses, helping thousands of our customers democratize AI within their organizations. We're building the products and infrastructure that power the next generation of AI.\n We're hiring across multiple teams in our AI Engineering org, including the FMAPI (Foundation Model APIs) team — the unified serving layer for large language models across real-time and batch inference, powering model inference at enterprise scale. We are looking to hire high-agency engineers who bridge the gap between technical execution and product strategy.\n The impact you will have: \n \n Build LLM infrastructure powering large-scale inference workloads for customers through partner models (OpenAI, Anthropic, Gemini) and self-hosted models (Qwen, GPT-OSS, Llama)\n Shape the direction of the FMAPI product — from roadmap to execution — by leveraging deep customer empathy and direct engagement with enterprise users and model providers\n Improve reliability, latency, and efficiency of distributed AI workloads\n Collaborate with platform, infra, and ML teams to deliver seamless end-to-end experiences\n Shape how developers and data scientists build and interact with AI on Databricks\n \n What we look for: \n \n 8+ years of experience in backend or infrastructure engineering\n Experience with distributed systems, scalable APIs, or cloud-native infrastructure\n Strong product and ownership mindset, with a focus on shipping user-facing value\n Experience with real-time serving, ML infrastructure, or GPU orchestration\n Familiarity with service-oriented architecture, deployment pipelines, and system observability\n Strong programming skills in Scala, Go, or Python\n \n Bonus points for: \n \n Exposure to platforms like SageMaker, Vertex AI, or Azure ML\n Built products that support AI workflows\n  \n Pay Range Transparency \n Databricks is committed to fair and equitable compensation practices. The pay range(s) for this role is listed below and represents the expected salary range for non-commissionable roles or on-target earnings for commissionable roles.  Actual compensation packages are based on several factors that are unique to each candidate, including but not limited to job-related skills, depth of experience, relevant certifications and training, and specific work location. Based on the factors above, Databricks anticipates utilizing the full width of the range. The total compensation package for this position may also include eligibility for annual performance bonus, equity, and the benefits listed above. For more information regarding which range your location is in visit our page here . \n  \n Local Pay Range\n $190,000 — $265,000 USD \n About Databricks \n Databricks is the Data and AI company. More than 20,000 organizations worldwide — including adidas, AT\u0026T, Bayer, Block, Mastercard, Rivian, Unilever, and 70% of the Fortune 500 — rely on the Databricks Data + AI Platform to build and scale data and AI apps, analytics and agents. Headquartered in San Francisco with 30+ offices around the globe, Databricks offers a unified platform that includes Genie, Lakebase, Agent Bricks, Lakeflow, Lakehouse, and Unity Catalog. To learn more, follow Databricks on LinkedIn , X , YouTube , and Instagram . Benefits At Databricks, we strive to provide comprehensive benefits and perks that meet the needs of all of our employees. For specific details on the benefits offered in your region click here . \n Our Commitment to Diversity and Inclusion \n At Databricks, we are committed to fostering a diverse and inclusive culture where everyone can excel. We take great care to ensure that our hiring practices are inclusive and meet equal employment opportunity standards. Individuals looking for employment at Databricks are considered without regard to age, color, disability, ethnicity, family or marital status, gender identity or expression, language, national origin, physical and mental ability, political affiliation, race, ","salary_min":190000,"salary_max":265000,"location":"San Francisco, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["generative-ai","cloud","mlops","agents","llm","distributed-systems","inference"],"apply_url":"https://databricks.com/company/careers/open-positions/job?gh_jid=8637143002","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-07-16T18:21:12Z","expires_at":"2026-09-29T13:32:36.095952Z","created_at":"2026-07-18T14:02:34.2202Z","updated_at":"2026-08-30T13:32:36.239982Z","company_name":"Databricks","company_slug":"databricks","company_logo_url":"https://www.google.com/s2/favicons?domain=databricks.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/3f34c43f-1307-4759-b68f-8073291a1c99"},{"id":"02fdc710-8e20-40fd-aedd-05f740fa50ac","company_id":"377b9ca2-ac79-48a5-8657-da630f9e447d","title":"Senior Staff / Principal Machine Learning Scientist, AI Inference \u0026 Optimization","slug":"senior-staff-principal-machine-learning-scientist-ai-inference-optimization-8c8ecaa7","description":"Join the Future of Security at Netskope\n Netskope (NASDAQ: NTSK) is a leader in modern security and networking for the cloud and AI era. We secure and accelerate cloud, data, and AI in real time, everywhere. Thousands of customers, including more than 30 of the Fortune 100, trust the Netskope One platform, its Zero Trust Engine, and the powerful NewEdge network to gain full visibility and control without performance trade-offs.\n At Netskope, our technology is driven by our greatest strength: our people. We believe that belonging powers innovation, and success is both personal and organizational. We embrace differences in gender, ethnicity, beliefs, ability, and identity, creating an environment where every voice is heard and respected. We empower our employees to bring their authentic selves to work, grow their careers through continuous education and mentorship, and lead with transparency and curiosity. Join a team where you belong, where you are encouraged to be an entrepreneur, and where together, we continue to redefine the landscape of security.\n Visit Careers at Netskope to learn more. Follow us on  LinkedIn and  Instagram .\n Positions are available at Senior Staff and above. Candidates are assessed individually and leveled according to their specific skills and background. \n About the role\n As a Senior Staff Machine Learning Scientist, you own the inference and optimization layer that makes AI in agentic workflows fast, efficient, and production-grade. You fine-tune and evaluate models, push latency and throughput on real hardware, and build the runtime that executes bounded AI tasks, validated against usage from Netskope’s large customer base so you optimize where the data points, not where you guess.\n What’s in it for you\n \n High-impact ownership. You own the model layer of a net-new product that changes the performance and economics of agentic AI.\n Cutting-edge, unusual stack. The hard, interesting inference problems live here: quantization, KV-cache and memory management, sparsity, fine-tuning, and hardware acceleration under real-world resource constraints.\n Real scale to build against. Netskope’s customer footprint gives you production signals most teams never see, so you deploy, validate, and iterate fast.\n \n What you will be doing\n \n Build and optimize the model inference path : quantization, KV-cache optimization, batching, and latency/memory/throughput tuning on constrained, commodity hardware.\n Fine-tune and evaluate models for bounded tasks; build eval harnesses that gate a capability to release on real accuracy, latency, and security relevance.\n Design and grow the task execution runtime (bounded sub-agents), pushing toward dynamic task generation and context compaction.\n Drive hardware acceleration / sparsity and support for larger models as the platform matures.\n Partner with the systems and backend engineers to ship capabilities end-to-end and iterate on real production signals.\n \n Required skills and experience\n \n 10+ years of overall industry experience , with 4+ years hands-on in ML/AI (model development, fine-tuning, and inference optimization).\n Hands-on with fine-tuning (e.g. LoRA/QLoRA), quantization (GGUF/AWQ/GPTQ), and inference runtimes (vLLM/SGLang, TensorRT-LLM, ONNX Runtime, llama.cpp, or MLX/CoreML). On-device or edge inference experience is a strong plus.\n Strong Python; comfort reaching into C++ for low-level interop is a plus.\n Solid grasp of transformer internals and the levers that move real inference performance and cost: KV cache, attention, batching, memory footprint.\n Fluency with agentic coding systems and genuine curiosity about agent harnesses like Claude Code, Pi, and Codex , so you should already be building with them, or itching to.\n Clear communication: able to distill a model or infra bottleneck into an actionable concept for cross-functional teammates.\n \n Education\n \n MS in Computer Science, Machine Learning, Electrical Engineering, or equivalent technical degree required, with a focus in AI/ML research; PhD in a related field strongly preferred.\n Compensation:  \n At Netskope, salary is one component of our competitive total rewards package. The salary range for this position is as listed below. This is a national range. For purposes of complying with applicable laws, the range applies to candidates in California, Colorado, Illinois, Maryland, New York, Washington, and other states. \n The successful candidate’s starting pay will also be determined based on job-related skills, experience, qualifications, location, and market conditions.  \n For all sales roles, the posted salary range is the On Target Earnings (OTE) range for the role, which is the sum of base salary and target commission amount at 100% goal achievement. \n In addition to salary, candidates may be eligible for other forms of compensation such as participation in a bonus plan (for non-sales roles) and a stock award program. Candidates may also be eligible for a comprehensive he","salary_min":124500,"salary_max":272000,"location":"San Jose, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"principal","tags":["llm","fine-tuning","cloud","agents","inference","machine-learning"],"apply_url":"https://www.netskope.com/company/careers/open-positions/?gh_jid=8063869","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-07-14T04:20:32Z","expires_at":"2026-09-29T13:39:57.880956Z","created_at":"2026-07-15T14:11:39.076302Z","updated_at":"2026-08-30T13:39:58.013353Z","company_name":"Netskope","company_slug":"netskope","company_logo_url":"https://www.google.com/s2/favicons?domain=netskope.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/02fdc710-8e20-40fd-aedd-05f740fa50ac"},{"id":"6e73bc75-a490-4b93-af2d-5d0040a7eb71","company_id":"6ea0f41a-b13e-481a-b410-5195f391f939","title":"Research Engineer, Post-Training Inference","slug":"research-engineer-post-training-inference-ff4ae18b","description":"About the role \n The Model Shaping team at Together AI works on products and research focused on tailoring open foundation models to downstream applications. We build services that enable machine learning developers to choose the best models for their tasks and further improve these models using domain-specific data. In addition, we develop new methods for more efficient model training and evaluation, drawing inspiration from a broad range of ideas across machine learning, natural language processing, and ML systems.\n As a Research Engineer within Model Shaping, you will develop a platform that enables users to customize open-source models with their own data. Working across the training and inference stacks, you will build and improve our Fine-Tuning, Reinforcement Learning, and Evaluation services – from ensuring a seamless path from post-training to production serving, to optimizing the inference engine for RL training workloads. You will collaborate closely with our product, research, and engineering teams to keep the API reliable, performant, and well integrated into the company's technical infrastructure. Above all, you will help build the foundational layer of the open-source AI ecosystem, enabling developers around the world to efficiently create high-quality models tailored to their specific applications.\n Responsibilities \n \n Design and build Together’s systems for customizing open-source models\n Build integrations between the Model Shaping and Inference platforms to ensure a seamless path from post-training to serving production workloads\n Add features to inference engines for large-scale post-training experiments, including optimizations for RL workloads\n Make sure the service is stable and robust, participating in an on-call rotation and ensuring 24/7 availability of our platform\n \n Requirements \n \n Have 2+ years of experience building and deploying machine learning-based services in a production environment\n Have hands-on experience with modern inference engines, such as SGLang, vLLM, and TensorRT-LLM\n Are familiar with the latest methods for fine-tuning LLMs and other AI models\n Have a strong software engineering background in Python or Go\n Stay up to date with the latest advances and trends in the machine learning community\n \n Experience in any of the following will make you stand out \n \n Serving low-precision (FP4/FP8) models, multiple LoRA adapters within one model instance (Multi-LoRA), or models distributed across several GPU nodes\n Optimizing the performance of RL training workloads\n Developing CUDA/Triton/CuTE DSL kernels for inference\n Developing large-scale and high-load production systems\n Maintaining or contributing to open-source ML projects\n Managing machine learning workloads on Kubernetes clusters\n \n About Together AI \n Together AI is a research-driven artificial intelligence company. We believe open and transparent AI systems will drive innovation and create the best outcomes for society, and together we are on a mission to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models. We have contributed to leading open-source research, models, and datasets to advance the frontier of AI, and our team has been behind technological advancement such as FlashAttention, ATLAS, RedPajama, and Mamba. We invite you to join a passionate group of researchers in our journey in building the next generation AI infrastructure.\n Compensation \n We offer competitive compensation, startup equity, health insurance, and other benefits. The US base salary range for this full-time position is $200,000 - $290,000. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge.\n Equal Opportunity \n Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.\n Please see our privacy policy at  https://www.together.ai/privacy","salary_min":200000,"salary_max":290000,"location":"San Francisco, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"junior","tags":["fine-tuning","search","llm","nlp","generative-ai","reinforcement-learning","gpu","inference"],"apply_url":"https://job-boards.greenhouse.io/togetherai/jobs/5179372007","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-07-06T18:21:40Z","expires_at":"2026-09-29T13:32:21.485882Z","created_at":"2026-07-09T14:02:08.323229Z","updated_at":"2026-08-30T13:32:21.629042Z","company_name":"Together AI","company_slug":"together-ai","company_logo_url":"https://www.google.com/s2/favicons?domain=together.ai\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/6e73bc75-a490-4b93-af2d-5d0040a7eb71"},{"id":"fa7a576b-4fca-402a-b5c9-b66e260cd3e9","company_id":"f5ee7284-a657-4da2-b351-cb806a3681cd","title":"Member of Technical Staff - RL Inference","slug":"member-of-technical-staff-rl-inference-cb7bd166","description":"SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge.  Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. \n ABOUT THE ROLE: \n The RL infrastructure team is looking for an engineer to help with low precision RL training and inference.\n RESPONSIBILITIES: \n \n Design and optimize our inference stack for all shapes of RL workloads at SpaceXAI, from small scale ablations to production training runs.\n Analyze, profile and address performance bottlenecks in large scale RL systems\n Work closely with the modelling team to efficiently implement novel RL techniques and algorithms\n \n BASIC QUALIFICATIONS: \n \n Experience in building, debugging, and optimizing efficiency of large-scale distributed systems\n Experience in LLM inference\n Proficiency in programming languages such as Python, C++ and/or Rust; frameworks such as PyTorch, Jax, CUDA\n Willingness to dive deep and solve hardcore problems at all levels of the stack\n \n PREFERRED SKILLS AND EXPERIENCE: \n \n Strong knowledge in quantization and numerics in LLM inference and training\n Experience in developing inference engines, e.g. SGLang, vLLM\n \n COMPENSATION AND BENEFITS: \n $180,000 - $440,000 USD\n Base salary is just one part of our total rewards package at SpaceXAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short \u0026 long-term disability insurance, life insurance, and various other discounts and perks.\n SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice .","salary_min":180000,"salary_max":440000,"location":"Palo Alto, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["gpu","llm","distributed-systems","pytorch","inference"],"apply_url":"https://job-boards.greenhouse.io/xai/jobs/5180223007","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-07-06T14:59:58Z","expires_at":"2026-09-29T13:33:49.351225Z","created_at":"2026-07-09T14:03:21.460874Z","updated_at":"2026-08-30T13:33:49.492616Z","company_name":"xAI","company_slug":"xai","company_logo_url":"https://www.google.com/s2/favicons?domain=x.ai\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/fa7a576b-4fca-402a-b5c9-b66e260cd3e9"},{"id":"8ff64df9-5c3c-4794-90fd-dc4cbcafc029","company_id":"a0000000-0000-0000-0000-000000000001","title":"Staff + Senior Software Engineer, Inference Deployment","slug":"staff-senior-software-engineer-inference-deployment-8d22b347","description":"About Anthropic \n Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.\n About the role\n Our Inference team is responsible for building and maintaining the critical systems that serve Claude to millions of users worldwide. We bring Claude to life by serving our models via the industry’s largest compute-agnostic inference deployments. We are responsible for the entire stack from intelligent request routing to fleet-wide orchestration across diverse AI accelerators.\n The team has a dual mandate: maximizing compute efficiency to reliably serve our explosive customer growth, while enabling breakthrough research by giving our scientists the high-performance inference infrastructure they need to develop next-generation models. We tackle complex, distributed systems challenges across multiple accelerator families and emerging AI hardware running in multiple cloud platforms.\n Inference systems are highly performance sensitive distributed systems. Inference serves hundreds of thousands of customers every day, and the size \u0026 span of the inference fleet requires sophisticated routing, scaling, and networking systems.\n Key responsibilities\n \n Design, build, and maintain the distributed systems that serve Claude to millions of users worldwide\n Develop resilient, flexible systems that adapt in real time to real world events\n Develop intelligent request routing, load balancing, and traffic management systems across thousands of accelerators\n Maximize compute efficiency across the fleet by autoscaling and orchestrating production, research, and experimental workloads\n Build and operate production-grade deployment pipelines for releasing new models to users\n Provide high-performance inference infrastructure that enables researchers to develop next-generation models\n Integrate new AI accelerator platforms and support inference for new model architectures\n \n Minimum qualifications\n \n Significant software engineering experience, particularly with distributed systems\n Results-oriented, with a bias towards flexibility and impact\n Willingness to pick up slack, even if it goes outside your job description\n Desire to learn more about machine learning systems and infrastructure\n Thrive in environments where technical excellence directly drives both business results and research breakthroughs\n Care about the societal impacts of your work\n \n Preferred qualifications\n \n Experience with high-performance, large-scale distributed systems\n Experience implementing and deploying machine learning systems at scale\n Experience with load balancing, request routing, or traffic management systems\n Familiarity with LLM inference optimization, batching, and caching strategies\n Experience with Kubernetes and cloud infrastructure (AWS, GCP, Azure)\n Proficiency in Python or Rust\n \n Representative projects\n \n Designing intelligent routing algorithms that optimize request distribution across many accelerators in different environments\n Autoscaling our compute fleet to dynamically match supply with demand across production, research, and experimental workloads\n Building production-grade deployment pipelines for releasing new models to millions of users reliably\n Contributing to new inference features\n Supporting inference for new model architectures\n Analyzing observability data to tune performance based on real-world production workloads\n Managing multi-region deployments and geographic routing for global customers\n \n Deadline to apply:  None. Applications will be reviewed on a rolling basis. \n The annual compensation range for this role is listed below. \n For sales roles, the range provided is the role’s On Target Earnings (\"OTE\") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role.\n Annual Salary:\n $320,000 — $485,000 USD \n Logistics \n Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience\n Required field of study:  A field relevant to the role as demonstrated through coursework, training, or professional experience\n Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position\n Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices.\n Visa sponsorship:  We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this.\n We encourage you to apply even if you do not believe you meet every sin","salary_min":320000,"salary_max":485000,"location":"San Francisco, CA","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["distributed-systems","alignment","cloud","llm","inference","infrastructure"],"apply_url":"https://job-boards.greenhouse.io/anthropic/jobs/5285557008","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-06-29T14:33:21Z","expires_at":"2026-09-29T13:30:36.83131Z","created_at":"2026-06-30T14:00:33.50732Z","updated_at":"2026-08-30T13:30:36.97278Z","company_name":"Anthropic","company_slug":"anthropic","company_logo_url":"https://www.google.com/s2/favicons?domain=anthropic.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/8ff64df9-5c3c-4794-90fd-dc4cbcafc029"},{"id":"a8d08f57-0c6f-40d4-b42c-4fa9061feaed","company_id":"2114efab-ea67-411b-bfb8-7899153105f3","title":"Member of Technical Staff, Inference ","slug":"member-of-technical-staff-inference-86b5008f","description":"Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.\n\n\n\n\nABOUT THE ROLE\n\nWe're looking for an inference runtime engineer to push the boundaries of what's possible in LLM and diffusion model serving. Models grow larger. Architectures shift: mixture-of-experts, multimodal, agentic. Every breakthrough demands innovations on the inference engine itself. You'll work at the core of vLLM, optimizing how models execute across diverse hardware and architectures. Your work will directly impact how the world runs AI inference. \n\n\n\n\nSKILLS AND QUALIFICATIONS\n\nMinimum qualifications:\n\n - Bachelor's degree or equivalent experience in computer science, engineering, or similar.\n\n - Deep understanding of transformer architectures and their variants.\n\n - Strong programming skills in Python with experience in PyTorch internals.\n\n - Experience with LLM inference systems (vLLM, TensorRT-LLM, SGLang, TGI).\n\n - Ability to read and implement model architectures and inference techniques from research papers.\n\n - Demonstrate the ability to contribute performant and maintainable code and debug in complex ML codebases.\n\nPreferred qualifications:\n\n - Deep understanding of KV-cache memory management, prefix caching, and hybrid model serving.\n\n - Familiarity with RL frameworks and algorithms for LLMs.\n\n - Experience with multimodal inference (audio/image/video/text).\n\n - Contributions to open-source ML or system infrastructure projects.\n\nBonus points if you have:\n\n - Implemented core features in vLLM or other inference engine projects.\n\n - Contributed to vLLM integrations (verl, OpenRLHF, Unsloth, LlamaFactory, etc).\n\n - Written widely-shared technical blogs or side projects on vLLM or LLM inference.\n   \n   \n\n\nLOGISTICS\n\n - Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.\n\n - Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.\n\n - Visa sponsorship: We sponsor visas on a case-by-case basis.\n\n - Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.","salary_min":200000,"salary_max":400000,"location":"San Francisco, CA","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["llm","reinforcement-learning","diffusion-models","agents","pytorch","mlops","research","inference"],"apply_url":"https://jobs.ashbyhq.com/inferact/43c0ca54-fcf5-41fa-83a1-38800c75ccc0/application","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-06-18T18:55:18.703Z","expires_at":"2026-09-29T13:41:27.162539Z","created_at":"2026-06-28T14:10:47.134304Z","updated_at":"2026-08-30T13:41:27.292841Z","company_name":"Inferact","company_slug":"inferact","company_logo_url":"https://www.google.com/s2/favicons?domain=inferact.ai\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/a8d08f57-0c6f-40d4-b42c-4fa9061feaed"},{"id":"da71ff33-42e4-48bf-87eb-04189fa35d9d","company_id":"a0000000-0000-0000-0000-000000000001","title":"Staff+ Software Engineer, Inference Velocity","slug":"staff-software-engineer-inference-runtime-0cb9cf6b","description":"About Anthropic \n Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.\n About the role\n Anthropic's Inference organization serves Claude to millions of users and enterprise customers with the speed, reliability, and efficiency that frontier AI demands. We build across GPUs, TPUs, and Trainium, and the complexity of our development environment grows with every platform we add. We're looking for a Staff engineer to be the technical lead for Inference Developer Productivity: the team that makes every engineer in the org dramatically more effective at building, testing, and shipping inference software.\n This is a senior IC role with broad technical ownership. You'll set technical direction for the team's toolchains, workflows, and feedback loops, and you'll be the one making the hard calls on architecture, prioritization, and tradeoffs across heterogeneous accelerator platforms. You'll pair with the team's Engineering Manager, who owns hiring and people development, while you own the technical roadmap and drive the work. You'll also partner closely with Anthropic's central Infrastructure org, where company-wide developer productivity lives, to make sure Inference's multi-accelerator reality is well served without duplicating effort.\n This role is for someone who has been the technical anchor on a platform or infrastructure team before, who thinks in systems and feedback loops, and who gets real satisfaction from the moment another engineer stops fighting their environment and starts shipping.\n Key responsibilities\n \n Set technical direction for Inference Developer Productivity, owning the architecture and roadmap for toolchains, dev environments, and CI/CD across GPU (CUDA), TPU, and Trainium platforms \n Be the technical owner of accelerator toolchain management: compilers, drivers, libraries, frameworks, kept current, compatible, and well-tested so Inference engineers focus on model serving instead of environment archaeology \n Design and build infrastructure for efficient accelerator usage during development, including devbox environments, pre- and post-land validation automation, and shared tooling that reduces the cost of working across heterogeneous hardware \n Define and instrument productivity metrics for the Inference org, building the dashboards and alerting that surface regressions early (smoke tests red for extended periods, build times creeping up, toolchain breakages) and drive them to resolution \n Proactively hunt down bottlenecks, toil, and friction across Inference engineering workflows, then design and build the systems that eliminate them \n Act as the technical counterpart to Anthropic's central Infrastructure org, aligning on shared developer productivity initiatives, contributing Inference-specific requirements, and making the call on build vs. adopt \n Mentor engineers on the team through design review, code review, and direct collaboration, raising the technical bar without owning headcount\n \n Minimum qualifications\n \n 8+ years of software engineering experience, with significant time as the technical lead or anchor on an infrastructure, platform, or developer productivity team \n Deep background in systems engineering, build/test infrastructure, or ML infrastructure, with the ability to go hands-on with toolchain issues, CI/CD pipelines, and developer workflow optimization \n Experience owning toolchains or development environments for compute-intensive workloads (ML training or inference, HPC, large-scale distributed systems) \n Real depth in at least one accelerator ecosystem (CUDA/GPU, TPU, or Trainium/AWS Neuron) and genuine appetite to learn the others \n A track record of defining and using engineering metrics to drive improvement: you've built dashboards, set SLOs on developer workflows, or led initiatives that measurably improved engineering velocity \n Experience driving technical alignment across organizational boundaries, advocating for your team's needs while contributing to shared infrastructure \n Strong written and verbal communication, and the ability to influence technical direction without formal authority\n \n Preferred qualifications\n \n Experience with ML compiler toolchains (XLA, Triton, NeuronX) or accelerator driver/firmware management at scale \n Background building or running shared development environments (devboxes, remote development, ephemeral environments) for hardware-dependent workflows \n Experience with CI/CD systems at scale, particularly for workloads involving accelerator hardware \n Familiarity with Kubernetes-based development and job scheduling environments \n Prior tech lead experience on a developer productivity or platform engineering team at a fast-growing AI/ML company\n The annual compensation","salary_min":405000,"salary_max":485000,"location":"San Francisco, CA","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["distributed-systems","cloud","gpu","alignment","mlops","inference","infrastructure"],"apply_url":"https://job-boards.greenhouse.io/anthropic/jobs/5257650008","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-06-12T20:17:38Z","expires_at":"2026-09-29T13:30:40.414519Z","created_at":"2026-06-28T14:00:35.550199Z","updated_at":"2026-08-30T13:30:40.558553Z","company_name":"Anthropic","company_slug":"anthropic","company_logo_url":"https://www.google.com/s2/favicons?domain=anthropic.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/da71ff33-42e4-48bf-87eb-04189fa35d9d"},{"id":"8dd71050-3131-45bd-a196-a71038f3f814","company_id":"b467c425-56b3-40ce-826a-e603e82a08bd","title":"Senior Product Manager, Inference Platform ","slug":"senior-product-manager-inference-platform-806705aa","description":"Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators.  \n At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there.  \n A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. \n About the Role \n Inference Platform is Roblox's multi-tier orchestration system for microservices, powering AI Inference across Roblox. Through a simple developer API, we deliver a \"deploy and forget\" runtime, hiding the complexity of scheduling, scaling, and reliably running services across our on-prem and multi-cloud footprint at global scale. As Senior Product Manager, Inference Platform, you'll take the helm at a defining moment - leading the charge as we scale the platform to become the default runtime for critical Roblox services worldwide.\n You Will \n \n Own Inference Platform end-to-end - set the multi-year vision for how Roblox engineers deploy, run, and scale AI and other services across our Core and Edge Datacenters, and cloud.          \n Power Roblox's AI future - build the platform that brings frontier models and next-gen AI workloads to life, with the primitives, scheduling guarantees, and resource classes AI teams need to move fast.                              \n Evolve the platform's technical core - define how Roblox prioritizes, isolates, and preempts workloads across shared fleet, and shape the entire system across global control plane, multi-region orchestration, distributed systems primitives, capacity/state management, DR/failover, and operator-based Kubernetes extensions.                                               \n Make developer experience the reason teams choose Jobs Platform - by delivering integrated tooling (CI/CD, telemetry, profiling, tracing) and managing robust SLOs via real-time fleet signals (health, queue depth, scheduling efficiency)—all while engineering a frictionless, effortless onboarding experience for teams.                                                        \n Architect intelligent, predictive scaling - advance Multi-Cloud Bursting from reactive to proactive with AI-powered prediction, and partner with Data Science and Finance on demand forecasting, cost-to-serve analytics, and AI-driven resource optimization.     \n Own the trade-off calls across efficiency, latency, cost-to-serve and time-to-ship, while maintaining reliability.\n Lead cross-functionally, acting as the connective tissue between Jobs Platform Users, and Jobs Platform, AI, and infrastructure engineers. \n \n You Have \n \n 7+ years of product management experience building resource orchestration, workload management platforms or large-scale distributed systems \n Track record of building developer-facing platforms that engineers genuinely love to use. You have a deep understanding of modern software development lifecycle and developer tooling with a history of abstracting complex infrastructure into clean, maintainable platform schemas.\n Built and operated Kubernetes and enterprise-grade service mesh at scale, and you understand the developer pain points around control plane mechanics like etcd scaling and API server bottlenecks. You've shipped custom Operators, Controllers, and CRDs, going beyond vanilla Kubernetes. \n Familiarity with GPU/accelerator architecture and the complex scheduling challenges that come with it.\n Background in AI model development, training, inference \n A builder mindset - you are passionate about prototyping, evolving products through rapid iteration, and leveraging AI for ideation and unlocking value for users.\n Experience building workload management or resource orchestration platforms in AWS, Google Cloud Platform (GCP), Azure or other cloud providers. (preferred)\n Kernel-level experience or familiarity with custom kernel drivers. (preferred)\n Experience building agentic systems for workload management or infrastructure. (preferred)\n For roles that are based at our headquarters in San Mateo, CA: The starting base pay for this position is as shown below. The actual base pay is dependent upon a variety of job-related factors such as professional background, training, work experience, location, business needs and market demand. Therefore, in some circumstances, the actual salary could fall outside of this expected range. This pay range is sub","salary_min":280540,"salary_max":330950,"location":"San Mateo, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"senior","tags":["agents","microservices","distributed-systems","generative-ai","cloud","inference"],"apply_url":"https://careers.roblox.com/jobs/7997195?gh_jid=7997195","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-06-10T18:48:38Z","expires_at":"2026-09-29T13:47:50.726155Z","created_at":"2026-08-25T18:33:06.215553Z","updated_at":"2026-08-30T13:47:50.853379Z","company_name":"Roblox","company_slug":"roblox","company_logo_url":"https://www.google.com/s2/favicons?domain=roblox.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/8dd71050-3131-45bd-a196-a71038f3f814"},{"id":"ddcf527b-a153-46b0-a4f1-94832080a914","company_id":"a0000000-0000-0000-0000-000000000001","title":"Staff + Senior Software Engineer, Inference","slug":"staff-senior-software-engineer-inference-bb6014f6","description":"About Anthropic \n Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.\n About the role\n Our Inference team is responsible for building and maintaining the critical systems that serve Claude to millions of users worldwide. We bring Claude to life by serving our models via the industry’s largest compute-agnostic inference deployments. We are responsible for the entire stack from intelligent request routing to fleet-wide orchestration across diverse AI accelerators.\n The team has a dual mandate: maximizing compute efficiency to reliably serve our explosive customer growth, while enabling breakthrough research by giving our scientists the high-performance inference infrastructure they need to develop next-generation models. We tackle complex, distributed systems challenges across multiple accelerator families and emerging AI hardware running in multiple cloud platforms.\n Inference systems are highly performance sensitive distributed systems. Inference serves hundreds of thousands of customers every day, and the size \u0026 span of the inference fleet requires sophisticated routing, scaling, and networking systems.\n Key responsibilities\n \n Design, build, and maintain the distributed systems that serve Claude to millions of users worldwide\n Develop resilient, flexible systems that adapt in real time to real world events\n Develop intelligent request routing, load balancing, and traffic management systems across thousands of accelerators\n Maximize compute efficiency across the fleet by autoscaling and orchestrating production, research, and experimental workloads\n Build and operate production-grade deployment pipelines for releasing new models to users\n Provide high-performance inference infrastructure that enables researchers to develop next-generation models\n Integrate new AI accelerator platforms and support inference for new model architectures\n \n Minimum qualifications\n \n Significant software engineering experience, particularly with distributed systems\n Results-oriented, with a bias towards flexibility and impact\n Willingness to pick up slack, even if it goes outside your job description\n Desire to learn more about machine learning systems and infrastructure\n Thrive in environments where technical excellence directly drives both business results and research breakthroughs\n Care about the societal impacts of your work\n \n Preferred qualifications\n \n Experience with high-performance, large-scale distributed systems\n Experience implementing and deploying machine learning systems at scale\n Experience with load balancing, request routing, or traffic management systems\n Familiarity with LLM inference optimization, batching, and caching strategies\n Experience with Kubernetes and cloud infrastructure (AWS, GCP, Azure)\n Proficiency in Python or Rust\n \n Representative projects\n \n Designing intelligent routing algorithms that optimize request distribution across many accelerators in different environments\n Autoscaling our compute fleet to dynamically match supply with demand across production, research, and experimental workloads\n Building production-grade deployment pipelines for releasing new models to millions of users reliably\n Contributing to new inference features\n Supporting inference for new model architectures\n Analyzing observability data to tune performance based on real-world production workloads\n Managing multi-region deployments and geographic routing for global customers\n \n Deadline to apply:  None. Applications will be reviewed on a rolling basis. \n The annual compensation range for this role is listed below. \n For sales roles, the range provided is the role’s On Target Earnings (\"OTE\") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role.\n Annual Salary:\n $320,000 — $485,000 USD \n Logistics \n Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience\n Required field of study:  A field relevant to the role as demonstrated through coursework, training, or professional experience\n Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position\n Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices.\n Visa sponsorship:  We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this.\n We encourage you to apply even if you do not believe you meet every sin","salary_min":320000,"salary_max":485000,"location":"San Francisco, CA","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["llm","alignment","cloud","distributed-systems","inference","infrastructure"],"apply_url":"https://job-boards.greenhouse.io/anthropic/jobs/5245851008","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-06-08T20:10:07Z","expires_at":"2026-09-29T13:30:36.630333Z","created_at":"2026-06-28T14:00:33.577053Z","updated_at":"2026-08-30T13:30:36.782854Z","company_name":"Anthropic","company_slug":"anthropic","company_logo_url":"https://www.google.com/s2/favicons?domain=anthropic.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/ddcf527b-a153-46b0-a4f1-94832080a914"},{"id":"fb7bb920-e27c-4c81-b15f-63d2c9803c44","company_id":"a0000000-0000-0000-0000-000000000001","title":"Staff + Sr. Software Engineer, Cloud Inference Launch Engineering","slug":"staff-sr-software-engineer-cloud-inference-launch-engineering-b39407f2","description":"About Anthropic \n Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.\n About the role\n The Cloud Inference team scales and optimizes Claude to serve the massive audiences of developers and enterprise companies across AWS, GCP, Azure, and future cloud service providers (CSPs). We own the end-to-end product of Claude on each cloud platform, from API integration and intelligent request routing to inference execution, capacity management, and day-to-day operations.\n Within Cloud Inference, the model \u0026 inference launch team owns the validation pipeline for our inference server and load balancer on these platforms. We're responsible for every inference change — model launches, performance improvements, safeguard integrations — landing on cloud platforms with correctness, performance, and reliability intact.\n This is high-leverage infrastructure work: validation has to be fast and cheap enough to run on the same accelerators that serve customers, trustworthy enough to replace manual checks, and consistent enough that a change working on Anthropic first-party means it works everywhere. This directly determines how fast frontier models and features ship to every cloud platform, and how quickly performance wins reach production — reclaiming capacity at a time when compute is our scarcest resource.\n Key responsibilities\n \n Be on the critical path for frontier model launches, bringing up inference for new model architectures and shipping them to cloud platforms in lockstep with our first-party platform\n Work with the core inference team to bring new inference features (e.g. structured sampling, prompt caching, and more) to cloud platforms, owning the platform-specific integration that gets them to production\n Identify and dive deep on the gaps that make inference behave differently across first-party and CSPs — config drift, observability, deployment patterns, hard cross-platform bugs — and fix them at the source rather than building platform-specific workarounds\n Design, build, and own the CI/CD infrastructure for the inference server and load balancer across cloud platforms, with shadow traffic, performance baselines (throughput and latency), and correctness checks that catch regressions before production\n Drive down merge-to-production cycle time by making validation faster, more parallel, and cost-effective enough to run on the same constrained accelerator pool that serves customers, without trading away reliability \n Analyze observability data across providers to identify performance bottlenecks, cost anomalies, and regressions, and drive remediation based on real-world production workloads\n \n Minimum qualifications\n \n Have a strong interest in LLM serving; prior inference or ML experience is not required \n Have significant software engineering experience, with a strong background in high-performance, large-scale distributed systems serving millions of users\n Have a track record of building automation or test infrastructure that measurably improved release velocity or reliability\n Have experience building or operating services on at least one major cloud platform (AWS, GCP, or Azure), with exposure to Kubernetes, Infrastructure as Code, or container orchestration\n Thrive in cross-functional collaboration with both internal teams and external partners\n Are a fast learner who can quickly ramp up on new technologies, hardware platforms, and provider ecosystems\n Are highly autonomous and take ownership of problems end-to-end, including work that falls outside your job description\n \n Preferred qualifications\n \n LLM inference optimization, batching, and caching strategies\n Capacity-constrained scheduling or shared-resource test infrastructure\n Solid understanding of multi-region deployments, request routing, load balancing, global traffic management\n Working with CSP partner teams to scale infrastructure across multiple platforms, navigating differences in networking, security, privacy, and managed service\n Proficiency in Python or Rust\n The annual compensation range for this role is listed below. \n For sales roles, the range provided is the role’s On Target Earnings (\"OTE\") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role.\n Annual Salary:\n $320,000 — $485,000 USD \n Logistics \n Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience\n Required field of study:  A field relevant to the role as demonstrated through coursework, training, or professional experience\n Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position\n Location-based hybri","salary_min":320000,"salary_max":485000,"location":"San Francisco, CA","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["alignment","llm","distributed-systems","infrastructure","inference"],"apply_url":"https://job-boards.greenhouse.io/anthropic/jobs/5238296008","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-06-03T15:30:49Z","expires_at":"2026-09-29T13:30:43.145116Z","created_at":"2026-07-15T14:00:39.741964Z","updated_at":"2026-08-30T13:30:43.287243Z","company_name":"Anthropic","company_slug":"anthropic","company_logo_url":"https://www.google.com/s2/favicons?domain=anthropic.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/fb7bb920-e27c-4c81-b15f-63d2c9803c44"},{"id":"30c4c03e-a463-4f2f-ae06-6d06ab1405ba","company_id":"a0000000-0000-0000-0000-000000000001","title":"Staff + Sr. Software Engineer, Cloud Inference","slug":"staff-sr-software-engineer-cloud-inference-02c87a14","description":"About Anthropic \n Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.\n About the role \n The Cloud Inference team scales and optimizes Claude to serve the massive audiences of developers and enterprise companies across AWS, GCP, Azure, and future cloud service providers (CSPs). We own the end-to-end product of Claude on each cloud platform, from API integration and intelligent request routing to inference execution, capacity management, and day-to-day operations.\n Our engineers are extremely high leverage: we simultaneously drive multiple major revenue streams while optimizing one of Anthropic's most precious resources: compute. As we expand to more cloud platforms, the complexity of managing inference efficiently across providers with different hardware, networking stacks, and operational models grows significantly. We need product-minded backend engineers who can navigate these platform differences, design the services and abstractions that work across providers, and make architectural decisions that keep us reliable and cost-effective at massive scale.\n Your work will increase the scale at which our services operate, accelerate our ability to reliably launch new frontier models and innovative features to customers across all platforms, and ensure our LLMs meet rigorous safety, performance, and security standards.\n Key responsibilities \n \n Design, build, and own backend services and infrastructure that serve Claude across multiple CSPs, accounting for differences in compute hardware, networking, APIs, and operational models\n Work cross-functionally with internal inference, product API, systems, and security teams, among others, and with CSP partners to stand up the full serving stack on new cloud platforms, resolve operational issues, and influence provider roadmaps\n Build and evolve CI/CD automation systems, including validation and deployment pipelines, that reliably ship new model versions to millions of users across cloud platforms without regressions\n Design interfaces and tooling abstractions across CSPs that enable cost-effective inference management, scale across providers, and reduce per-platform complexity\n Contribute to capacity planning, autoscaling, and workload routing strategies that match supply with demand and direct requests to the most cost-effective accelerator and region\n Analyze observability data across providers to identify performance bottlenecks, cost anomalies, and regressions, and drive remediation based on real-world production workloads\n \n Minimum qualifications \n \n Have significant software engineering experience, with a strong background in high-performance, large-scale distributed systems serving millions of users\n Have experience building or operating services on at least one major cloud platform (AWS, GCP, or Azure), with exposure to Kubernetes, Infrastructure as Code, or container orchestration\n Are curious about LLM serving; prior inference or ML experience is not required\n Thrive in cross-functional collaboration with both internal teams and external partners\n Have experience working with external partners to align goals and deliver impact\n Are a fast learner who can quickly ramp up on new technologies, hardware platforms, and provider ecosystems\n Are highly autonomous and take ownership of problems end-to-end, including work that falls outside your job description\n \n Preferred qualifications \n \n Direct experience working with CSPs to scale infrastructure or products across multiple platforms, navigating differences in networking, security, privacy, billing, and managed service offerings\n Hands-on experience with capacity management, cost optimization, or resource planning at scale across heterogeneous environments\n Solid understanding of multi-region deployments, geographic routing, and global traffic management\n Proficiency in Python or Rust\n The annual compensation range for this role is listed below. \n For sales roles, the range provided is the role’s On Target Earnings (\"OTE\") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role.\n Annual Salary:\n $320,000 — $485,000 USD \n Logistics \n Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience\n Required field of study:  A field relevant to the role as demonstrated through coursework, training, or professional experience\n Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position\n Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices.\n Vis","salary_min":320000,"salary_max":485000,"location":"San Francisco, CA","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["distributed-systems","alignment","payments","llm","inference","infrastructure"],"apply_url":"https://job-boards.greenhouse.io/anthropic/jobs/5231496008","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-06-03T15:30:48Z","expires_at":"2026-09-29T13:30:43.054059Z","created_at":"2026-06-28T14:00:36.239717Z","updated_at":"2026-08-30T13:30:43.194546Z","company_name":"Anthropic","company_slug":"anthropic","company_logo_url":"https://www.google.com/s2/favicons?domain=anthropic.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/30c4c03e-a463-4f2f-ae06-6d06ab1405ba"}],"page":1,"per_page":20,"total":104,"total_is_exact":true,"total_pages":6}
