{"access":{"catalog_url":"https://aidevboard.com/api/v1/catalog","description":"Public read endpoints are open and free. API keys are optional for stable agent identity and keyed hourly throttling.","docs_url":"https://aidevboard.com/docs","employer_pilot_url":"https://aidevboard.com/verified-interview-pilot","mode":"open","register_url":"https://aidevboard.com/api/v1/register"},"candidate_resume_action":{"application_authorized":false,"candidate_charge":0,"endpoint":"https://aidevboard.com/api/v1/candidate/resume-preview","job_id_json_path":"jobs[].id","method":"POST","preview_requires_identity":false,"required_body_fields":["job_id","evidence_bullets"],"requires_explicit_human_review":true,"saved_artifact_protocol":"mcp","saved_artifact_requires_verified_human":true,"saved_artifact_tool":"compile_job_specific_resume","search_requires_identity":false,"status":"available_after_candidate_selects_job","submission_performed":false,"uses_candidate_verified_evidence":true},"degraded":false,"estimated":false,"has_next":false,"jobs":[{"id":"fad2f7a2-bd11-4215-b4fb-85ede2f803fe","company_id":"a0000000-0000-0000-0000-000000000001","title":"Staff+ Software Engineer, Kubernetes Platform","slug":"staff-software-engineer-kubernetes-platform-34576b6d","description":"About Anthropic \n Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.\n About the role \n We run one of the industry's largest AI compute fleets, spanning multiple cloud providers and datacenters, to train, research, and serve frontier AI models. Those fleets run on Kubernetes, and the Kubernetes Platform team owns the control plane that makes them work.\n We are operating at a scale where the defaults stop working. We own the scheduler and extend it to place topology-sensitive ML workloads across thousands of accelerators at once. We scale the control plane itself — apiserver, etcd, controllers — so it stays responsive as object counts and node counts grow by orders of magnitude. And we build the core cluster services every workload depends on, like service discovery, so they hold up under the same pressure.\n We make sure the control plane is fast, correct, and always available. Your work will directly determine whether Anthropic can keep reliably and safely training frontier models as our compute footprint continues to grow.\n Key responsibilities \n \n Own, operate, and extend the Kubernetes scheduler for Anthropic's accelerator fleets, including custom scheduling plugins and policies for gang scheduling, topology awareness, and preemption\n Scale the Kubernetes control plane (apiserver, etcd, controller-manager) to support clusters far beyond typical limits, and find the next bottleneck before it finds us\n Design, build, and operate core cluster services such as service discovery that every workload in the fleet depends on\n Build and maintain custom controllers, operators, and CRDs\n Partner with research, training, and inference to understand workload shapes and turn their requirements into platform capabilities\n Collaborate with cloud providers on required features and escalations\n Participate in on-call, lead incident response, and design processes (postmortems, runbooks, SLOs) that help the team avoid repeating failures\n \n Minimum qualifications \n \n Significant software engineering experience building and operating production distributed systems\n Proficiency in at least one systems-appropriate language (e.g., Go, Python, Rust, or C++)\n Deep, hands-on Kubernetes experience (well beyond \"user of”) into scheduler, controllers, apiserver, or operating large multi-tenant clusters\n Demonstrated ability to debug complex issues across the stack, from API behavior down to node and network-level root causes\n A track record of designing for reliability, correctness, and clear failure semantics in systems other engineers depend on\n Strong written and verbal communication; comfort building consensus with internal stakeholders\n \n Preferred qualifications \n \n Experience with Kubernetes internals or contributions: kube-scheduler / scheduling framework, apiserver, etcd, client-go, controller-runtime, or similar\n Experience building or operating cluster schedulers or batch systems (e.g., Kueue, Volcano, Slurm, or in-house equivalents)\n Background scaling control planes or coordination systems (etcd, ZooKeeper, Consul, or large DNS/service-mesh deployments)\n Familiarity with ML infrastructure: GPUs, TPUs, or Trainium; gang scheduling; topology-aware placement; collective networking such as NCCL\n Experience with GCP and/or AWS, including GKE/EKS internals and Infrastructure as Code\n Low-level systems experience such as Linux kernel tuning, cgroups, or eBPF\n 10+ years of relevant industry experience, including time leading large, ambiguous infrastructure projects\n The annual compensation range for this role is listed below. \n For sales roles, the range provided is the role’s On Target Earnings (\"OTE\") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role.\n Annual Salary:\n $320,000 — $485,000 USD \n Logistics \n Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience\n Required field of study:  A field relevant to the role as demonstrated through coursework, training, or professional experience\n Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position\n Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices.\n Visa sponsorship:  We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this.\n We encourage you to apply even if you do not believe you meet every sin","salary_min":320000,"salary_max":485000,"location":"San Francisco, CA","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["distributed-systems","alignment","cloud","infrastructure","kubernetes"],"apply_url":"https://job-boards.greenhouse.io/anthropic/jobs/5211241008","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-05-06T07:14:11Z","expires_at":"2026-09-29T13:30:40.877316Z","created_at":"2026-05-06T14:00:40.021287Z","updated_at":"2026-08-30T13:30:41.018521Z","company_name":"Anthropic","company_slug":"anthropic","company_logo_url":"https://www.google.com/s2/favicons?domain=anthropic.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/fad2f7a2-bd11-4215-b4fb-85ede2f803fe"},{"id":"134572d2-b112-45db-8b0e-fe6ebb53bbcc","company_id":"a0000000-0000-0000-0000-000000000013","title":"Platform Engineer - AI/ML Infrastructure (Kubernetes \u0026 Terraform)","slug":"site-reliability-engineer-ai-ml-infrastructure-kubernetes-aws-terraform-1154b02a","description":"COMPANY OVERVIEW\n\nDeepgram is the leading platform underpinning the emerging trillion-dollar Voice AI economy, providing real-time APIs for speech-to-text (STT), text-to-speech (TTS), and building production-grade voice agents at scale. More than 200,000 developers and 1,300+ organizations build voice offerings that are ‘Powered by Deepgram’, including Twilio, Cloudflare, Sierra, Decagon, Vapi, Daily, Cresta, Granola, and Jack in the Box. Deepgram’s voice-native foundation models are accessed through cloud APIs or as self-hosted and on-premises software, with unmatched accuracy, low latency, and cost efficiency. Backed by a recent Series C led by leading global investors and strategic partners, Deepgram has processed over 50,000 years of audio and transcribed more than 1 trillion words. There is no organization in the world that understands voice better than Deepgram.\n\n\n\n\nCOMPANY OPERATING RHYTHM\n\nAt Deepgram, we expect an AI-first mindset—AI use and comfort aren’t optional, they’re core to how we operate, innovate, and measure performance.\n\nEvery team member who works at Deepgram is expected to actively use and experiment with advanced AI tools, and even build your own into your everyday work. We measure how effectively AI is applied to deliver results, and consistent, creative use of the latest AI capabilities is key to success here. Candidates should be comfortable adopting new models and modes quickly, integrating AI into their workflows, and continuously pushing the boundaries of what these technologies can do.\n\nAdditionally, we move at the pace of AI. Change is rapid, and you can expect your day-to-day work to evolve just as quickly. This may not be the right role if you’re not excited to experiment, adapt, think on your feet, and learn constantly, or if you’re seeking something highly prescriptive with a traditional 9-to-5.\n\n\n\nOpportunity:\n\nWe're looking for an experienced Platform Engineer to build and operate the hybrid infrastructure foundation for our advanced AI/ML research and product development. You'll architect, build, and run the platform spanning AWS and our bare metal data centers, empowering our teams to train and deploy complex models at scale. This role is focused on creating a robust, self-service environment using Kubernetes, AWS, and Infrastructure-as-Code (Terraform), and orchestrating high-demand GPU workloads using schedulers like Slurm.\n\n\n\nWhat You’ll Do\n\n - Architect and maintain our core computing platform using Kubernetes on AWS and on-premise, providing a stable, scalable environment for all applications and services.\n\n - Develop and manage our entire infrastructure using Infrastructure-as-Code (IaC) principles with Terraform, ensuring our environments are reproducible, versioned, and automated.\n\n - Design, build, and optimize our AI/ML job scheduling and orchestration systems, integrating Slurm with our Kubernetes clusters to efficiently manage GPU resources.\n\n - Provision, manage, and maintain our on-premise bare metal server infrastructure for high-performance GPU computing.\n\n - Implement and manage the platform's networking (CNI, service mesh) and storage (CSI, S3) solutions to support high-throughput, low-latency workloads across hybrid environments.\n\n - Develop a comprehensive observability stack (monitoring, logging, tracing) to ensure platform health, and create automation for operational tasks, incident response, and performance tuning.\n\n - Collaborate with AI researchers and ML engineers to understand their infrastructure needs and build the tools and workflows that accelerate their development cycle.\n\n - Automate the life cycle of single-tenant, managed deployments\n\n\n\nYou’ll Love This Role If You\n\n - Are passionate about building platforms that empower developers and researchers.\n\n - Enjoy creating elegant, automated solutions for complex infrastructure challenges in both cloud and data center environments.\n\n - Thrive on optimizing hybrid infrastructure for performance, cost, and reliability.\n\n - Are excited to work at the intersection of modern platform engineering and cutting-edge AI.\n\n - Love to treat infrastructure as a product, continuously improving the developer experience.\n\n\n\nIt’s Important To Us That You Have\n\n - 5+ years of experience in Platform Engineering, DevOps, or Site Reliability Engineering (SRE).\n\n - Proven, hands-on experience building and managing production infrastructure with Terraform.\n\n - Expert-level knowledge of Kubernetes architecture and operations in a large-scale environment.\n\n - Strong scripting and automation skills (e.g., Python, Go, Bash).\n\n - Experience with CI/CD systems (e.g., GitLab CI, Jenkins, ArgoCD) and building developer tooling.\n\n  \n\nIt Would Be Great if You Had \n\n - Experience with high-performance compute (HPC) job schedulers, specifically Slurm, for managing GPU-intensive AI workloads.\n\n - Experience managing bare metal infrastructure, including server provisioning (e.g., PXE boot, MAAS), config","location":"United States","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"senior","tags":["cloud","speech","generative-ai","machine-learning","kubernetes","platform","infrastructure"],"apply_url":"https://jobs.ashbyhq.com/deepgram/f424ef6a-c27f-4984-9e77-40a1ad16ae28/application","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-07-03T18:10:26.112Z","expires_at":"2026-09-29T13:34:51.1558Z","created_at":"2026-04-13T09:40:05.854489Z","updated_at":"2026-08-30T13:34:51.292903Z","company_name":"Deepgram","company_slug":"deepgram","company_logo_url":"https://www.google.com/s2/favicons?domain=deepgram.com\u0026sz=128","quality_score":60,"url":"https://aidevboard.com/job/134572d2-b112-45db-8b0e-fe6ebb53bbcc"},{"id":"3edbec0f-7c89-48a6-8cb4-552274585aa0","company_id":"a0b04b48-9259-414d-93bd-ae677520bef1","title":"Software Infrastructure Kubernetes Engineer ","slug":"software-infrastructure-kubernetes-engineer-a7ebbd87","description":"About Graphcore  \n Graphcore is one of the world’s leading innovators in Artificial Intelligence compute.  \n It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry.  \n As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone.   \n Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation.  \n Summary  \n Join our dynamic Software Infrastructure team and take a pivotal role in scaling and managing our infrastructure. You will develop essential tools and services that empower our broader software team. Your contributions will enhance the build, test, deployment, and productisation processes of our Machine Learning Software components. Work with our High-Performance Computing (HPC) AI platforms and gain invaluable experience in distributed system\n The Team \n The Software Infrastructure team provides critical platforms and services for software development teams across the business. Our responsibilities include managing the CI platform and services, build engineering, component integration, and packaging and release systems. We operate in squads, fostering a culture of service ownership and empowerment for our engineers. We focus on long-term engineering solutions and strive to eliminate toil wherever possible.   \n Responsibilities and Duties   \n \n Develop, own, and maintain tools and services to support the software org   \n \n \n Deploy and maintain Kubernetes infrastructure to develop, test, and scale Graphcore hardware and its software stack   \n \n \n Manage our Cloud Infrastructure using tools such as Terraform   \n \n Candidate Profile   \n Essential:   \n \n Practical experience developing in Go   \n \n \n Familiarity with cloud services (AWS preferred)   \n \n \n Experience managing or developing in Linux environments   \n \n \n Understanding of CI/CD principles   \n \n \n Strong experience of Kubernetes (k8s) development and deployment   \n \n Desirable   \n \n Experience developing Kubernetes Controllers   \n \n \n Experience with Infrastructure as Code (IaC) tools (e.g. Terraform/OpenTofu)   \n \n \n Experience with GitHub Actions   \n \n \n Experience with distributed HPC systems   \n \n \n Experience with modern observability tooling (e.g. Prometheus)   \n \n \n Knowledge of Python/C++ (or similar language)   \n \n Benefits \n In addition to a competitive salary, Graphcore offers flexible working, a generous annual leave policy, private medical insurance and health cash plan, a dental plan, pension (matched up to 5%), life assurance and income protection. We have a generous parental leave policy and an employee assistance programme (which includes health, mental wellbeing, and bereavement support). We offer a range of healthy food and snacks at our central Bristol office and have our own barista bar! We welcome people of different backgrounds and experiences; we’re committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments.\n Applicants for this position must hold the right to work in the UK. Unfortunately at this time, we are unable to provide visa sponsorship or support for visa applications","location":"London, UK","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"mid","tags":["cloud","distributed-systems","kubernetes","infrastructure"],"apply_url":"https://job-boards.greenhouse.io/graphcore/jobs/8605734002","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-06-24T14:23:57Z","expires_at":"2026-09-29T13:43:48.540398Z","created_at":"2026-06-28T14:13:06.934074Z","updated_at":"2026-08-30T13:43:48.672968Z","company_name":"Graphcore","company_slug":"graphcore","company_logo_url":"https://www.google.com/s2/favicons?domain=graphcore.ai\u0026sz=128","quality_score":60,"url":"https://aidevboard.com/job/3edbec0f-7c89-48a6-8cb4-552274585aa0"},{"id":"745b1f76-7018-4e26-9515-3cedc12affcc","company_id":"a0b04b48-9259-414d-93bd-ae677520bef1","title":"Software Infrastructure Kubernetes Engineer ","slug":"software-infrastructure-kubernetes-engineer-1e340fca","description":"About Graphcore  \n Graphcore is one of the world’s leading innovators in Artificial Intelligence compute.  \n It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry.  \n As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone.   \n Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation.  \n Summary  \n Join our dynamic Software Infrastructure team and take a pivotal role in scaling and managing our infrastructure. You will develop essential tools and services that empower our broader software team. Your contributions will enhance the build, test, deployment, and productisation processes of our Machine Learning Software components. Work with our High-Performance Computing (HPC) AI platforms and gain invaluable experience in distributed system\n The Team \n The Software Infrastructure team provides critical platforms and services for software development teams across the business. Our responsibilities include managing the CI platform and services, build engineering, component integration, and packaging and release systems. We operate in squads, fostering a culture of service ownership and empowerment for our engineers. We focus on long-term engineering solutions and strive to eliminate toil wherever possible.   \n Responsibilities and Duties   \n \n Develop, own, and maintain tools and services to support the software org   \n \n \n Deploy and maintain Kubernetes infrastructure to develop, test, and scale Graphcore hardware and its software stack   \n \n \n Manage our Cloud Infrastructure using tools such as Terraform   \n \n Candidate Profile   \n Essential:   \n \n Practical experience developing in Go   \n \n \n Familiarity with cloud services (AWS preferred)   \n \n \n Experience managing or developing in Linux environments   \n \n \n Understanding of CI/CD principles   \n \n \n Strong experience of Kubernetes (k8s) development and deployment   \n \n Desirable   \n \n Experience developing Kubernetes Controllers   \n \n \n Experience with Infrastructure as Code (IaC) tools (e.g. Terraform/OpenTofu)   \n \n \n Experience with GitHub Actions   \n \n \n Experience with distributed HPC systems   \n \n \n Experience with modern observability tooling (e.g. Prometheus)   \n \n \n Knowledge of Python/C++ (or similar language)   \n \n Benefits \n In addition to a competitive salary, Graphcore offers flexible working, a generous annual leave policy, private medical insurance and health cash plan, a dental plan, pension (matched up to 5%), life assurance and income protection. We have a generous parental leave policy and an employee assistance programme (which includes health, mental wellbeing, and bereavement support). We offer a range of healthy food and snacks at our central Bristol office and have our own barista bar! We welcome people of different backgrounds and experiences; we’re committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments.\n Applicants for this position must hold the right to work in the UK. Unfortunately at this time, we are unable to provide visa sponsorship or support for visa applications","location":"Cambridge, UK","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"mid","tags":["cloud","distributed-systems","infrastructure","kubernetes"],"apply_url":"https://job-boards.greenhouse.io/graphcore/jobs/8605733002","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-06-24T14:23:22Z","expires_at":"2026-09-29T13:43:48.632947Z","created_at":"2026-06-28T14:13:07.097496Z","updated_at":"2026-08-30T13:43:48.765491Z","company_name":"Graphcore","company_slug":"graphcore","company_logo_url":"https://www.google.com/s2/favicons?domain=graphcore.ai\u0026sz=128","quality_score":60,"url":"https://aidevboard.com/job/745b1f76-7018-4e26-9515-3cedc12affcc"},{"id":"5faa90ac-966c-4f0a-84d5-b65964818b3a","company_id":"a0b04b48-9259-414d-93bd-ae677520bef1","title":"Software Infrastructure Kubernetes Engineer ","slug":"software-infrastructure-kubernetes-engineer-def89290","description":"About Graphcore  \n Graphcore is one of the world’s leading innovators in Artificial Intelligence compute.  \n It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry.  \n As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone.   \n Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation.  \n Summary  \n Join our dynamic Software Infrastructure team and take a pivotal role in scaling and managing our infrastructure. You will develop essential tools and services that empower our broader software team. Your contributions will enhance the build, test, deployment, and productisation processes of our Machine Learning Software components. Work with our High-Performance Computing (HPC) AI platforms and gain invaluable experience in distributed system\n The Team \n The Software Infrastructure team provides critical platforms and services for software development teams across the business. Our responsibilities include managing the CI platform and services, build engineering, component integration, and packaging and release systems. We operate in squads, fostering a culture of service ownership and empowerment for our engineers. We focus on long-term engineering solutions and strive to eliminate toil wherever possible.   \n Responsibilities and Duties   \n \n Develop, own, and maintain tools and services to support the software org   \n \n \n Deploy and maintain Kubernetes infrastructure to develop, test, and scale Graphcore hardware and its software stack   \n \n \n Manage our Cloud Infrastructure using tools such as Terraform   \n \n Candidate Profile   \n Essential:   \n \n Practical experience developing in Go   \n \n \n Familiarity with cloud services (AWS preferred)   \n \n \n Experience managing or developing in Linux environments   \n \n \n Understanding of CI/CD principles   \n \n \n Strong experience of Kubernetes (k8s) development and deployment   \n \n Desirable   \n \n Experience developing Kubernetes Controllers   \n \n \n Experience with Infrastructure as Code (IaC) tools (e.g. Terraform/OpenTofu)   \n \n \n Experience with GitHub Actions   \n \n \n Experience with distributed HPC systems   \n \n \n Experience with modern observability tooling (e.g. Prometheus)   \n \n \n Knowledge of Python/C++ (or similar language)   \n \n Benefits \n In addition to a competitive salary, Graphcore offers annual leave policy, medical and dental health plans, a gym card, and employee pension (matched up to 4%). We review our benefits on a yearly basis to ensure we offer a valuable and rewarding benefits programme to our employees. We welcome people of different backgrounds and experiences; we’re committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments.","location":"Gdańsk, Pomeranian Voivodeship, Poland","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"mid","tags":["distributed-systems","cloud","kubernetes","infrastructure"],"apply_url":"https://job-boards.greenhouse.io/graphcore/jobs/8605723002","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-06-24T14:23:20Z","expires_at":"2026-09-29T13:43:48.356893Z","created_at":"2026-06-28T14:13:07.182068Z","updated_at":"2026-08-30T13:43:48.486037Z","company_name":"Graphcore","company_slug":"graphcore","company_logo_url":"https://www.google.com/s2/favicons?domain=graphcore.ai\u0026sz=128","quality_score":60,"url":"https://aidevboard.com/job/5faa90ac-966c-4f0a-84d5-b65964818b3a"},{"id":"35f31c85-304c-4e62-bf4e-be7bf77daa15","company_id":"a0000000-0000-0000-0000-000000000001","title":"Staff Software Engineer, Kubernetes Platform","slug":"staff-software-engineer-kubernetes-platform-44feb137","description":"About Anthropic \n Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.\n About the role \n We run one of the industry's largest AI compute fleets, spanning multiple cloud providers and datacenters, to train, research, and serve frontier AI models. Those fleets run on Kubernetes, and the Kubernetes Platform team owns the control plane that makes them work.\n We are operating at a scale where the defaults stop working. We own the scheduler and extend it to place topology-sensitive ML workloads across thousands of accelerators at once. We scale the control plane itself — apiserver, etcd, controllers — so it stays responsive as object counts and node counts grow by orders of magnitude. And we build the core cluster services every workload depends on, like service discovery, so they hold up under the same pressure.\n We make sure the control plane is fast, correct, and always available. Your work will directly determine whether Anthropic can keep reliably and safely training frontier models as our compute footprint continues to grow.\n Key responsibilities \n \n Own, operate, and extend the Kubernetes scheduler for Anthropic's accelerator fleets, including custom scheduling plugins and policies for gang scheduling, topology awareness, and preemption\n Scale the Kubernetes control plane (apiserver, etcd, controller-manager) to support clusters far beyond typical limits, and find the next bottleneck before it finds us\n Design, build, and operate core cluster services such as service discovery that every workload in the fleet depends on\n Build and maintain custom controllers, operators, and CRDs\n Partner with research, training, and inference to understand workload shapes and turn their requirements into platform capabilities\n Collaborate with cloud providers on required features and escalations\n Participate in on-call, lead incident response, and design processes (postmortems, runbooks, SLOs) that help the team avoid repeating failures\n \n Minimum qualifications \n \n Significant software engineering experience building and operating production distributed systems\n Proficiency in at least one systems-appropriate language (e.g., Go, Python, Rust, or C++)\n Deep, hands-on Kubernetes experience (well beyond \"user of”) into scheduler, controllers, apiserver, or operating large multi-tenant clusters\n Demonstrated ability to debug complex issues across the stack, from API behavior down to node and network-level root causes\n A track record of designing for reliability, correctness, and clear failure semantics in systems other engineers depend on\n Strong written and verbal communication; comfort building consensus with internal stakeholders\n \n Preferred qualifications \n \n Experience with Kubernetes internals or contributions: kube-scheduler / scheduling framework, apiserver, etcd, client-go, controller-runtime, or similar\n Experience building or operating cluster schedulers or batch systems (e.g., Kueue, Volcano, Slurm, or in-house equivalents)\n Background scaling control planes or coordination systems (etcd, ZooKeeper, Consul, or large DNS/service-mesh deployments)\n Familiarity with ML infrastructure: GPUs, TPUs, or Trainium; gang scheduling; topology-aware placement; collective networking such as NCCL\n Experience with GCP and/or AWS, including GKE/EKS internals and Infrastructure as Code\n Low-level systems experience such as Linux kernel tuning, cgroups, or eBPF\n 12+ years of relevant industry experience, including time leading large, ambiguous infrastructure projects\n The annual compensation range for this role is listed below. \n For sales roles, the range provided is the role’s On Target Earnings (\"OTE\") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role.\n Annual Salary:\n £325,000 — £485,000 GBP \n Logistics \n Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience\n Required field of study:  A field relevant to the role as demonstrated through coursework, training, or professional experience\n Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position\n Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices.\n Visa sponsorship:  We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this.\n We encourage you to apply even if you do not believe you meet every s","location":"London, UK","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["cloud","alignment","distributed-systems","kubernetes","infrastructure"],"apply_url":"https://job-boards.greenhouse.io/anthropic/jobs/5211305008","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-05-07T09:54:50Z","expires_at":"2026-09-29T13:30:40.782946Z","created_at":"2026-05-07T14:00:27.744446Z","updated_at":"2026-08-30T13:30:40.926729Z","company_name":"Anthropic","company_slug":"anthropic","company_logo_url":"https://www.google.com/s2/favicons?domain=anthropic.com\u0026sz=128","quality_score":60,"url":"https://aidevboard.com/job/35f31c85-304c-4e62-bf4e-be7bf77daa15"},{"id":"53f0c2dc-1441-4135-8ea6-fede8b6a6baf","company_id":"a0b04b48-9259-414d-93bd-ae677520bef1","title":"Software Infrastructure Kubernetes Engineer ","slug":"software-infrastructure-kubernetes-engineer-d4c3baf3","description":"About Graphcore  \n Graphcore is one of the world’s leading innovators in Artificial Intelligence compute.  \n It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry.  \n As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone.   \n Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation.  \n Summary  \n Join our dynamic Software Infrastructure team and take a pivotal role in scaling and managing our infrastructure. You will develop essential tools and services that empower our broader software team. Your contributions will enhance the build, test, deployment, and productisation processes of our Machine Learning Software components. Work with our High-Performance Computing (HPC) AI platforms and gain invaluable experience in distributed system\n The Team \n The Software Infrastructure team provides critical platforms and services for software development teams across the business. Our responsibilities include managing the CI platform and services, build engineering, component integration, and packaging and release systems. We operate in squads, fostering a culture of service ownership and empowerment for our engineers. We focus on long-term engineering solutions and strive to eliminate toil wherever possible.   \n Responsibilities and Duties   \n \n Develop, own, and maintain tools and services to support the software org   \n \n \n Deploy and maintain Kubernetes infrastructure to develop, test, and scale Graphcore hardware and its software stack   \n \n \n Manage our Cloud Infrastructure using tools such as Terraform   \n \n Candidate Profile   \n Essential:   \n \n Practical experience developing in Go   \n \n \n Familiarity with cloud services (AWS preferred)   \n \n \n Experience managing or developing in Linux environments   \n \n \n Understanding of CI/CD principles   \n \n \n Strong experience of Kubernetes (k8s) development and deployment   \n \n Desirable   \n \n Experience developing Kubernetes Controllers   \n \n \n Experience with Infrastructure as Code (IaC) tools (e.g. Terraform/OpenTofu)   \n \n \n Experience with GitHub Actions   \n \n \n Experience with distributed HPC systems   \n \n \n Experience with modern observability tooling (e.g. Prometheus)   \n \n \n Knowledge of Python/C++ (or similar language)   \n \n Benefits \n In addition to a competitive salary, Graphcore offers flexible working, a generous annual leave policy, private medical insurance and health cash plan, a dental plan, pension (matched up to 5%), life assurance and income protection. We have a generous parental leave policy and an employee assistance programme (which includes health, mental wellbeing, and bereavement support). We offer a range of healthy food and snacks at our central Bristol office and have our own barista bar! We welcome people of different backgrounds and experiences; we’re committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments.\n Applicants for this position must hold the right to work in the UK. Unfortunately at this time, we are unable to provide visa sponsorship or support for visa applications","location":"Bristol, UK","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"mid","tags":["distributed-systems","cloud","infrastructure","kubernetes"],"apply_url":"https://job-boards.greenhouse.io/graphcore/jobs/8420432002","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-02-12T16:03:26Z","expires_at":"2026-09-29T13:43:48.446578Z","created_at":"2026-04-16T14:45:51.882574Z","updated_at":"2026-08-30T13:43:48.577912Z","company_name":"Graphcore","company_slug":"graphcore","company_logo_url":"https://www.google.com/s2/favicons?domain=graphcore.ai\u0026sz=128","quality_score":60,"url":"https://aidevboard.com/job/53f0c2dc-1441-4135-8ea6-fede8b6a6baf"},{"id":"9a6520c0-2554-494a-9f5c-bb70ade691c0","company_id":"7551b4ca-b2b0-493a-ab58-a15bd9c50393","title":"Infrastructure Software Engineer (Kubernetes)","slug":"staff-infrastructure-software-engineer-kubernetes-e6290ed9","description":"Cresta unlocks the true potential of the customer experience, turning every conversation into a competitive advantage. Cresta’s unified AI platform combines conversational AI agents, real-time human agent augmentation, and comprehensive conversation intelligence to drive revenue and efficiency gains across every channel. The world’s leading companies, including United Airlines, Cox Communications, and Marriott, use Cresta to power world-class customer experiences every day. \n Born from the Stanford AI Lab, Cresta has raised more than $270 million from the world’s leading investors, including a16z, Greylock, and Sequoia. Cresta’s leadership includes some of the leading minds in AI today. Our CEO, Ping Wu , founded and led Google's Contact Center AI and Vertex AI platforms before joining Cresta to build the future of AI-driven customer experiences.\n Over the next few years, AI is going to redefine how people all over the world interact with businesses every day. Come build that future at Cresta.\n \n About the role: \n As a member of the infrastructure team you are responsible for designing, building, and advancing our core infrastructure that allows the engineering team to execute quickly, productively, and securely. You will join a collaborative but highly autonomous working environment in which each member has a defined role with clear expectations, as well as the freedom to pursue projects they find interesting.\n Responsibilities:\n \n Developer Toolchain . Partner with engineers to build dev tools that empower developer workflows and deployment infrastructure.\n Ensure reliability of multi-cloud Kubernetes clusters and pipelines.\n Metrics, logging, analytics, and alerting for performance and security across all endpoints and applications.\n Infrastructure-as-code deployment tooling and supporting services on multiple cloud providers.\n Automate operations and engineering . Focus on automation so we can spend energy where it matters.\n Building machine learning infrastructure that enables AI teams to train, test, and deploy on large-scale datasets.\n \n What we are looking for:\n \n \n 5+ years experience in DevOps, Site Reliability Engineering, Production Engineering, or equivalent field.\n Deep proficiency with coding languages such as Golang or Python.\n Deep familiarity with container-related security best practices.\n Production experience working with Kubernetes, and a deep understanding of the Kubernetes ecosystem, including popular open-source tooling such as cert-manager or external-dns.  Experience with GPU-enabled clusters is a bonus.\n Production experience with Kubernetes templating tools such as Helm or Kustomize.\n Production experience with IAC tools such as Terraform or CloudFormation.\n Production experience working with AWS and services such as IAM, S3, EC2, and EKS.\n Production experience with other cloud providers such as Google Cloud and Azure is a bonus.\n Production experience with database software such as PostgreSQL\n Experience with GitOps tooling such as Flux or Argo.\n Experience with CI/CD such as GitHub Actions.\n \n Compensation for this position includes a base salary, equity, and a variety of benefits. Actual base salaries will be based on candidate-specific factors, including experience, skillset, and location, and local minimum pay requirements as applicable. \n We have noticed a rise in recruiting impersonations across the industry, where scammers attempt to access candidates' personal and financial information through fake interviews and offers. All Cresta recruiting email communications will always come from the @cresta.ai domain. Any outreach claiming to be from Cresta via other sources should be ignored.  If you are uncertain whether you have been contacted by an official Cresta employee, reach out to recruiting@cresta.ai","location":"Romania","workplace":"remote","remote_scope":"unknown","job_type":"full-time","experience_level":"senior","tags":["cloud","agents","infrastructure","kubernetes"],"apply_url":"https://job-boards.greenhouse.io/cresta/jobs/4802840008","is_featured":false,"is_sticky":false,"status":"active","published_at":"2025-07-11T10:13:31Z","expires_at":"2026-09-29T13:34:38.737261Z","created_at":"2026-04-14T14:04:56.974846Z","updated_at":"2026-08-30T13:34:38.875678Z","company_name":"Cresta","company_slug":"cresta","company_logo_url":"https://www.google.com/s2/favicons?domain=cresta.com\u0026sz=128","quality_score":60,"url":"https://aidevboard.com/job/9a6520c0-2554-494a-9f5c-bb70ade691c0"},{"id":"01385ff5-12a0-4394-a2ae-29187bfc3b1a","company_id":"7551b4ca-b2b0-493a-ab58-a15bd9c50393","title":"Infrastructure Software Engineer (Kubernetes)","slug":"staff-infrastructure-software-engineer-kubernetes-d3988d09","description":"Cresta unlocks the true potential of the customer experience, turning every conversation into a competitive advantage. Cresta’s unified AI platform combines conversational AI agents, real-time human agent augmentation, and comprehensive conversation intelligence to drive revenue and efficiency gains across every channel. The world’s leading companies, including United Airlines, Cox Communications, and Marriott, use Cresta to power world-class customer experiences every day. \n Born from the Stanford AI Lab, Cresta has raised more than $270 million from the world’s leading investors, including a16z, Greylock, and Sequoia. Cresta’s leadership includes some of the leading minds in AI today. Our CEO, Ping Wu , founded and led Google's Contact Center AI and Vertex AI platforms before joining Cresta to build the future of AI-driven customer experiences.\n Over the next few years, AI is going to redefine how people all over the world interact with businesses every day. Come build that future at Cresta.\n \n About the role: \n As a member of the infrastructure team you are responsible for designing, building, and advancing our core infrastructure that allows the engineering team to execute quickly, productively, and securely. You will join a collaborative but highly autonomous working environment in which each member has a defined role with clear expectations, as well as the freedom to pursue projects they find interesting.\n Responsibilities:\n \n Developer Toolchain . Partner with engineers to build dev tools that empower developer workflows and deployment infrastructure.\n Ensure reliability of multi-cloud Kubernetes clusters and pipelines.\n Metrics, logging, analytics, and alerting for performance and security across all endpoints and applications.\n Infrastructure-as-code deployment tooling and supporting services on multiple cloud providers.\n Automate operations and engineering . Focus on automation so we can spend energy where it matters.\n Building machine learning infrastructure that enables AI teams to train, test, and deploy on large-scale datasets.\n \n What we are looking for:\n \n \n 5+ years experience in DevOps, Site Reliability Engineering, Production Engineering, or equivalent field.\n Deep proficiency with coding languages such as Golang or Python.\n Deep familiarity with container-related security best practices.\n Production experience working with Kubernetes, and a deep understanding of the Kubernetes ecosystem, including popular open-source tooling such as cert-manager or external-dns.  Experience with GPU-enabled clusters is a bonus.\n Production experience with Kubernetes templating tools such as Helm or Kustomize.\n Production experience with IAC tools such as Terraform or CloudFormation.\n Production experience working with AWS and services such as IAM, S3, EC2, and EKS.\n Production experience with other cloud providers such as Google Cloud and Azure is a bonus.\n Production experience with database software such as PostgreSQL\n Experience with GitOps tooling such as Flux or Argo.\n Experience with CI/CD such as GitHub Actions.\n \n Perks \u0026 Benefits: \n \n Paid parental leave to support you and your family\n Monthly Health \u0026 Wellness allowance\n PTO: 28 days in Berlin \n \n Compensation for this position includes a base salary, equity, and a variety of benefits. Actual base salaries will be based on candidate-specific factors, including experience, skillset, and location, and local minimum pay requirements as applicable. Your recruiter can provide further details.\n We have noticed a rise in recruiting impersonations across the industry, where scammers attempt to access candidates' personal and financial information through fake interviews and offers. All Cresta recruiting email communications will always come from the @cresta.ai domain. Any outreach claiming to be from Cresta via other sources should be ignored.  If you are uncertain whether you have been contacted by an official Cresta employee, reach out to recruiting@cresta.ai","location":"Germany","workplace":"remote","remote_scope":"unknown","job_type":"full-time","experience_level":"senior","tags":["agents","cloud","kubernetes","infrastructure"],"apply_url":"https://job-boards.greenhouse.io/cresta/jobs/4535898008","is_featured":false,"is_sticky":false,"status":"active","published_at":"2025-02-13T00:26:00Z","expires_at":"2026-09-29T13:34:38.832328Z","created_at":"2026-04-14T14:04:56.897063Z","updated_at":"2026-08-30T13:34:38.96853Z","company_name":"Cresta","company_slug":"cresta","company_logo_url":"https://www.google.com/s2/favicons?domain=cresta.com\u0026sz=128","quality_score":60,"url":"https://aidevboard.com/job/01385ff5-12a0-4394-a2ae-29187bfc3b1a"},{"id":"f58187fb-02e2-46eb-af57-49be218405d8","company_id":"dccc92b1-e96d-42a6-b302-5ec74e525e12","title":"Senior Staff Infrastructure Engineer – Kubernetes Platform","slug":"staff-infrastructure-engineer-kubernetes-platform-7523b645","description":"About TensorWave\n\nOur mission is simple: deliver seamless, secure, reliable, and resilient AI compute at scale. We've built a versatile cloud platform that eliminates infrastructure barriers, empowering builders to focus on innovation instead of fighting their stack. Because breakthrough AI should move at the speed of ideas, not infrastructure.\n\n \n\nAbout the Role\n\nWe’re looking for a Kubernetes Platform Staff Infrastructure Engineer to join our team during an exciting phase of growth. In this role, you’ll be responsible for owning the design, evolution, and operational reliability of our Kubernetes control plane architecture, working closely with cross-functional partners to support business objectives while upholding our standards for excellence, collaboration, and impact.\n\n \n\nWhat You’ll Do\n\nPlatform Architecture \u0026 Strategy\n\n - Design and evolve Kubernetes control plane architecture across regions\n\n - Define and implement multi-tenant cluster models, including shared control planes, virtual cluster approaches (e.g., vcluster, Kamaji)\n\n - Drive transition from standalone clusters to regionally managed platform models\n\n - Define standards for isolation boundaries, resource segmentation, policy enforcement\n\nPlatform Ownership \u0026 Operations\n\n - Own the reliability and behavior of Kubernetes platforms in production\n\n - Participate in on-call rotation and lead incident response\n\n - Diagnose and resolve control plane instability, API server saturation, scheduling and resource contention issues\n\n - Ensure consistent lifecycle management across clusters - provisioning, upgrades, scaling\n\nMulti-Region Scaling\n\n - Design and implement strategies for regional scaling, multi-data center cluster deployments\n\n - Ensure consistent behavior and reliability across environments\n\n - Define cluster topology and failure domain strategies\n\nNetworking \u0026 Data Plane Integration\n\n - Design ingress and egress architectures at cluster level and regional level\n\n - Troubleshoot and optimize pod-to-pod networking, north-south traffic flows, CNI behavior (Cilium preferred)\n\n - Collaborate with network engineering on high-performance networking integration\n\nObservability \u0026 Reliability\n\n - Improve observability across control plane components, cluster health and performance\n\n - Define and implement resilience strategies aligned with platform goals\n\n - Lead root cause analysis for production incidents\n\nCross-Team Collaboration\n\n - Work closely with DevOps engineers (automation and CI/CD) and Infrastructure teams (compute, storage, networking)\n\n - Align Kubernetes platform design with underlying infrastructure capabilities\n\n \n\nWho You Are\n\nRequired Qualifications\n\n - 7+ years of experience in infrastructure, platform engineering, or distributed systems\n\n - Deep experience operating Kubernetes at scale in production environments\n\n - Experience in CSP, hyperscale, or equivalent large-scale environments strongly preferred\n\n - Proven experience scaling Kubernetes across:\n   \n   - Multiple clusters\n   \n   - Multiple regions or data centers\n\n - Strong understanding of Kubernetes internals:\n   \n   - API server\n   \n   - Scheduler\n   \n   - Controller manager\n   \n   - etcd\n\n - Experience designing or evolving:\n   \n   - Control plane architectures\n   \n   - Multi-tenant cluster models\n\nTechnical Depth\n\n - Strong Linux systems expertise\n\n - Deep troubleshooting ability across:\n   \n   - Kubernetes\n   \n   - Container runtime\n   \n   - Networking stack\n\n - Experience with CNI plugins (Cilium preferred)\n\n - Strong understanding of:\n   \n   - Networking and traffic patterns\n   \n   - Resource isolation and scheduling\n\nPreferred Qualifications\n\n - Experience with virtual cluster technologies (vcluster, Kamaji, or similar)\n\n - Experience supporting GPU workloads in Kubernetes\n\n - Familiarity with:\n   \n   - NUMA-aware scheduling\n   \n   - Topology-aware workloads\n\n - Awareness of RDMA and high-throughput networking environments\n\n - Experience with observability platforms (Prometheus, Grafana, etc.)\n\n \n\nWhat We Offer\n\n - Stock Options\n\n - 100% paid Medical, Dental, and Vision insurance for Employees\n\n - Company Health Savings Account Contributions\n\n - 100% paid Short Term and Long Term Disability Insurance for Employees\n\n - Life and Voluntary Supplemental Insurance Options\n\n - Other Insurance Options, such as Pet \u0026 Legal Insurance\n\n - Various Supplementary Health Benefits, such as discounted Virtual Healthcare Appointments and Serious Illness Support\n\n - Flexible Spending Account\n\n - 401(k)\n\n - Employee Assistance Program\n\n - Flexible PTO\n\n - Paid Holidays\n\n - Parental Leave\n\n - Other In-Office Perks\n\n \n\nEqual Employment Opportunity\n\nTensorWave is an Equal Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate on the basis of any protected status under applicable law.\n\n \n\nReasonable Accommodations\n\nTensorWave provides reasonable accommodations in accordance with applicab","location":"Remote","workplace":"remote","remote_scope":"unknown","job_type":"full-time","experience_level":"lead","tags":["healthcare","distributed-systems","infrastructure","kubernetes","platform"],"apply_url":"https://jobs.ashbyhq.com/tensorwave/4932c835-6d13-4770-b3e6-802bb48cfc57/application","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-07-08T15:43:54.678Z","expires_at":"2026-09-29T13:48:37.557564Z","created_at":"2026-05-12T14:21:51.804779Z","updated_at":"2026-08-30T13:48:37.728765Z","company_name":"TensorWave","company_slug":"tensorwave","company_logo_url":"https://www.google.com/s2/favicons?domain=tensorwave.com\u0026sz=128","quality_score":55,"url":"https://aidevboard.com/job/f58187fb-02e2-46eb-af57-49be218405d8"}],"page":1,"per_page":20,"total":10,"total_is_exact":true,"total_pages":1}
