{"access":{"catalog_url":"https://aidevboard.com/api/v1/catalog","description":"Public read endpoints are open and free. API keys are optional for stable agent identity and keyed hourly throttling.","docs_url":"https://aidevboard.com/docs","employer_pilot_url":"https://aidevboard.com/verified-interview-pilot","mode":"open","register_url":"https://aidevboard.com/api/v1/register"},"candidate_resume_action":{"application_authorized":false,"candidate_charge":0,"endpoint":"https://aidevboard.com/api/v1/candidate/resume-preview","job_id_json_path":"jobs[].id","method":"POST","preview_requires_identity":false,"required_body_fields":["job_id","evidence_bullets"],"requires_explicit_human_review":true,"saved_artifact_protocol":"mcp","saved_artifact_requires_verified_human":true,"saved_artifact_tool":"compile_job_specific_resume","search_requires_identity":false,"status":"available_after_candidate_selects_job","submission_performed":false,"uses_candidate_verified_evidence":true},"degraded":false,"estimated":false,"has_next":true,"jobs":[{"id":"4c376d37-cd6f-463d-a12e-7f7436246713","company_id":"e8c9f3a5-9310-43f5-9341-321fe6d93a92","title":"Site Reliability Engineering Manager, Vehicle Software","slug":"site-reliability-engineering-manager-vehicle-software-967fc97f","description":"About us    \n Founded in 2017, Wayve is the leading developer of Embodied AI technology.  Our advanced AI software and foundation models enable vehicles to perceive, understand, and navigate any complex environment, enhancing the usability and safety of automated driving systems.\n Our vision is to create autonomy that propels the world forward.  Our intelligent, mapless, and hardware-agnostic AI products are designed for automakers, accelerating the transition from assisted to automated driving.  In our fast-paced environment big problems ignite us—we embrace uncertainty, leaning into complex challenges to unlock groundbreaking solutions. We aim high and stay humble in our pursuit of excellence, constantly learning and evolving as we pave the way for a smarter, safer future.\n At Wayve, your contributions matter.  We value diversity, embrace new perspectives, and foster an inclusive work environment; we back each other to deliver impact.  \n Make Wayve the experience that defines your career!  \n The role\n As SRE Manager, you'll build the Vehicle Software SRE team from the ground up — defining its charter, hiring its founding engineers, establishing the operating model, and creating the technical strategy that makes reliability a first-class property of the software running on our vehicles.\n You'll work in a production environment unlike most: a globally distributed fleet of autonomous vehicles operating at the intersection of software, hardware, networking, sensors, and the physical world. Failures are often intermittent, hard to reproduce, and distributed across ownership boundaries. You'll move reliability upstream — from reactive field support to prevention through architecture, automation, observability, and disciplined production readiness.\n You'll embed your team within Vehicle Software, partnering with product teams who retain ownership of what they build while your team provides the reliability engineering, standards, and leverage that help them operate fleet-critical software safely at scale. You'll stay hands-on throughout — writing code, reviewing critical designs, and leading the investigations that matter most.\n The systems you help harden will connect Wayve's AI to physical vehicles and underpin the transition from engineering fleets to commercial operations. Few engineering leadership roles offer this combination of zero-to-one team building, deep systems work, and direct influence on the safety and scalability of autonomous mobility.\n  \n Key Responsibilities\n \n Team building \u0026 leadership : Build and lead a new SRE team from the ground up, staying hands-on as a player-coach on the team's most consequential work.\n Reliability strategy: Own technical direction for vehicle software reliability across deployment, service health, telemetry, and diagnostics; define what production-ready means at Wayve.\n Production readiness: Define SLIs, SLOs, and error budgets for fleet-critical workflows; drive release criteria, automated gates, rollback strategies, and fault-injection practices across Vehicle Software.\n Observability \u0026 tooling: Design and implement the observability and automation that shortens the path from vehicle symptom to root cause, cuts the manual toil between failure and fix, and shapes systems for robustness, recoverability, and debuggability.\n Incident response: Lead investigations into complex failures, ensure every incident produces a durable fix, and strengthen on-call practices and escalation paths across service-owning teams.\n Mentorship \u0026 communication: Mentor engineers and emerging leaders, and give senior leadership the clarity on reliability health, risks, and investment they need to make good decisions.\n \n  \n About you\n Essential \n \n 8+ years building and operating production software systems with strong depth in SRE, production engineering, platform engineering, embedded systems, or robotics, and a recent track record of writing production-quality code and leading architecture reviews across Linux-based, distributed, or hardware-software systems.\n 3+ years in people leadership with a track record of hiring, coaching, and growing engineers across levels while staying actively engaged in coding, design, and code review; experience forming a new team or capability from scratch is a strong plus.\n Proven experience with SLOs, error budgets, production-readiness standards, observability, incident management, postmortems, and toil-reduction programmes, with measurable outcomes to show for it.\n A track record of turning ambiguous, cross-functional problems into clear ownership, sequenced plans, and reliable delivery without relying on formal authority.\n Hands-on experience building production software, automation, and diagnostic tooling in C++, Rust, Python, or Go, with familiarity with CI/CD, release systems, telemetry pipelines, and modern observability tooling.\n Calm and structured during incidents, with clear communication across software, hardware, operations","salary_min":276100,"salary_max":311400,"location":"Sunnyvale, CA","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["autonomous-vehicles","generative-ai","robotics","devops"],"apply_url":"https://wayve.firststage.co/jobs?gh_jid=8728803002","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-25T20:00:30Z","expires_at":"2026-09-29T13:43:28.958489Z","created_at":"2026-08-26T13:43:37.362178Z","updated_at":"2026-08-30T13:43:29.088912Z","company_name":"Wayve","company_slug":"wayve","company_logo_url":"https://www.google.com/s2/favicons?domain=wayve.ai\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/4c376d37-cd6f-463d-a12e-7f7436246713"},{"id":"3707734c-f7ec-4eee-bd06-3bd04c3d3355","company_id":"52f44519-9f93-4eac-ae0b-8be13e385ebe","title":"Cloud DevOps Engineer","slug":"cloud-devops-engineer-c42c2a76","description":"CLOUD DEVOPS ENGINEER\n\n\n\nYou'll build the cloud infrastructure that turns the open web into data — the platform beneath Firecrawl's crawling, scraping, and search products. We need engineers who can make that foundation — Kubernetes, storage, networking, deployments — fast, reliable, and cheap at web scale. You'll own real infrastructure from day one — not tickets in a backlog.\n\n \n\nSalary Range: $246,000–$271,000/year\n\nEquity Range: Competitive equity — details shared during the process.\n\nLocation: San Francisco, CA (Onsite)\n\nJob Type: Full-Time \n\nExperience: 5+ years in DevOps, Platform Engineering or Cloud Infrastructure \n\nVisa: Must be legally authorized to work in the United States. We're not able to sponsor visas right now, though that may change down the line.\n\n\n\n\nABOUT FIRECRAWL\n\nFirecrawl is the easiest way to turn the web into data AI agents can use. One API call converts any URL into clean, LLM-ready markdown or structured data - the boring-hard problem everyone building with LLMs eventually hits, solved.\n\nWe hit 8 figures in ARR in year one and more than doubled it in year two. We have 170k+ GitHub stars, and developers, agents, and category-defining AI companies build on us every day. Growth like this is rare, and we're just getting started.\n\nWe're a small team punching far above our weight. Everyone here owns a real piece of the product and company, end to end, and runs it themselves - no hiding behind process or headcount.\n\nThis is a place for people who want to work at the frontier: an AI company building the infrastructure other AI companies run on, not one bolting AI onto an existing product. We move fast, go deep, and are building the tools superintelligence will rely on to gather data from the web.\n\n\n\n\nWHAT YOU'LL DO\n\n - Design, build, and operate GCP and on-prem infrastructure behind Firecrawl's products\n\n - Run large stateful and high-throughput workloads on Kubernetes — search clusters, crawling fleets, queues, and databases — with zero-downtime upgrades\n\n - Own CI/CD, Infrastructure as Code, automations, and containerized deployments across the platform\n\n - Drive down infrastructure cost per request while traffic and data volume grow\n\n - Build the observability that keeps latency, throughput, and reliability predictable and define the SLIs, SLOs, and customer-facing SLAs we hold ourselves to.\n\n - Build and support our enterprise controls — SSO/SAML, RBAC, audit logging, tenant isolation, private networking, and the infrastructure behind SOC 2 and customer security reviews\n\n - Work directly with product and search engineers to productionize new services, retrieval, and ML workloads\n\n - Own the incident lifecycle with engineers — from on-call and triage process to postmortems and resolution.\n\n\n\n\nWHAT WE'RE LOOKING FOR\n\n - You've operated stateful distributed systems on Kubernetes at real scale — not just stateless services\n\n - You have deep experience with a major cloud (e.g. GCP, AWS, Azure), Docker, and Terraform; MLOps or ML-serving infrastructure experience (GPU workloads, model deployment pipelines) is a plus\n\n - You've run large-scale, data-heavy systems in production — search platforms, crawling or ingestion pipelines, or comparable. Hands-on experience operating Vespa https://github.com/vespa-engine/vespa is a strong plus.\n\n - You care about latency, cost, and reliability in equal measure\n\n - Experience with security and compliance infrastructure (SSO/SAML, audit logging, network isolation, SOC 2) is a strong plus — especially for the Core Platform focus\n\n - You're comfortable owning ambiguous problems and turning them into shipped infrastructure\n\n\n\n\nWHAT WE'RE NOT LOOKING FOR\n\n - Someone who needs a fully-specced ticket to start\n\n - Someone who wants to specialize narrowly and hand off everything else\n\n - Someone who optimizes for process over shipping\n\n\n\n\nA NOTE ON PACE\n\nWe operate at an absurd level of urgency because the window for what we're building won't stay open forever. If that excites you, keep reading. If it doesn't, no hard feelings — but this role probably isn't for you.\n\n\n\n\nBENEFITS \u0026 PERKS\n\n\n\n\nAVAILABLE TO ALL EMPLOYEES\n\n - Salary that makes sense — $246,000–$271,000/year, based on impact, not tenure\n\n - Own a piece — Gain competitive equity in what you're helping build\n\n - Generous PTO — 15 days mandatory, anything after 24 days, just ask (holidays excluded); take the time you need to recharge\n\n - Parental leave — 12 weeks fully paid, for all parents\n\n - Wellness stipend — $100/month for the gym, therapy, massages, or whatever keeps you human\n\n - Learning \u0026 Development — Expense up to $1,000/year toward anything that helps you grow professionally\n\n - Team offsites — A change of scenery, minus the trust falls\n\n - Sabbatical — 3 paid months off after 4 years, do something fun and new\n\n\n\n\nAVAILABLE TO US-BASED FULL-TIME EMPLOYEES\n\n - Full coverage, no red tape — Medical, dental, and vision (100% for employees, 50% fo","salary_min":246000,"salary_max":271000,"location":"San Francisco, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"senior","tags":["llm","distributed-systems","agents","mlops","search","cloud","devops"],"apply_url":"https://jobs.ashbyhq.com/firecrawl/fe538f2b-7dd5-4d8d-941e-f8014a911652/application","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-20T22:04:31.381Z","expires_at":"2026-09-29T13:45:44.202795Z","created_at":"2026-08-25T18:32:19.567715Z","updated_at":"2026-08-30T13:45:44.333747Z","company_name":"Firecrawl","company_slug":"firecrawl","company_logo_url":"https://www.google.com/s2/favicons?domain=firecrawl.dev\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/3707734c-f7ec-4eee-bd06-3bd04c3d3355"},{"id":"e83f9b8b-565f-4a94-b2f4-03b79a6f7843","company_id":"2114efab-ea67-411b-bfb8-7899153105f3","title":"Member of Technical Staff, Site Reliability Engineer","slug":"member-of-technical-staff-site-reliability-engineer-369fc8ef","description":"Overview\n\nInferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.\n\nAbout the Role\n\nWe're looking for a Site Reliability Engineer to help make vLLM-powered inference systems reliable, observable, and operationally simple at production scale. This role is for someone who thinks about failure before launch, designs systems that are easier to operate, and knows how to turn incidents into durable improvements rather than one-off fixes.\n\nYou'll work across engineering and infrastructure to define SLOs, improve monitoring and alerting, strengthen incident response, drive post-mortems, and reduce operational risk before it reaches users. Your work will directly impact the reliability, availability, and production readiness of the systems powering AI inference at scale.\n\n\n\nSkills and Qualifications\n\nMinimum qualifications:\n\n - Bachelor's degree or equivalent experience in computer science, engineering, systems, infrastructure, or similar.\n\n - Strong experience operating production systems with meaningful traffic, user impact, or infrastructure criticality.\n\n - Deep understanding of SLOs, SLIs, error budgets, alerting, incident response, and post-mortem processes.\n\n - Experience live-fighting major production incidents, including mitigation, root cause analysis, escalation, and follow-through on prevention work.\n\n - Strong Linux, networking, systems debugging, observability, and distributed systems fundamentals.\n\n - Ability to design operationally simple systems and identify likely failure modes before launch.\n\n - Strong programming or scripting ability in Python, Go, Bash, or similar for automation, tooling, and reliability improvements.\n\nPreferred qualifications:\n\n - Experience supporting ML infrastructure, inference systems, GPU workloads, Kubernetes-based platforms, or high-scale backend services.\n\n - Experience building or improving observability systems using metrics, logs, traces, dashboards, alerts, and runbooks.\n\n - Experience with Kubernetes, Docker, Terraform, cloud infrastructure, service meshes, CI/CD systems, or production deployment platforms.\n\n - Experience driving incident review culture, post-mortem processes, reliability reviews, and prevention-oriented engineering work.\n\n - Ability to partner with engineering teams to improve service design, release safety, capacity planning, and operational readiness.\n\nBonus points if you have:\n\n - Owned reliability for high-throughput, latency-sensitive, or mission-critical production systems.\n\n - Supported AI inference, model serving, GPU clusters, ML platforms, or distributed serving infrastructure.\n\n - Built automation that reduced toil, improved recovery time, or prevented repeat incidents.\n\n - Led incident response for severe outages with clear communication across engineering and leadership.\n\n - Created practical SLOs, dashboards, alerts, runbooks, or release gates that improved production reliability.\n\n\n\nLogistics\n\n - Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.\n\n - Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.\n\n - Visa sponsorship: We sponsor visas on a case-by-case basis.\n\n - Benefits: We offers generous health, dental, and vision benefits as well as 401(k) company match.","salary_min":200000,"salary_max":400000,"location":"San Francisco, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["llm","mlops","gpu","cloud","distributed-systems","devops","research"],"apply_url":"https://jobs.ashbyhq.com/inferact/ad992ead-2a9a-4694-8fca-0504354548cd/application","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-20T18:25:22.552Z","expires_at":"2026-09-29T13:41:27.521095Z","created_at":"2026-08-25T18:30:22.617589Z","updated_at":"2026-08-30T13:41:27.654002Z","company_name":"Inferact","company_slug":"inferact","company_logo_url":"https://www.google.com/s2/favicons?domain=inferact.ai\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/e83f9b8b-565f-4a94-b2f4-03b79a6f7843"},{"id":"fcda730f-0dab-4f16-a21e-6906a54e407c","company_id":"3029e985-56bf-4ac2-9ae1-df4cdd53b12f","title":"AI DevOps Engineer","slug":"ai-devops-engineer-40d84475","description":"About Zscaler \n Zscaler accelerates digital transformation to ensure our customers can be more agile, efficient, resilient, and secure. As an AI-forward enterprise , we are constantly pushing the envelope, leveraging the world’s largest security data lake to power our cloud-native Zero Trust Exchange platform. This innovation protects our customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location.\n Here, impact in your role matters more than title and trust is built on results. We say, impact over activity. We seek innovators who actively use AI to amplify their impact and who thrive in an environment where we leverage intelligent systems to stay ahead of evolving threats. We believe in transparency and value constructive, honest debate —we’re focused on getting to the best ideas, faster. We build high-performing teams that can make an impact quickly and with high quality. To do this, we are building a culture of execution centered on customer obsession , collaboration, ownership, and accountability.\n We value high-impact, high-accountability with a sense of urgency where you’re enabled to do your best work and embrace your potential. If you’re driven by purpose, thrive on solving complex challenges, and want to be part of the team that’s helping to secure the AI age, we invite you to bring your talents to Zscaler and help shape the future of cybersecurity.\n Role  \n We are looking for an AI DevOps Engineer to join our team. This is a Remote within the United States (with a hybrid preference for San Jose, CA) role, reporting to the Manager, IT Cloud Operations in the Cloud Platform Engineering department. Our team builds and operates the internal cloud platform that powers Zscaler's product and corporate infrastructure across AWS, GCP, and Azure. In this role, you will design and ship automation that provisions cloud environments, enforces security baselines, and integrates AI-assisted tooling and agentic workflows to accelerate delivery.\n What you’ll do (Role Expectations) \n \n Build and extend Day 1 automation, including infrastructure provisioning pipelines, account vending, golden repo templates, and CI/CD components\n Build and extend Day 2 automation for lifecycle management, upgrade pipelines, dependency scanning, drift detection, and automated remediation workflows\n Write and maintain Terraform modules, GitLab CI/CD components, and Python automation to establish the platform's paved road\n Integrate AI and agentic tooling into operational workflows to reduce manual toil and increase operational consistency\n Collaborate with Security, IAM, Network, and FinOps teams to translate cross-functional requirements into automated guardrails\n \n Who You Are (Success Profile) \n \n You thrive in ambiguity. You are comfortable building the path as you walk it, viewing dynamic environments as raw material to build something meaningful.\n You act like an owner. Your passion for the mission fuels your bias for action, seamlessly navigating between high-level strategy and hands-on execution.\n You are a problem-solver. You seek out challenges because you are energized by finding solutions, knowing that solving hard problems delivers maximum impact.\n You are a high-trust collaborator. You embrace a challenge culture by giving and receiving ongoing feedback with clarity, respect, and candor.\n You are a learner. You bring a true growth mindset and actively seek feedback to continuously develop yourself and support your team.\n \n What We’re Looking for (Minimum Qualifications) \n \n Demonstrated curiosity and active exploration of AI tools, with a proven history of integrating new technologies to enhance daily workflows and augment problem-solving\n 5+ years of experience building and operating multi-cloud infrastructure at scale across major cloud platforms\n Hands-on expertise with Terraform, including module design, state management, and CI-driven workflows\n Proficiency in scripting and automation using Python, Bash, or equivalent languages\n Working knowledge of CI/CD pipeline design and Kubernetes operations\n Understanding of identity federation, secrets management, and least-privilege security patterns\n \n What Will Make You Stand Out (Preferred Qualifications) \n \n Experience with multi-account governance tooling such as AWS Control Tower, Organizations, SCPs, or RCPs\n Experience with GitOps patterns and automated infrastructure lifecycle tooling such as Renovate, Dependabot, or ArgoCD\n Experience with MLOps frameworks: kubeflow, ML flow, and model services like Bedrock, Sagemaker, Vertex AI, Gemini Enterprise\n \n #LI-Remote #LI-YC2\n Zscaler’s salary ranges are benchmarked and are determined by role and level. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations and could be higher or lower based on a multitude of factors, including job-related skills, experience, and ","salary_min":140000,"salary_max":175000,"location":"Remote (US)","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"senior","tags":["fine-tuning","security","data-pipeline","mlops","agents","cloud","devops"],"apply_url":"https://job-boards.greenhouse.io/zscaler/jobs/5208829007","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-20T15:00:13Z","expires_at":"2026-09-29T13:39:53.846816Z","created_at":"2026-08-25T18:29:22.952569Z","updated_at":"2026-08-30T13:39:53.98003Z","company_name":"Zscaler","company_slug":"zscaler","company_logo_url":"https://www.google.com/s2/favicons?domain=zscaler.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/fcda730f-0dab-4f16-a21e-6906a54e407c"},{"id":"e2ee6535-e8ad-47aa-80bf-88a13e3a97fc","company_id":"a0000000-0000-0000-0000-000000000001","title":"Staff+ Site Reliability Engineer, Safeguards ML Infra","slug":"staff-site-reliability-engineer-safeguards-ml-infra-ebc3e8d7","description":"About Anthropic \n Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.\n About the role: \n The Safeguards ML Infra team designs, builds, and operates the production infrastructure that powers Claude's safety systems. We own the critical backend services that ensure safety on the token generation path, and we own the operational work of getting those systems safely into production: standing up safeguards for every new model launch, and deploying new safety classifiers as they ship. Every frontier model release runs through this team – we configure, verify, and roll out safeguards across every platform Claude runs on (1P, AWS Bedrock, GCP Vertex, etc.), and we lead incident response when issues arise.\n This role sits at the center of that operational work. You'll ensure safeguards are properly configured and deployed for model launches and own the off-cycle deployment of new safety classifiers — canarying changes, verifying that the right safeguards are provably live on the right models, and holding rollback authority when something looks wrong. Every launch should also shrink the checklist, and the manual verifications should evolve into a system that runs itself. You'll turn launch runbooks into tooling, hand-built checks into continuous validation, and one-off deploys into a repeatable pipeline.\n We're looking for engineers with deep experience in production change management at scale — people who have owned deploy pipelines, config management systems, rollout safety, or launch readiness for systems under real production pressure. Familiarity with ML research or transformer architectures is not required — you will learn that on the job. What we prioritize is production judgment: a track record of shipping changes to critical systems safely, and of automating yourself out of the work you did last quarter.\n What you'll do: \n \n Launch captain model releases: stand up, configure, and verify safeguards for every new model, and serve as the safeguards point of contact in the launch room during release windows.\n Own the off-cycle deployment of new safety classifiers as they ship from research — canarying rollouts, running post-deploy validations, and investigating discrepancies when something looks wrong.\n Verify that the right safeguards are provably live on the right models across every deployment platform (1P, AWS Bedrock, GCP Vertex, etc.), and detect and eliminate configuration drift between them.\n Automate yourself out of last quarter's work: turn launch runbooks into tooling, hand-built checks into continuous validation, and one-off deploys into a repeatable pipeline.\n \n Plan to use Claude aggressively to do this! And be a trailblazer that paves the path for safe agentic operations of safety-critical systems.\n \n Build and maintain a safeguards registry with full provenance — what is running in production, on which model, on which platform, and when and by whom it was deployed.\n Participate in on-call and operational-duty rotations covering service incidents, model provisioning, and time-sensitive research and safety launches.\n \n You may be a good fit if you: \n \n Have owned production change management at scale — deploy pipelines, config management systems, canary analysis — and have strong opinions about what \"verified\" means.\n Have run high-stakes releases: served as a launch captain, incident commander, or release owner for systems where a bad deploy has real consequences, and are energized rather than drained by being in the critical path.\n Have meaningful on-call experience for production systems, including incident response and postmortem-driven improvements — and a track record of turning (and fixing!) postmortem action items into process and tooling changes.\n Have a desire to close the gap where nobody has yet raised their hand, even if it requires manually hand-holding processes until automation and tooling can be built.\n Have hands-on experience deploying and operating on cloud platforms (AWS, GCP) at scale.\n Are proficient in Python; experience with Rust is a plus but not required.\n \n Strong candidates may also have: \n \n 8+ years of industry software engineering or site reliability engineering experience.\n A demonstrated history of reducing operational toil through automation, including transitioning teams from manual deployment processes to self-serve pipelines.\n Experience running launch or production-readiness review processes across multiple teams.\n Familiarity with LLM inference systems and the operational characteristics of transformer-based models.\n The annual compensation range for this role is listed below. \n For sales roles, the range provided is the role’s On Target Earnings (\"OTE\") range, meaning th","salary_min":405000,"salary_max":485000,"location":"San Francisco, CA","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["llm","agents","alignment","cloud","rust","infrastructure","devops"],"apply_url":"https://job-boards.greenhouse.io/anthropic/jobs/5230394008","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-11T00:28:36Z","expires_at":"2026-09-29T13:30:36.924785Z","created_at":"2026-08-25T18:26:19.565568Z","updated_at":"2026-08-30T13:30:37.072954Z","company_name":"Anthropic","company_slug":"anthropic","company_logo_url":"https://www.google.com/s2/favicons?domain=anthropic.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/e2ee6535-e8ad-47aa-80bf-88a13e3a97fc"},{"id":"017828a4-b1c3-40ba-80bf-94bd8d9cacc2","company_id":"f8de0913-0ef7-4e72-a9cf-81f8513ec624","title":"Software Engineer, DevOps","slug":"software-engineer-devops-e13f556d","description":"FieldAI’s Irvine team is where embodied AI meets real robots, real sensors, and real field deployments. Based in the heart of Southern California’s robotics ecosystem, we build risk-aware, reliable, field-ready AI systems that solve the hardest problems in robotics and unlock the full potential of embodied intelligence. If you want your work to ship, get tested on hardware, and improve through real deployments, Irvine is the place. We go beyond typical data-driven approaches or pure transformer-only architectures, combining rigorous engineering with learning systems proven in globally deployed solutions that deliver results today and get better every time our robots run in the field.\n","salary_min":115000,"salary_max":170000,"location":"Irvine, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"mid","tags":["robotics","infrastructure","devops","platform"],"apply_url":"https://jobs.lever.co/field-ai/b6a02c7e-8726-47b7-b0b0-899644613c47/apply","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-01T00:42:16.075Z","expires_at":"2026-09-29T13:46:13.324635Z","created_at":"2026-08-25T18:32:29.478451Z","updated_at":"2026-08-30T13:46:13.457753Z","company_name":"Field AI","company_slug":"field-ai","company_logo_url":"https://www.google.com/s2/favicons?domain=field.ai\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/017828a4-b1c3-40ba-80bf-94bd8d9cacc2"},{"id":"6109f5c9-3808-4bdf-a225-316edd299a6b","company_id":"6734f15a-40ed-4186-ae4a-d774c655ae58","title":"Staff/Principal DevOps Engineer, AI Inference","slug":"staffprincipal-devops-engineer-ai-inference-09cb28c0","description":"Your Impact at LILA \n The Staff/Principal DevOps Engineer - AI Inference will drive the design, implementation, and optimization of infrastructure purpose-built for serving machine learning models at scale. This role bridges platform engineering, site reliability, and ML infrastructure, building the systems that power low-latency, high-throughput inference across GPU clusters and cloud accelerators. You will collaborate with ML engineers, research scientists, and software engineers to build inference platforms that serve models reliably to production users while maximizing compute efficiency.\n What You'll Be Building \n \n GPU/accelerator infrastructure on Kubernetes: scheduling, resource isolation, multi-tenant GPU sharing, device plugins, and topology-aware placement for inference workloads\n Model serving platforms using frameworks such as vLLM, Triton Inference Server, TGI, or custom serving stacks with optimized batching, caching, and request routing\n Intelligent request routing and load balancing across heterogeneous accelerator fleets (NVIDIA GPUs, AWS Inferentia/Trainium) to maximize utilization and minimize latency\n Autoscaling systems that dynamically match inference compute supply with demand across production, research, and experimental workloads\n Production-grade deployment pipelines for ML models: canary rollouts, A/B testing, model versioning, and safe rollback across multi-region deployments\n Infrastructure-as-code with Terraform and Helm for GPU-accelerated EKS clusters, including node pools, spot/on-demand strategies, and accelerator-specific networking\n Observability and performance optimization: GPU utilization monitoring, inference latency profiling, token throughput dashboards, and SLO/SLI tracking for model endpoints\n CI/CD pipelines for model artifacts: container image builds with CUDA/driver dependencies, model registry integration, and automated inference benchmarking in CI\n AWS cloud infrastructure for ML: EKS with GPU node groups, EC2 accelerated instances (P4/P5, Inf2, Trn1), S3 model storage, EFA/high-bandwidth networking, and IAM least privilege\n Cost optimization and capacity planning: right-sizing accelerator instances, spot instance strategies for inference, and fleet-wide efficiency reporting\n \n What You'll Need to Succeed \n \n Expertise in DevOps, SRE, or Platform Engineering with significant experience operating GPU/accelerator infrastructure at scale\n Deep experience with Kubernetes for ML workloads: GPU scheduling, resource quotas, node affinity, and accelerator device management\n Strong proficiency deploying to AWS using infrastructure-as-code (Terraform, Helm) with hands-on experience managing GPU-based compute (EKS, EC2 P-series/Inf/Trn instances)\n Experience with model serving infrastructure: inference servers, request batching, KV-cache optimization, or LLM serving frameworks\n Strong understanding of networking for distributed inference: high-bandwidth interconnects, NCCL, VPC/PrivateLink, and load balancing at L4/L7\n Strong proficiency in Python for automation, tooling, and integration with ML frameworks\n \n Bonus Points For \n \n Experience with LLM inference optimization: continuous batching, speculative decoding, quantization (GPTQ, AWQ, FP8), tensor parallelism, and pipeline parallelism\n Hands-on experience with multiple accelerator families (NVIDIA A100/H100, AWS Inferentia2, Trainium, AMD MI300X) and maintaining hardware-agnostic serving infrastructure\n Multi-region deployment experience with geographic routing and failover for latency-sensitive inference endpoints\n Proficiency in Rust or Go for performance-critical infrastructure components\n SRE practices for ML systems: chaos engineering on GPU workloads, incident management, capacity modeling for bursty inference traffic\n Experience with model registries, artifact versioning, and ML supply chain security\n Observability platform expertise: building custom metrics for token-level throughput, time-to-first-token, and per-request GPU memory profiling\n Prior startup/high-growth experience balancing velocity with reliability in rapidly scaling AI systems\n \n \n Compensation \n We offer competitive base compensation with bonus potential and generous early-stage equity. Your final offer will reflect your background, expertise, and expected impact.\n U.S. Benefits. Full-time U.S. employees receive a comprehensive benefits program including medical, dental, and vision coverage; employer-paid life and disability insurance; flexible time off with generous company wide holidays; paid parental leave; an educational assistance program; commuter benefits, including bike share memberships for office based employees; and a company subsidized lunch program.\n International Benefits. Full-time employees outside the U.S. receive a comprehensive benefits program tailored to their region. USD salary ranges apply only to U.S.-based positions; international salaries are set to local market.\n Expected Base Salary Range\n $192,000 — $272,000","salary_min":192000,"salary_max":272000,"location":"Boston, MA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"principal","tags":["llm","mlops","gpu","cloud","devops","inference"],"apply_url":"https://job-boards.greenhouse.io/lilasciences/jobs/4248032009","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-07-28T15:38:59Z","expires_at":"2026-09-29T13:48:28.77755Z","created_at":"2026-07-29T14:18:39.098522Z","updated_at":"2026-08-30T13:48:28.909371Z","company_name":"Lila Sciences","company_slug":"lila-sciences","company_logo_url":"https://www.google.com/s2/favicons?domain=lila.ai\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/6109f5c9-3808-4bdf-a225-316edd299a6b"},{"id":"c5a82787-a19d-4022-80ca-e5dead08f53b","company_id":"63bced38-3605-4e57-99f3-e213b2d40bf3","title":"Site Reliability Engineer ll","slug":"site-reliability-engineer-ll-fc478b73","description":"Opportunity Overview:  \n This is a remote-first role that may require travel to Boston, MA for new hire onboarding and occasional in-person team meetings and company events.\n We are seeking an operational-focused Site Reliability Engineer (SRE) to maximize the availability, performance, and resilience of our production healthcare systems. In this role, you will bridge the gap between AWS cloud infrastructure, MERN stack applications, and large-scale data workflows. You will spend roughly 60% of your time on live incident remediation, data pipeline operations, and Node.js/Python infrastructure tuning, and 40% on engineering automated solutions to eliminate operational toil.\n What you’ll do: \n \n Production Operations: Maintain the continuous uptime, scalability, and security of our AWS-hosted MERN applications and backend data architectures.\n Serverless Execution: Manage, optimize, and troubleshoot event-driven architectures running on AWS Lambda, focusing on cold-start mitigation, memory allocation, and execution timeouts.\n Data Pipeline Execution: Monitor scheduled PySpark data workflows, execute standard operating procedures (SOPs) for large-scale data ingestion, and rapidly triage, rerun, or patch failed data processing jobs.\n Incident Management: Participate in a collaborative on-call rotation to rapidly triage, debug, and mitigate live application outages and data flow bottlenecks.\n Healthcare Compliance: Maintain strict HIPAA, SOC2, and HITRUST compliance profiles across all runtime environments, storage systems, and data pipelines handling Protected Health Information (PHI).\n Toil Elimination: Engineer automated workflows to eliminate repetitive tasks like manual data seeding, infrastructure provisioning, and routine PySpark pipeline recovery steps.\n Observability Engineering: Build specialized dashboards and alerts to monitor Node.js event loops, PySpark job execution stages, driver/worker memory leaks, and data pipeline throughput anomalies.\n Post-Mortem Culture: Lead blameless post-mortems for operational and data processing failures, translating system crashes into permanent structural fixes.\n \n What you’ll need: \n \n SaaS Platform Experience: Minimum of 3+ years of hands-on experience operating multi-tenant, cloud-hosted, or cloud-native SaaS platforms at scale.\n AWS Cloud Engineering: Deep expertise operating AWS core services, specifically AWS Lambda, Amazon ECS/EKS, Amazon EMR or AWS Glue (for Spark), EC2, VPC networking, IAM permissions, and CloudWatch.\n Automation \u0026 Data Languages: Professional competency in writing, debugging, and maintaining automation scripts and data tools using Python (including PySpark APIs) and Node.js.\n Data Operations: Experience managing and troubleshooting distributed data orchestration pipelines, ETL tools, message queues (e.g., AWS SQS/SNS, RabbitMQ), or stream processing frameworks.\n MERN Stack Operations: Deep understanding of the operational lifecycle of JavaScript/TypeScript applications, including memory management, asynchronous runtimes, and Node.js clustering.\n Database Administration: Practical experience managing, sharding, indexing, and optimizing production-grade MySQL DB \u0026 Athena (RDS or self-hosted).\n Infrastructure as Code: Proven ability to deploy and maintain immutable infrastructure utilizing Terraform or OpenTofu.\n Healthcare Experience: Minimum 1 year working within HIPAA-regulated environments. Direct experience securing data-at-rest and data-in-transit containing sensitive patient records is preferred.\n Education \u0026 Experience: Minimum of 4 years of software/systems experience, with at least 1-2 years focused on live cloud operations and distributed data workflow management is preferred.\n Crisis Management: Calm under pressure with a methodical approach to identifying and isolating PySpark driver OOM (Out of Memory) errors or data corruption during high-stress outages. Attention to detail and effective communications skills will be critical in working with clients and internal stakeholders is preferred.\n \n  \n Pay \u0026 Perks: \n 💻 Fully remote opportunity with about 5% travel\n 🩺 Medical, dental, vision, life, disability insurance, and Employee Assistance Program \n 📈 401K retirement plan with company match; flexible spending and health savings account \n 🏝️ Flex Time Off + company holidays\n 👶 Up to 14 weeks of paid parental leave \n 🐶 Pet insurance  \n  \n The salary range for this position is $100,000 to $110,000 annually; as part of a total benefits package which includes health insurance, 401k and bonus. In accordance with state applicable laws, Cohere is required to provide a reasonable estimate of the compensation range for this role. Individual pay decisions are ultimately based on a number of factors, including but not limited to qualifications for the role, experience level, skillset, and internal alignment. \n This role is not eligible for hire in: CA\n  \n Interview Process*: \n \n Connect with Talent Acquisition ","salary_min":100000,"salary_max":110000,"location":"United States","workplace":"remote","remote_scope":"restricted","job_type":"full-time","experience_level":"mid","tags":["healthcare","data-pipeline","payments","cloud","agents","devops"],"apply_url":"https://job-boards.greenhouse.io/coherehealth/jobs/7807972003","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-07-17T14:50:55Z","expires_at":"2026-09-29T13:36:34.285011Z","created_at":"2026-07-18T14:06:45.54803Z","updated_at":"2026-08-30T13:36:34.424934Z","company_name":"Cohere Health","company_slug":"cohere-health","company_logo_url":"https://www.google.com/s2/favicons?domain=coherehealth.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/c5a82787-a19d-4022-80ca-e5dead08f53b"},{"id":"cea172f2-7ff5-4ae5-9400-c763a96f22dc","company_id":"714f360f-a244-487d-b3f0-0c43518a9e66","title":"Sr. Data Scientist, Infrastructure","slug":"sr-data-scientist-infrastructure-269a1d04","description":"About Pinterest: \n Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product.\n Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the  flexibility to do your best work. Creating a career you love? It’s Possible.\n At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI.\n Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here .\n Pinterest brings millions of people the inspiration to create a life they love. Behind that experience is a complex infrastructure ecosystem that powers reliability, performance, measurement, and efficiency across the platform. As Pinterest grows, it’s increasingly important that we understand these systems clearly so we can make smarter decisions for both Pinners and the business.\n  \n We’re looking for a Data Scientist to join our Infrastructure Data Science team. In this role, you’ll partner with engineering and cross-functional teams to make Pinterest’s infrastructure more measurable, intelligible, and actionable. Depending on the area, your work may span app performance, shopping infrastructure, metrics quality, infrastructure governance, or site reliability. You’ll help build the data foundations, measurement systems, and analytical frameworks that enable Pinterest to optimize core technical systems and make better product and infrastructure decisions.\n  \n What you’ll do: \n In this role, you will partner closely with engineering and cross-functional teams to improve how Pinterest measures, understands, and optimizes its infrastructure:\n \n Partner with engineering teams to define, measure, and improve the health, quality, and efficiency of Pinterest’s infrastructure systems.\n Build and refine metrics, dashboards, and analytical frameworks that make complex technical systems more understandable and actionable.\n Strengthen data foundations by improving metric definitions, auditing data quality, and contributing to pipeline and measurement improvements where needed.\n Design and analyze experiments, investigations, and deep dives to quantify the impact of infrastructure changes on user experience, reliability, and business outcomes.\n Translate ambiguous technical problems into clear analyses and actionable recommendations for engineering and platform partners.\n Support high-priority investigations and decision-making related to infrastructure performance, reliability, cost, and measurement quality.\n Identify opportunities to improve how Pinterest measures and optimizes infrastructure across a range of domains, such as performance, shopping infrastructure, governance, metrics quality, and site reliability.\n \n  \n What we’re looking for: \n \n 4+ years of combined post-graduate academic and industry experience applying scientific methods to solve real-world problems with large-scale data.\n Bachelor’s/Master’s degree in a relevant field such as Computer Science, or equivalent experience.”\n Strong SQL and analytical programming skills, with experience working through messy, imperfect data and building reliable metrics and datasets.\n Experience partnering on or contributing to production-ready data pipelines, measurement systems, or foundational data work that improves data quality and usability.\n Solid foundation in experimentation and measurement, with the ability to design analyses, interpret results rigorously, and partner effectively with engineers and other cross-functional stakeholders.\n Demonstrated ability to translate ambiguous problems into clear analytical workstreams and actionable recommendations.\n Strong cross-functional communication skills, with the ability to explain technical findings clearly to engineering, product, and platform stakeholders.\n Ability to operate independently, prioritize across both longer-term projects and fast-turn inbound requests, and drive work forward in a dynamic environment.\n Curiosity and a builder mindset, with excitement for improving messy systems and creating more scalable, trustworthy measurement foundations.\n \n  \n In-Office Requirement Statement: \n \n We recognize that the ideal environment for work is situational and may differ across departments. What this looks like day-to-day can vary based on the needs","salary_min":139764,"salary_max":287749,"location":"San Francisco, CA","workplace":"remote","remote_scope":"unknown","job_type":"full-time","experience_level":"senior","tags":["data-pipeline","infrastructure","data-science","devops"],"apply_url":"https://www.pinterestcareers.com/jobs/?gh_jid=8024966","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-07-07T20:19:01Z","expires_at":"2026-09-29T13:38:53.944326Z","created_at":"2026-07-09T14:08:38.685575Z","updated_at":"2026-08-30T13:38:54.079802Z","company_name":"Pinterest","company_slug":"pinterest","company_logo_url":"https://www.google.com/s2/favicons?domain=www.pinterest.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/cea172f2-7ff5-4ae5-9400-c763a96f22dc"},{"id":"f559fa07-ecbc-46d6-a526-8cc8c9dd7d69","company_id":"e3915539-5a8f-4461-9f26-06366a918674","title":"Senior Site Reliability Engineer","slug":"senior-site-reliability-engineer-40b58c81","description":"Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the expertise, technology, and business model of the 21st century’s most innovative companies to the defense industry, Anduril is changing how military systems are designed, built and sold. Anduril’s family of systems is powered by Lattice OS, an AI-powered operating system that turns thousands of data streams into a realtime, 3D command and control center. As the world enters an era of strategic competition, Anduril is committed to bringing cutting-edge autonomy, AI, computer vision, sensor fusion, and networking technology to the military in months, not years.\n ABOUT THE TEAM \n We are seeking a highly skilled and mission-driven Site Reliability Engineer (SRE) to join our Mission Autonomy team. In this critical role, you will be responsible for ensuring the reliability, scalability, performance, and operational excellence of our cutting-edge autonomous systems. This isn't just about keeping servers up; it's about building and maintaining the resilient backbone for systems where failure is not an option, and mission success directly impacts national security. You will embed with our autonomy software development teams, acting as a bridge between development and operations. Your work will directly enable our Mission Autonomy software and control systems to operate flawlessly, whether in cloud-based simulation environments, hardware-in-the-loop devices or air-gapped environments\n What You’ll Do \n \n Manage and expand specialized on-site infrastructure: Administer and grow on-premises developer servers, Hardware-in-the-Loop (HITL) systems, and other compute resources.\n Design, implement, and maintain highly available, fault-tolerant, and resilient autonomous systems\n Identify and eliminate performance bottlenecks in software and infrastructure, ensuring low-latency, high-throughput, and real-time responsiveness for mission-critical operations.\n Develop and implement comprehensive monitoring, logging, tracing, and alerting solutions to provide deep insights into system health and behavior at scale.\n Automate away manual operational tasks, from provisioning and deployment to testing and recovery.\n Develop and implement strategies for scaling our services and infrastructure to meet evolving mission demands, including distributed systems and edge deployments.\n Work closely with security teams to integrate best practices into our operational processes and infrastructure, ensuring the integrity and confidentiality of our autonomous systems.\n Create clear, concise, and comprehensive documentation, runbooks, and playbooks for operational procedures.\n Integrate open-source, commercial, and Anduril-internal tooling to create effective solutions for software delivery.\n Collaborate with Anduril's Developer Platform, Networking, and Security teams to support integration with broader Anduril systems.\n Work with a multi-disciplinary team on challenging problems in a fast-paced environment.\n \n Required Qualifications \n \n Bachelor of Science degree in Computer Science, Engineering or a related field, or equivalent work experience.\n 5+ years of experience in Site Reliability Engineering, DevOps, or a similar role focused on security for mission-critical applications\n Strong proficiency in at least one modern programming language (Python, Go ) .\n Experience with automation tools (Ansible, Puppet or Terraform)\n Deep expertise with Linux operating systems and strong command-line skills.\n Knowledge of secure coding practices and experience implementing security controls in cloud and on-premise environments.\n Solid understanding of networking fundamentals (TCP/IP, DNS, HTTP, load balancing) and their impact on system reliability.\n Proficiency with containerization technologies (Docker) and orchestration platforms (Kubernetes).\n Strong analytical, problem-solving, and debugging skills, with a methodical approach to complex system issues.\n Excellent communication skills and the ability to work effectively in cross-functional teams.\n Must be a U.S. Person due to required access to U.S. export controlled information or facilities.\n Active U.S. Security Clearance.\n \n Preferred Qualifications \n \n Experience with edge computing, mesh networks, or highly distributed autonomous systems.\n Experience with embedded Linux systems development and associated tools.\n Experience troubleshooting and analyzing remotely deployed software systems.\n Familiarity with monitoring and logging tools (like auditd, journald, selinux, Splunk).\n Prior experience in defense, aerospace, robotics, or other mission-critical domains\n Extensive experience with cloud platforms (AWS, Azure, or GCP) and understanding of their core services.\n US Salary Range\n $166,000 — $220,000 USD \n The salary range for this role is an estimate based on a wide range of compensation factors, inclusive of base salary only. Actual ","salary_min":166000,"salary_max":220000,"location":"Costa Mesa, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"senior","tags":["payments","robotics","distributed-systems","computer-vision","cloud","devops"],"apply_url":"https://boards.greenhouse.io/andurilindustries/jobs/5124136007?gh_jid=5124136007","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-05-11T19:42:36Z","expires_at":"2026-09-29T13:37:24.422738Z","created_at":"2026-05-12T14:08:08.32887Z","updated_at":"2026-08-30T13:37:24.556192Z","company_name":"Anduril","company_slug":"anduril","company_logo_url":"https://www.google.com/s2/favicons?domain=anduril.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/f559fa07-ecbc-46d6-a526-8cc8c9dd7d69"},{"id":"8cf7148e-6bb7-4dd6-a94f-25991b9cdc54","company_id":"73b8a49c-c986-41f7-9066-23b04b8632bb","title":"Software Engineer, DevOps","slug":"software-engineer-devops-b3898fcf","description":"ABOUT EMA\n\nEma is building the world’s leading Agentic AI platform to transform enterprise productivity. We enable organizations to delegate repetitive tasks to Ema, the Universal AI Employee, delivering 10x gains in workforce efficiency, across functions. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs.\n\nWe are backed by industry leading investors including Accel, Naspers/Prosus, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production — we ship real systems that run real business processes at scale.\n\n\nWHO YOU ARE\n\nWe are seeking an experienced DevOps Engineer to join our growing team and play a pivotal role in designing and building our platform and infrastructure as we continue to scale our product and user base. As a part of our team, you will be working in a dynamic, fast-paced environment to ensure the reliability, scalability, and performance of our systems, while focusing on service architecture and deployment, query optimization, distributed systems, data and machine learning infrastructure, and security and authentication. Most importantly, you are excited to be part of a mission-oriented, fast-paced, high-growth startup that can create a lasting impact.\n\n\n\n\nYOU WILL:\n\n 1. Partner with product teams to architect, design, and build the foundational infrastructure for our products.\n\n 2. Design, develop, and deploy highly available and scalable Multi-tenant SaaS solutions on any one of the public cloud networks like AWS, Azure and GCP. Leverage technologies such as Kubernetes, Helm, Terraform, and Istio to achieve infrastructure resilience.\n\n 3. Drive the automation of infrastructure tasks, from provisioning to configuration management and deployment, utilizing tools like Terraform, Ansible, and Kubernetes.\n\n 4. Collaborate closely with the software development team to refine CI/CD pipelines, e.g., using GitHub Actions and Cloud Build tools, enhance service interfaces, and improve the overall developer experience.\n\n 5. Architect and implement advanced observability solutions using tools like Prometheus and Grafana. Ensure real-time alerting and error tracking with Sentry and Pagerduty to maintain system health and performance.\n\n 6. Deploy comprehensive testing frameworks, including tools like Selenium for end-to-end testing. Ensure robust integration and system testing to maintain software quality.\n\n 7. Performance Analysis: Regularly monitor system health, analyze performance metrics, and recommend enhancements. This includes optimizing database queries and ensuring peak database performance.\n\n\n\n\nNICE TO HAVE\n\n 1. ML/OPs experience\n\n 2. Experience with Postgres query optimization and related performance improvement techniques.\n\n 3. Experience with event-driven data and machine learning infrastructure, including streaming pipelines, database systems, model training\n\n 4. Experience with air-gapped cloud environments or private clouds\n\n 5. Experience administering complex deployments on Azure, especially AKS\n    \n    \n\n\nQUALIFICATIONS:\n\n - Bachelor's or Master's degree in Computer Science or related field.\n\n - 3+ years of experience in Infrastructure engineering, or a similar role, \n\n - Excellent problem-solving skills and the ability to work under pressure in a fast-paced environment.\n\n - Ability to work independently and as part of a team\n\n - Experience working with global teams\n\n\n\nFor California based candidates:\nThe standard base salary for this position is $135,000-$225,000 annually. \n\nCompensation offered will be determined by factors such as location, level, job-related knowledge, skills, and experience. Certain roles may be eligible for variable compensation, equity, and benefits.\n\nEma Unlimited is an equal opportunity employer and is committed to providing equal employment opportunities to all employees and applicants for employment without regard to race, color, religion, sex, national origin, age, disability, sexual orientation, gender identity, or genetics.","salary_min":135000,"salary_max":225000,"location":"San Francisco, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"mid","tags":["cloud","distributed-systems","agents","devops"],"apply_url":"https://jobs.ashbyhq.com/ema/6394f5e3-6952-4f0e-9e6c-4e9556549f3a/application","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-04-30T16:14:54.083Z","expires_at":"2026-09-29T13:44:15.928218Z","created_at":"2026-05-06T14:19:09.551668Z","updated_at":"2026-08-30T13:44:16.061945Z","company_name":"Ema","company_slug":"ema","company_logo_url":"https://www.google.com/s2/favicons?domain=ema.co\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/8cf7148e-6bb7-4dd6-a94f-25991b9cdc54"},{"id":"68470449-f449-4862-9c0d-f15fee9aad3e","company_id":"a0000000-0000-0000-0000-000000000003","title":"DevOps Engineer, Infrastructure \u0026 Security","slug":"devops-engineer-infrastructure-security-a4b4c992","description":"As Scale's product portfolio and customer base expand, we are seeking skilled DevOps Engineers, Public Sector to be at the forefront of building out and enhancing our CI/CD pipelines. You will play a crucial role in streamlining our Software Development Life Cycle (SDLC) through collaborative efforts, moving us from a state of manual, disparate deployments to a more unified and automated system.\n These engineers will gain a deep understanding of our core products' architecture and composition, enabling them to effectively deploy and manage these systems when needed. A critical aspect of this role will be seamlessly integrating various machine learning (ML) tasks and updates into our SDLC, transforming currently separate ML components into a cohesive and automated workflow. While direct ML expertise is not required, a desire to learn and integrate ML components into the lifecycle is essential.\n Must have: \n \n At least an active TS/SCI clearance and the ability \u0026 willingness to up level to CI Poly. This is a requirement and candidates will not be considered who do not hold at least a TS/SCI clearance. \n \n You will: \n \n Design, develop, and maintain robust CI/CD pipelines to automate the deployment of our lowside and highside products.\n Collaborate closely with product and engineering teams to enhance existing application code for improved compatibility and streamlined integration within automated pipelines.\n Contribute to the overall architecture and design of our deployment systems, bringing new ideas to life for increased efficiency and reliability.\n Troubleshoot and resolve complex deployment issues, ensuring minimal disruption to development cycles.\n Develop a deep understanding of our product and ML architectures to facilitate seamless integration and deployment.\n Document pipeline processes and configurations to ensure maintainability and knowledge transfer.\n Proactively incorporate security best practices into all stages of the CI/CD pipeline, building security into our development processes.\n Drive standardization and foster collaboration across different product teams to achieve a unified and efficient SDLC.\n \n Ideally you'd have: \n \n 2-3 years of experience as a DevOps Engineer, DevSecOps Engineer, Software Engineer with a strong focus on CI/CD, or a similar role.\n Proven track record of building or significantly enhancing CI/CD pipelines.\n Experience configuring and adapting application code to integrate seamlessly with evolving CI/CD environments.\n Experience working fluently with standard containerization \u0026 deployment technologies like Kubernetes, Terraform, Docker, etc.\n Familiarity with cloud platforms (e.g., AWS, Azure, GCP).\n Strong proficiency in scripting and automation (e.g., Python, Bash, PowerShell).\n Familiarity with various CI/CD platforms (e.g., Jenkins, GitLab CI, GitHub Actions, Azure DevOps).\n Knowledge of software architecture, system design, and version control systems.\n Comfort with rapidly changing, fast-paced environments and a passion for finding automated solutions to complex problems.\n Basic understanding of security best practices in software development and an eagerness to integrate them.\n A hunger for learning new technologies, particularly in the realm of integrating ML into automated workflows.\n Strong problem-solving, analytical, collaboration, and communication skills.\n \n Nice to haves: \n \n Experience with containerization technologies (e.g., Docker, Kubernetes).\n Exposure to machine learning lifecycles or MLOps concepts.\n Prior experience in classified environments.\n Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position and may be inclusive of several career levels at Scale; it will be determined during the interview process based on work location and additional factors, including job-related skills, experience, qualifications, interview performance, and relevant education or training. Scale employees in eligible roles are also granted equity based compensation, subject to Board of Director approval. Your recruiter can share more about the specific salary range for your preferred location during the hiring process, and confirm whether the hired role will be eligible for equity grant. You'll also receive benefits including, but not limited to: comprehensive health, dental and vision coverage, retirement benefits, a learning and development stipend, and generous PTO. Additionally, this role may be eligible for additional benefits such as a commuter stipend. \n Please reference the job posting's subtitle for where this position will be located. For pay transparency purposes, the base salary range for this full-time position in the locations of San Francisco, New York, Seattle is:\n $198,400 — $311,000 USD \n The base salary range for this full-time position in the locations of Hawaii, Washington DC, Texas, Colorado","salary_min":148800,"salary_max":233000,"location":"Washington, DC","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"junior","tags":["cloud","mlops","fine-tuning","infrastructure","devops"],"apply_url":"https://job-boards.greenhouse.io/scaleai/jobs/4674863005","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-03-18T21:34:04Z","expires_at":"2026-09-29T13:31:36.502223Z","created_at":"2026-04-13T09:36:42.046044Z","updated_at":"2026-08-30T13:31:36.643935Z","company_name":"Scale AI","company_slug":"scale-ai","company_logo_url":"https://www.google.com/s2/favicons?domain=scale.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/68470449-f449-4862-9c0d-f15fee9aad3e"},{"id":"21dabcf4-0b7a-4145-b578-311c0eb3f3ed","company_id":"f8de0913-0ef7-4e72-a9cf-81f8513ec624","title":"Software Engineer, DevOps","slug":"software-engineer-devops-46ebf5d8","description":"FieldAI’s Irvine team is where embodied AI meets real robots, real sensors, and real field deployments. Based in the heart of Southern California’s robotics ecosystem, we build risk-aware, reliable, field-ready AI systems that solve the hardest problems in robotics and unlock the full potential of embodied intelligence. If you want your work to ship, get tested on hardware, and improve through real deployments, Irvine is the place. We go beyond typical data-driven approaches or pure transformer-only architectures, combining rigorous engineering with learning systems proven in globally deployed solutions that deliver results today and get better every time our robots run in the field.\n","salary_min":115000,"salary_max":170000,"location":"Irvine, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"mid","tags":["robotics","devops","platform","infrastructure"],"apply_url":"https://jobs.lever.co/field-ai/208e9ad5-cf16-4fe8-b991-cec7296ca46b/apply","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-03-09T21:20:45.301Z","expires_at":"2026-09-29T13:46:13.229857Z","created_at":"2026-04-16T19:55:12.723688Z","updated_at":"2026-08-30T13:46:13.364269Z","company_name":"Field AI","company_slug":"field-ai","company_logo_url":"https://www.google.com/s2/favicons?domain=field.ai\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/21dabcf4-0b7a-4145-b578-311c0eb3f3ed"},{"id":"2e2b96a1-0862-4862-acb8-adb6282b70f0","company_id":"7551b4ca-b2b0-493a-ab58-a15bd9c50393","title":"Senior Infrastructure Engineer/SRE ","slug":"senior-infrastructure-engineersre-33a53499","description":"Cresta unlocks the true potential of the customer experience, turning every conversation into a competitive advantage. Cresta’s unified AI platform combines conversational AI agents, real-time human agent augmentation, and comprehensive conversation intelligence to drive revenue and efficiency gains across every channel. The world’s leading companies, including United Airlines, Cox Communications, and Marriott, use Cresta to power world-class customer experiences every day. \n Born from the Stanford AI Lab, Cresta has raised more than $270 million from the world’s leading investors, including a16z, Greylock, and Sequoia. Cresta’s leadership includes some of the leading minds in AI today. Our CEO, Ping Wu , founded and led Google's Contact Center AI and Vertex AI platforms before joining Cresta to build the future of AI-driven customer experiences.\n Over the next few years, AI is going to redefine how people all over the world interact with businesses every day. Come build that future at Cresta.\n \n \n About the role: \n As a member of the infrastructure team you are responsible for designing, building, and advancing our core infrastructure that allows the engineering team to execute quickly, productively, and securely. You will join a collaborative but highly autonomous working environment in which each member has a defined role with clear expectations, as well as the freedom to pursue projects they find interesting.\n Responsibilities:\n \n Developer Toolchain . Partner with engineers to build dev tools that empower developer workflows and deployment infrastructure.\n Ensure reliability  of multi-cloud Kubernetes clusters and pipelines.\n Metrics, logging, analytics, and alerting  for performance and security across all endpoints and applications.\n Infrastructure-as-code  deployment tooling and supporting services on multiple cloud providers.\n Automate operations and engineering . Focus on automation so we can spend energy where it matters.\n Building machine learning infrastructure  that enables AI teams to train, test, and deploy on large-scale datasets.\n \n What we are looking for:\n \n \n 5+ years experience in DevOps, Site Reliability Engineering, Production Engineering, or equivalent field.\n Deep proficiency with coding languages such as Golang or Python.\n Deep familiarity with container-related security best practices.\n Production experience working with Kubernetes, and a deep understanding of the Kubernetes ecosystem, including popular open-source tooling such as cert-manager or external-dns.  Experience with GPU-enabled clusters is a bonus.\n Production experience with Kubernetes templating tools such as Helm or Kustomize.\n Production experience with IAC tools such as Terraform or CloudFormation.\n Production experience working with AWS and services such as IAM, S3, EC2, and EKS.\n Production experience with other cloud providers such as Google Cloud and Azure is a bonus.\n Production experience with database software such as PostgreSQL\n Experience with GitOps tooling such as Flux or Argo.\n Experience with CI/CD such as GitHub Actions.\n \n Perks \u0026 Benefits: \n We offer a comprehensive and people-first benefits package to support you at work and in life:\n \n Comprehensive medical, dental, and vision coverage with plans to fit you and your family\n Flexible PTO to take the time you need, when you need it\n Paid parental leave for all new parents welcoming a new child\n Retirement savings plan to help you plan for the future\n Remote work setup budget to help you create a productive home office\n Monthly wellness and communication stipend to keep you connected and balanced\n In-office meal program and commuter benefits provided for onsite employees\n \n Compensation at Cresta:  \n Cresta’s approach to compensation is simple: recognize impact, reward excellence, and invest in our people. We offer competitive, location-based pay that reflects the market and what each individual brings to the table.\n The posted base salary range represents what we expect to pay for this role in a given location. Final offers are shaped by factors like experience, skills, education, and geography. In addition to base pay, total compensation includes equity and a comprehensive benefits package for you and your family.\n OTE Range : $205,000–$270,000 + Offers Equity\n We have noticed a rise in recruiting impersonations across the industry, where scammers attempt to access candidates' personal and financial information through fake interviews and offers. All Cresta recruiting email communications will always come from the @cresta.ai domain. Any outreach claiming to be from Cresta via other sources should be ignored.  If you are uncertain whether you have been contacted by an official Cresta employee, reach out to  recruiting@cresta.ai","salary_min":205000,"salary_max":270000,"location":"United States","workplace":"remote","remote_scope":"restricted","job_type":"full-time","experience_level":"senior","tags":["agents","cloud","devops","infrastructure"],"apply_url":"https://job-boards.greenhouse.io/cresta/jobs/5137153008","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-03-01T23:53:42Z","expires_at":"2026-09-29T13:34:40.188372Z","created_at":"2026-04-13T09:39:51.526402Z","updated_at":"2026-08-30T13:34:40.324573Z","company_name":"Cresta","company_slug":"cresta","company_logo_url":"https://www.google.com/s2/favicons?domain=cresta.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/2e2b96a1-0862-4862-acb8-adb6282b70f0"},{"id":"6d2662e2-42a0-4e50-8dc6-8be51b005dfe","company_id":"63839083-85dd-4aa0-b128-254fc82866e5","title":"Senior / Staff Software Engineer (Observability / SRE)","slug":"senior-staff-software-engineer-observability-sre-cc261653","description":"Waabi, founded by AI visionary Raquel Urtasun, is the leader in Physical AI. With a world-class team, we're unlocking the next era of autonomous transportation with technology that's powering commercial autonomous trucks and robotaxis. Waabi is backed by and partners with world leaders in AI, automotive, logistics, and deep tech.\n\nWith offices in Toronto, San Francisco, Dallas, and Pittsburgh, Waabi is growing quickly and looking for diverse, innovative and collaborative candidates who want to impact the world in a positive way. To learn more visit: www.waabi.ai\n\n\nYou will..\n- Design and lead the architecture and development of Waabi’s monitoring and observability stack, used to monitor the health and performance of cloud and on-prem environments.\n- Develop and extend workloads and benchmarks (compute, storage, network, ML/AI) and integrate stress, chaos, and regression tests to validate hardware and platform choices.\n- Analyze and optimize end-to-end performance across hardware, firmware, Linux kernel, runtimes, and distributed services using advanced profiling tools (perf, eBPF, flamegraphs, tracing frameworks).\n- Build automation and observability tooling (Go/Python/Java, Kubernetes/Docker) for CI/CD-based performance regression detection, telemetry, alerting, and anomaly detection.\n- Work with client teams to support their applications’ observability requirements.\n- Influence system architecture and tooling decisions that improve how Waabi builds, monitors, and scales its infrastructure.\n- Drive execution and quality, writing design docs, setting milestones, mentoring ICs, and communicating insights and results to stakeholders and leadership.\n \nQualifications:\n- 5+ years software engineering or systems/performance engineering experience (BS in CS/EE or related), with demonstrated end-to-end ownership of complex projects.\n- Proficient in at least one of: Python, Rust, C/C++; strong CS fundamentals and system design skills.\n- Hands-on with Linux internals (CPU scheduling, memory, I/O, networking) and perf tooling (perf, eBPF, flamegraphs, tracing frameworks).\n- Experience with Kubernetes, microservices, and distributed systems; comfort building production services and pipelines.\n- Proven track record of clear communication, writing design docs, and leading cross-functional efforts.\n \nBonus: \n- Experience deploying and managing observability platforms (OpenTelemetry, Grafana OSS).\n- Performance tuning for databases/streaming/batch/ML platforms; GPU/xPU or Arm performance exposure.\n- Experience tuning stream processing, batch or ML platforms (e.g. Argo Workflows, PyTorch).\n- Familiarity with microservices debugging and distributed tracing (OpenTelemetry, Prometheus).\n","salary_min":148000,"salary_max":249000,"location":"Toronto, Canada","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["distributed-systems","microservices","pytorch","devops"],"apply_url":"https://jobs.lever.co/waabi/17347bcc-7c94-4817-b7dc-28acebba05e1/apply","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-02-12T03:54:29.737Z","expires_at":"2026-09-29T13:36:14.463665Z","created_at":"2026-04-13T09:41:54.073204Z","updated_at":"2026-08-30T13:36:14.596705Z","company_name":"Waabi","company_slug":"waabi","company_logo_url":"https://www.google.com/s2/favicons?domain=waabi.ai\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/6d2662e2-42a0-4e50-8dc6-8be51b005dfe"},{"id":"c9923cc2-371f-4bb8-b7cb-84b54c2f3619","company_id":"66e863fb-9aaf-40df-996c-eb439e6f857e","title":"Lead Site Reliability Engineer","slug":"lead-site-reliability-engineer-249dfb54","description":"About Glean: \n  \n Glean is the Work AI platform that helps everyone work smarter with AI. What began as the industry’s most advanced enterprise search has evolved into a full-scale Work AI ecosystem, powering intelligent Search, an AI Assistant, and scalable AI agents on one secure, open platform. With over 100 enterprise SaaS connectors, flexible LLM choice, and robust APIs, Glean gives organizations the infrastructure to govern, scale, and customize AI across their entire business - without vendor lock-in or costly implementation cycles. \n  \n At its core, Glean is redefining how enterprises find, use, and act on knowledge. Its Enterprise Graph and Personal Knowledge Graph map the relationships between people, content, and activity, delivering deeply personalized, context-aware responses for every employee. This foundation powers Glean’s agentic capabilities - AI agents that automate real work across teams by accessing the industry’s broadest range of data: enterprise and world, structured and unstructured, historical and real-time. The result: measurable business impact through faster onboarding, hours of productivity gained each week, and smarter, safer decisions at every level. \n  \n Recognized by Fast Company as one of the World’s Most Innovative Companies (Top 10, 2025), by CNBC’s Disruptor 50, Bloomberg’s AI Startups to Watch (2026), Forbes AI 50, and Gartner’s Tech Innovators in Agentic AI, Glean continues to accelerate its global impact. With customers across 50+ industries and 1,000+ employees in more than 25 countries, we’re helping the world’s largest organizations make every employee AI-fluent, and turning the superintelligent enterprise from concept into reality. \n  \n If you’re excited to shape how the world works, you’ll help build systems used daily across Microsoft Teams, Zoom, ServiceNow, Zendesk, GitHub, and many more - deeply embedded where people get things done. You’ll ship agentic capabilities on an open, extensible stack, with the craft and care required for enterprise trust, as we bring Work AI to every employee, in every company. \n  \n About the Role: \n Glean is seeking a Site Reliability Engineering Lead to foster a culture of engineering excellence, drive technical strategy, and develop a high-performing, collaborative team. Your role is pivotal in ensuring our services meet stringent Service Level Objectives (SLOs) and in building resilient, automated production environments in the cloud. You'll lead a team and be responsible for products globally, providing technical leadership to key projects and empowering your team to do the same. \n Much of our software development focuses on building infrastructure to scale our operations in a hybrid cloud environment and eliminating work through automation. On the SRE team, you’ll have the opportunity to manage the complex challenges of scale and fast growth which are unique to Glean, while using your expertise in coding, algorithms, problem-solving, and SRE practices. We keep Glean applications up and running, ensuring our customers have the best and most reliable experience possible. \n You will: \n \n Technical Leadership and Mentorship : Play a key role in driving technical excellence and fostering a culture of reliability across engineering teams. You will lead by example, setting best practices for incident management, performance optimization, and automation. Influence best practices, drive cross-team collaborations, and contribute to the execution of key objectives in alignment with engineering leadership and cross-functional partners. Establish strong technical credibility, shaping architectural decisions and ensuring the delivery of high-quality, reliable systems. \n Ensure High Availability: Implement and maintain resilient cloud architectures, monitor system performance, and proactively identify and resolve potential bottlenecks or points of failure.  \n Incident Management: Participate in primary oncall rotation; cultivate technical curiosity and growth mindset, and a blameless postmortem culture within the team. Continuously optimize the on-call process for sustainability and efficiency. \n Automation and Tooling: Develop and maintain automation scripts, tools, and processes to streamline system deployment, monitoring, and management tasks. Your contributions will be vital in efficiently scaling cloud operations. \n Performance Optimization: Optimize cloud infrastructure and applications for performance, scalability, and cost-effectiveness. \n Security and Compliance: Collaborate with security engineers to implement best practices and ensure compliance with security standards and policies. \n Monitoring and Alerting: Design and configure advanced monitoring systems to gain insights into system behavior, set up alerts, and respond proactively to potential issues. Create and maintain comprehensive dashboards and playbooks for production on-call. \n Software Development Consultation: Engage actively in the en","salary_min":200000,"salary_max":260000,"location":"Mountain View, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["llm","distributed-systems","cloud","agents","security","devops"],"apply_url":"https://job-boards.greenhouse.io/gleanwork/jobs/4654833005","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-02-03T23:00:42Z","expires_at":"2026-09-29T13:34:06.249986Z","created_at":"2026-04-13T09:38:55.541153Z","updated_at":"2026-08-30T13:34:06.387503Z","company_name":"Glean","company_slug":"glean","company_logo_url":"https://www.google.com/s2/favicons?domain=glean.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/c9923cc2-371f-4bb8-b7cb-84b54c2f3619"},{"id":"faef06c2-1ca0-46fa-8726-198802cb9f93","company_id":"386fe9d9-0b35-4d37-bdcf-c61d636cf918","title":"Senior DevOps Engineer","slug":"senior-devops-engineer-9801c68e","description":"About EliseAI\n\nAt EliseAI, we're improving the industries that matter most: housing and healthcare. Everyone needs a place to live and access to quality healthcare, yet both are often harder to secure than they should be. \n\nBy integrating AI agents deeply into existing workflows, we make them more efficient, reduce costs, and improve the experience for everyone.\n\n\n\n - Housing: We simplify how renters tour apartments, sign leases, submit maintenance requests, and stay connected with their property team—bringing everything they need for their home into one place.\n\n - Healthcare: We make it easy to schedule appointments, complete intake forms, and we help patients communicate with providers, so everyone can focus on health instead of paperwork.\n   \n   \n\nWith EliseAI, organizations reduce manual work, improve accessibility, and deliver a seamless experience across essential services. We recently raised a $250 million Series E round https://www.eliseai.com/blog/eliseai-raises-250m-series-e led by Andreessen Horowitz to accelerate this mission.\n\n\n\nAbout The Role\n\nAs a DevOps Engineer at EliseAI, you will own the systems and processes that support reliable software deployment across multiple environments. You’ll be responsible for managing configuration, maintaining deployment workflows, and ensuring operational consistency as our infrastructure scales. This role requires close collaboration across engineering, product, and platform teams to support end-to-end delivery—from development through production. You’ll help build the foundation for how we deploy, monitor, and scale our systems as the company continues to grow.\n\n\n\nKey Responsibilities\n\n - Build, maintain, and improve infrastructure using AWS and modern DevOps practices\n\n - Design and implement monitoring, alerting, and incident response systems to ensure high availability\n\n - Automate deployment pipelines and manage CI/CD workflows\n\n - Collaborate with engineers to identify and resolve performance, scalability, and reliability issues\n\n - Improve system security and auditability across environments\n\n - Evaluate and introduce new tools and technologies to enhance operations\n\n\n\nMove at rocket speed, build something massive.\n\nWe’re scaling fast, solving real client problems with precision and ambition. Here, you own your impact; full autonomy, no micromanagement, no fluff. We hire the best, expect the best, and give you the masterclass of your career. It’s hard, it’s intense, and it’s the most rewarding work you’ll ever do. If you’re hungry, driven, and ready to build something massive, climb aboard.\n\n\n\nRequirements\n\n - 3+ years of DevOps or infrastructure engineering experience, preferably at a high-growth startup\n\n - Strong AWS experience, including services like EC2, ECS, RDS, Lambda, and IAM\n\n - Proficiency in scripting languages (preferably Python) and infrastructure-as-code tools (e.g., Terraform)\n\n - Strong software engineering fundamentals and ability to debug and optimize complex systems\n\n - Experience with CI/CD systems such as GitHub Actions or similar\n\n - Ability to thrive in a fast-paced environment and take ownership of large initiatives from day one\n\n - Willingness to work in person at our office 4-5 days a week\n\n\n\nWhy Join\n\nGrowth and impact. It’s not often that you can get in on the ground floor of a funded (unicorn! https://www.eliseai.com/blog/eliseai-raises-250m-series-e) startup that’s scaling so fast. That means that instead of following a playbook, you’ll be writing it. Every single day you will be challenged to identify how we can scale and execute on it. You’ll learn what works when you succeed and what doesn’t when you fail. Either way, the rest of the team will be here to support you.\n\n\n\nBenefits\n\nIn addition to the growth and impact you’ll have at EliseAI, we offer competitive salaries along with the following benefits:\n\n - Equity in the company\n\n - Medical, Dental and Vision premiums covered at 100%\n\n - Fully paid parental leave\n\n - Commuter benefits\n\n - 401k benefits\n\n - Fitness \u0026 home services stipend to cover part of your expenses so you can focus on what matters\n\n - A collaborative in-office environment with an open floor plan, fully stocked kitchen, and all meals covered in the office\n\n - Unlimited vacation and paid holidays\n\n - We'll cover relocation packages and make the move exciting, not painful!\n\n\n\nJob Compensation Range\n\nThe salary range for this role is $230,000 - $320,000. EliseAI offers a competitive total rewards package which includes base salary, equity, and a comprehensive benefits \u0026 perks package. Exact compensation is determined based on a number of factors including experience, skill level, location and qualifications which are assessed during the interview process. Additional details about total compensation and benefits will be provided by our Recruiting Team during the hiring process.\n\n\n\nEliseAI provides equal employment opportunities to all employees and applicants for emplo","salary_min":230000,"salary_max":320000,"location":"New York, NY","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"senior","tags":["cloud","healthcare","agents","devops"],"apply_url":"https://jobs.ashbyhq.com/eliseai/fe19cade-c6ec-4552-8b49-3f45107f2466/application","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-01-13T22:45:55.201Z","expires_at":"2026-09-29T13:48:12.700089Z","created_at":"2026-04-17T02:26:11.392316Z","updated_at":"2026-08-30T13:48:12.832583Z","company_name":"EliseAI","company_slug":"eliseai","company_logo_url":"https://www.google.com/s2/favicons?domain=eliseai.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/faef06c2-1ca0-46fa-8726-198802cb9f93"},{"id":"1e2f1419-efb8-473a-ad2c-2ab9293416cc","company_id":"ec4a8bb4-3840-4054-8ccd-77e81db037af","title":"Senior/Lead Site Reliability Engineer – Federal","slug":"seniorlead-site-reliability-engineer-federal-ddd04a64","description":"C3 AI (NYSE: AI), is the Enterprise AI application software company. C3 AI delivers a family of fully integrated products including the C3 Agentic AI Platform, an end-to-end platform for developing, deploying, and operating enterprise AI applications, C3 AI applications, a portfolio of industry-specific SaaS enterprise AI applications that enable the digital transformation of organizations globally, and C3 Generative AI, a suite of domain-specific generative AI offerings for the enterprise. Learn more at: C3 AI \n C3 AI is seeking a  Senior/Lead Site Reliability Engineer - Federal  to join our team in Tysons, VA or Redwood City, CA. \n This role requires US Citizenship.   Active US Government Security Secret clearance or higher is required (Top Secret or higher is preferred).  \n Responsibilities :\n \n Work with Federal customers to design and implement customized installations of the C3 AI Platform that meet unique access and security requirements of Federal environments\n Maximize system uptime and availability, ensuring functional and performance SLAs\n Establish end-to-end monitoring and alerting on all critical aspects\n Solve complex problems for critical services and build automation to prevent problem recurrence\n Initiate and lead scripting and automation to streamline system updates and upgrades\n Set up critical infrastructure, tools, and framework to streamline the deployment cycle\n Work cross-functionally with Services and Engineering teams\n Travel to customer site (up to 50%)\n \n Qualifications: \n \n Bachelor’s degree in a Science, Technology, Engineering or Mathematics (STEM), or comparable area of study\n An active U.S. Government security clearance (Top Secret preferred)\n Demonstrated experience in deploying, managing, and operating scalable and fault-tolerant Kubernetes-based infrastructure in AWS and Azure clouds; on-premise deployment experience preferred\n Expertise in Linux Operating Systems, Networking, and Database concepts\n Expertise in cloud providers, such as Amazon Web Services, Azure, and GCP\n Experience with Infrastructure-as-Code configurations such as Terraform, Ansible, or Puppet\n Experience in Ruby, Bash, or Python; to automate and monitor systems\n Excellent problem-solving, critical thinking, and communication skills\n Experience supporting as a DevOps or sys admin for commercial SaaS solutions. Customer facing experience is a plus.\n \n Candidates must be authorized to work in the United States without the need for current or future company sponsorship. \n C3 AI provides excellent benefits, a competitive compensation package and generous equity plan. \n California Base Pay Range\n $159,000 — $230,000 USD \n C3 AI is proud to be an Equal Opportunity and Affirmative Action Employer. We do not discriminate on the basis of any legally protected characteristics, including disabled and veteran status.","salary_min":159000,"salary_max":230000,"location":"Tysons, VA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["agents","cloud","generative-ai","devops"],"apply_url":"https://c3.ai/job-description/8198282002?gh_jid=8198282002","is_featured":false,"is_sticky":false,"status":"active","published_at":"2025-10-03T21:36:09Z","expires_at":"2026-09-29T13:40:13.129074Z","created_at":"2026-04-13T15:01:26.553639Z","updated_at":"2026-08-30T13:40:13.267042Z","company_name":"C3 AI","company_slug":"c3-ai","company_logo_url":"https://www.google.com/s2/favicons?domain=c3.ai\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/1e2f1419-efb8-473a-ad2c-2ab9293416cc"},{"id":"444bc7d2-2c8a-45e9-ae3f-d005f70c2dd2","company_id":"861968d1-d9f8-4217-9873-ce4b24851abc","title":"Senior Manager of Data Science Production Engineering, DevOps","slug":"senior-manager-of-data-science-production-engineering-devops-2f3e42d5","description":"This is an exciting opportunity to lead a DevOps team at Natera, a global leader in precision medicine and genomics testing.  As a Senior Manager of the DevOps team inside our Data Science Production Engineering (DSPE) department, you will have the opportunity to work with Natera's diverse and ultra-large data sets (up to PB size), build upon the latest information and cloud technologies, create solutions for the most cutting-edge medical diagnostics and genomics applications, and make real impact by delivering reliable, efficient, high performance data/cloud/automation systems to support Natera's Lab Operations and Genetic Counselor/Lab Director teams, and ultimately impact patients' medical outcomes.\n PRIMARY RESPONSIBILITIES:  \n Leadership \n \n Lead a DevOps team, working closely with other DSPE teams, to develop, test, deploy, maintain, and update DSPE's data platforms, cloud platforms, automation systems, and analytics systems\n Manage the DevOps team, to achieve the timely and efficient delivery of software products and systems with correct and robust functionality.\n Train and coach team members on new technologies, procedures, and guidelines/best practices.\n Create and maintain operational visibility (e.g., workload, productivity, etc.) of the DSPE DevOps team.\n \n Technical \n \n Translate DSPE's business needs and project requirements into technical specifications and implementation plans.\n Working with DSPE leadership teams, create suitable system architecture and infrastructure design documents that can be efficiently implemented and maintained.\n Maintain and optimize infrastructure performance, monitoring key system metrics including network I/O, disk I/O, memory allocation, and compute resource utilization.\n Contribute to key development, testing, deployment, validation, maintenance, and update activities.\n \n Quality, Compliance, and Documentation \n \n Improve and enforce procedures for work efficiency, quality, and compliance.\n Lead and manage the technical documentation's creation and maintenance/update for all DSPE DevOps owned software tools and systems, for example, design specifications, software verification and validation (V\u0026V) protocols and reports, deployment/update/maintenance records, user guides, etc.\n Create training materials, and contribute to training activities, on DSPE DevOps owned software tools and systems.\n \n Cross-Functional \n \n Engage proactively and directly with engineering and operational stakeholders to align infrastructure roadmaps, resolve technical dependencies, and report platform performance metrics.\n Represent the DSPE DevOps team in stakeholder meetings.\n \n QUALIFICATIONS: \n \n Minimal Bachelor's Degree in Computer Science, Software Engineering, Bioinformatics, or a related field.  Advanced degrees preferred.\n Minimal 10 years of relevant industry experience for candidates with only a Bachelor's Degree.  Minimal 8 years of relevant industry experience for candidates with a Master's Degree.  Minimal 4 years of relevant industry experience for candidates with a Doctorate degree.\n Minimal 3 years of working experience in a highly regulated environment, e.g., CLIA and/or FDA compliant settings.\n Minimal 3 years of management experience leading a DevOps team, or a software development team.\n Demonstrable experience in designing, delivering, and maintaining/updating custom-developed software tools or systems for production environments, preferable in high-throughput settings (e.g., heavy traffic and/or high computation load).\n Demonstrable experience in creating software design specification documents, verification and validation documents, and user manuals.\n \n KNOWLEDGE, SKILLS, AND ABILITIES: \n \n Excellent Python programming skills, with substantial cloud application and/or automation system development experience.\n Proficient with RDBMS systems such as PostgreSQL and/or MySQL, and substantial working experience in for-production database-based application development.\n Working knowledge in optimizing system architecture and individual components for performance, scalability, and/or reliability.\n Proficiency in tool/system performance evaluation methods, and in-depth understanding of key factors, such as RAM, Disk IO, Network IO's impact on system/tool's performance.\n Proficiency with AWS, and substantial working experience in developing for-production software tools/systems using AWS services.\n Proficiency in SDLC and software development and testing methodologies.\n Proficient Linux command line skills, and past Linux application development experience.\n Past experience and working knowledge in genomics and bioinformatics technologies.\n Good understanding and past working experience with genomics data, such as FASTQ, BAM, VCF files.\n Proficiency in Git workflow and version control systems like GitHub and/or GitLab.\n Past experience in developing automation solutions.\n Preferred but not required: IaC skills and experiences, such as CloudFormation, Terraform, and/or ","salary_min":179900,"salary_max":224850,"location":"San Carlos, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"senior","tags":["cloud","devops","data-science"],"apply_url":"https://job-boards.greenhouse.io/natera/jobs/6148527004","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-21T18:58:58Z","expires_at":"2026-09-29T13:40:55.056852Z","created_at":"2026-08-25T18:29:44.02412Z","updated_at":"2026-08-30T13:40:55.190658Z","company_name":"Natera","company_slug":"natera","company_logo_url":"https://www.google.com/s2/favicons?domain=natera.com\u0026sz=128","quality_score":85,"url":"https://aidevboard.com/job/444bc7d2-2c8a-45e9-ae3f-d005f70c2dd2"},{"id":"4aa13c5b-6572-4969-a74e-c9398da2f32b","company_id":"861968d1-d9f8-4217-9873-ce4b24851abc","title":"Senior Manager of Data Science Production Engineering, DevOps","slug":"senior-manager-of-data-science-production-engineering-devops-a9884303","description":"This is an exciting opportunity to lead a DevOps team at Natera, a global leader in precision medicine and genomics testing.  As a Senior Manager of the DevOps team inside our Data Science Production Engineering (DSPE) department, you will have the opportunity to work with Natera's diverse and ultra-large data sets (up to PB size), build upon the latest information and cloud technologies, create solutions for the most cutting-edge medical diagnostics and genomics applications, and make real impact by delivering reliable, efficient, high performance data/cloud/automation systems to support Natera's Lab Operations and Genetic Counselor/Lab Director teams, and ultimately impact patients' medical outcomes.\n PRIMARY RESPONSIBILITIES:  \n Leadership \n \n Lead a DevOps team, working closely with other DSPE teams, to develop, test, deploy, maintain, and update DSPE's data platforms, cloud platforms, automation systems, and analytics systems\n Manage the DevOps team, to achieve the timely and efficient delivery of software products and systems with correct and robust functionality.\n Train and coach team members on new technologies, procedures, and guidelines/best practices.\n Create and maintain operational visibility (e.g., workload, productivity, etc.) of the DSPE DevOps team.\n \n Technical \n \n Translate DSPE's business needs and project requirements into technical specifications and implementation plans.\n Working with DSPE leadership teams, create suitable system architecture and infrastructure design documents that can be efficiently implemented and maintained.\n Maintain and optimize infrastructure performance, monitoring key system metrics including network I/O, disk I/O, memory allocation, and compute resource utilization.\n Contribute to key development, testing, deployment, validation, maintenance, and update activities.\n \n Quality, Compliance, and Documentation \n \n Improve and enforce procedures for work efficiency, quality, and compliance.\n Lead and manage the technical documentation's creation and maintenance/update for all DSPE DevOps owned software tools and systems, for example, design specifications, software verification and validation (V\u0026V) protocols and reports, deployment/update/maintenance records, user guides, etc.\n Create training materials, and contribute to training activities, on DSPE DevOps owned software tools and systems.\n \n Cross-Functional \n \n Engage proactively and directly with engineering and operational stakeholders to align infrastructure roadmaps, resolve technical dependencies, and report platform performance metrics.\n Represent the DSPE DevOps team in stakeholder meetings.\n \n QUALIFICATIONS: \n \n Minimal Bachelor's Degree in Computer Science, Software Engineering, Bioinformatics, or a related field.  Advanced degrees preferred.\n Minimal 10 years of relevant industry experience for candidates with only a Bachelor's Degree.  Minimal 8 years of relevant industry experience for candidates with a Master's Degree.  Minimal 4 years of relevant industry experience for candidates with a Doctorate degree.\n Minimal 3 years of working experience in a highly regulated environment, e.g., CLIA and/or FDA compliant settings.\n Minimal 3 years of management experience leading a DevOps team, or a software development team.\n Demonstrable experience in designing, delivering, and maintaining/updating custom-developed software tools or systems for production environments, preferable in high-throughput settings (e.g., heavy traffic and/or high computation load).\n Demonstrable experience in creating software design specification documents, verification and validation documents, and user manuals.\n \n KNOWLEDGE, SKILLS, AND ABILITIES: \n \n Excellent Python programming skills, with substantial cloud application and/or automation system development experience.\n Proficient with RDBMS systems such as PostgreSQL and/or MySQL, and substantial working experience in for-production database-based application development.\n Working knowledge in optimizing system architecture and individual components for performance, scalability, and/or reliability.\n Proficiency in tool/system performance evaluation methods, and in-depth understanding of key factors, such as RAM, Disk IO, Network IO's impact on system/tool's performance.\n Proficiency with AWS, and substantial working experience in developing for-production software tools/systems using AWS services.\n Proficiency in SDLC and software development and testing methodologies.\n Proficient Linux command line skills, and past Linux application development experience.\n Past experience and working knowledge in genomics and bioinformatics technologies.\n Good understanding and past working experience with genomics data, such as FASTQ, BAM, VCF files.\n Proficiency in Git workflow and version control systems like GitHub and/or GitLab.\n Past experience in developing automation solutions.\n Preferred but not required: IaC skills and experiences, such as CloudFormation, Terraform, and/or ","salary_min":151000,"salary_max":188700,"location":"Remote (US)","workplace":"remote","remote_scope":"restricted","job_type":"full-time","experience_level":"senior","tags":["cloud","data-science","devops"],"apply_url":"https://job-boards.greenhouse.io/natera/jobs/6137673004","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-21T18:58:56Z","expires_at":"2026-09-29T13:40:54.954898Z","created_at":"2026-08-25T18:29:44.016104Z","updated_at":"2026-08-30T13:40:55.095479Z","company_name":"Natera","company_slug":"natera","company_logo_url":"https://www.google.com/s2/favicons?domain=natera.com\u0026sz=128","quality_score":85,"url":"https://aidevboard.com/job/4aa13c5b-6572-4969-a74e-c9398da2f32b"}],"page":1,"per_page":20,"total":94,"total_is_exact":true,"total_pages":5}
