{"access":{"catalog_url":"https://aidevboard.com/api/v1/catalog","description":"Public read endpoints are open and free. API keys are optional for stable agent identity and keyed hourly throttling.","docs_url":"https://aidevboard.com/docs","employer_pilot_url":"https://aidevboard.com/verified-interview-pilot","mode":"open","register_url":"https://aidevboard.com/api/v1/register"},"candidate_resume_action":{"application_authorized":false,"candidate_charge":0,"endpoint":"https://aidevboard.com/api/v1/candidate/resume-preview","job_id_json_path":"jobs[].id","method":"POST","preview_requires_identity":false,"required_body_fields":["job_id","evidence_bullets"],"requires_explicit_human_review":true,"saved_artifact_protocol":"mcp","saved_artifact_requires_verified_human":true,"saved_artifact_tool":"compile_job_specific_resume","search_requires_identity":false,"status":"available_after_candidate_selects_job","submission_performed":false,"uses_candidate_verified_evidence":true},"degraded":false,"estimated":false,"has_next":true,"jobs":[{"id":"ef4c6406-333f-41ab-8184-e9bd7afe7c50","company_id":"a0000000-0000-0000-0000-000000000009","title":"Forward Deployed Engineer, Infrastructure Specialist (North America)","slug":"forward-deployed-engineer-infrastructure-specialist-north-america-85c940b9","description":"Who are we?\n\nCohere is the leading security-first enterprise AI company.  We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.\n\nWe’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that.\n\nWe obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft.\n\nWe are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us!\n\n\nABOUT NORTH:\n\nNorth https://cohere.com/north is Cohere's cutting-edge AI workspace platform, designed to revolutionize the way enterprises utilize AI. It offers a secure and customizable environment, allowing companies to deploy AI while maintaining control over sensitive data. North integrates seamlessly with existing workflows, providing a trusted platform that connects AI agents with workplace tools and applications.\n\n\n\n\nWHY THIS ROLE?\n\nThis role offers a unique opportunity to shape how enterprises harness the power of AI in real-world applications. As a bridge between our core North product and our clients’ engineering teams, you’ll be at the forefront of solving complex problems and securely integrating AI into critical sectors such as finance, healthcare, and telecommunications. Our esteemed clients include industry leaders like RBC, Dell, and LG CNS.\n\nWe are seeking engineers who deeply care about customers and want to work at the cutting edge of Agentic AI.\n\n\n\nIn this role, you will:\n\n - Lead end-to-end deployment of North in private cloud and on-premises environments, including planning, configuration, testing, and rollout.\n\n - Partner with enterprise IT teams to assess infrastructure, security requirements, and data management practices.\n\n - Experiment at a high velocity and with a high level of quality to engage our customers and ultimately deliver solutions that exceed their expectations\n\n - Design and implement deployment strategies tailored to client needs, ensuring compliance with data privacy and security standards.\n\n - Troubleshoot and resolve deployment-related technical issues, providing timely solutions to minimize downtime.\n\n\n\nYou may be a good fit if:\n\n - You have experience with and enjoy working directly with customers\n\n - You have experience deploying enterprise software in private/hybrid cloud environments\n\n - You have proven experience administering production Kubernetes clusters and expertise with Helm\n\n - Familiarity with DevOps practices, CI/CD pipelines, and tools like Git for version control\n\n - You have strong expertise in cloud infrastructure (Azure, AWS, GCP), networking, and virtualization\n\n - You excel in fast-paced environments and can execute while priorities and objectives are a moving target\n\nCohere is committed to fair and transparent pay practices. The salary range listed for this role reflects the expected base compensation. Actual compensation offered will be determined by factors such as location, level, job-related knowledge, skills, education, and experience.\n\n - United States:\n   \n   - For candidates based in California, New York and Washington States, the compensation range is: $140,000 - $325,000 USD\n   \n   - For candidates based elsewhere in the US, the compensation range is: $120,000 – $275,000 USD\n\n - Canada:\n   \n   - For candidates in Canada, the Compensation Range is : $175,000 - $385,000 CAD\n\n\n\n\nFULL-TIME EMPLOYEES AT COHERE ENJOY THESE PERKS:\n\n - A weekly lunch stipend of $75/£75 or equivalent in your local currency for lunch.\n\n - Full health and dental benefits, including a separate budget for mental health.\n\n - RRSP matching, 401K, Pension Scheme.\n\n - 100% Parental Leave top-up for up to 6 months, for either parent.\n\n - Annual enrichment benefits:\n   \n   Arts \u0026 culture, fitness/wellness, quality time, and a workspace improvement credit.\n   \n   Education \u0026 learning stipend for conferences, courses, and coaching.\n\n - 6 weeks of paid vacation (30 working days!)\n\n - Budget for traveling to other offices if you are remote, plus an annual company offsite.\n\n\n\n\nHOW AND WHERE WE WORK:\n\n - Cohere is remote-friendly, but we also have offices in Toronto, London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul with more opening soon.\n\n - For those in the office: a daily lunch program, plenty of snacks, and regular community and social events.\n\n - For those not near an office: a co-working benefit so you can work alongside others in your city.\n\n - Everyone receives a $500 home office stipend to set up your workspace properly.\n   \n   \n\nIf any of the above doesn’t line up exactly with your exp","salary_min":175000,"salary_max":385000,"location":"Toronto, Canada","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"senior","tags":["healthcare","payments","agents","cloud","infrastructure"],"apply_url":"https://jobs.ashbyhq.com/cohere/be48aafc-9610-4ebd-8414-a0722a3cd59a/application","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-28T20:36:13.897Z","expires_at":"2026-09-29T13:31:52.797456Z","created_at":"2026-04-13T09:36:53.959441Z","updated_at":"2026-08-30T13:31:52.939334Z","company_name":"Cohere","company_slug":"cohere","company_logo_url":"https://www.google.com/s2/favicons?domain=cohere.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/ef4c6406-333f-41ab-8184-e9bd7afe7c50"},{"id":"17669270-764c-4f80-a496-6cf3009fcf77","company_id":"a0000000-0000-0000-0000-000000000009","title":"Forward Deployed Engineer, Infrastructure Specialist (Public Sector)","slug":"forward-deployed-engineer-infrastructure-specialist-public-sector-4cfcc1d7","description":"Who are we?\n\nCohere is the leading security-first enterprise AI company.  We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.\n\nWe’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that.\n\nWe obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft.\n\nWe are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us!\n\n\n\n\nABOUT NORTH:\n\nNorth https://cohere.com/north is Cohere's cutting-edge AI workspace platform, designed to revolutionize the way enterprises utilize AI. It offers a secure and customizable environment, allowing companies to deploy AI while maintaining control over sensitive data. North integrates seamlessly with existing workflows, providing a trusted platform that connects AI agents with workplace tools and applications.\n\n\n\n\nWHY THIS ROLE?\n\nCohere’s team partners with Canadian public sector organisations to unlock transformative value through secure, ethical deployment of Generative AI (GenAI) solutions. We work collaboratively to address complex societal challenges while maintaining the highest standards of data security and compliance. You will work directly with public sector customers to quickly understand their greatest problems and design and implement solutions using Cohere's stack.\n\nThis role offers a unique opportunity to shape how enterprises harness the power of AI in real-world applications. As a bridge between our core North product and our clients’ engineering teams, you’ll be at the forefront of solving complex problems and securely integrating AI into critical sectors.\n\nWe are seeking engineers with diverse skill sets, including backend, infrastructure, agent development, and deployments, who deeply care about customers and want to work at the cutting edge of Agentic AI.\n\n\n\nLocation: Ottawa, 20-40% travel anticipated.\n\n\nSecurity Clearance: Active Top Secret clearance strongly preferred; candidates eligible and willing to obtain clearance will also be considered. If you are ineligible for clearance there are other positions on our careers site that do not have this requirement.\n\nMore information about Canadian Security Clearance can be found here https://www.canada.ca/en/public-services-procurement/services/industrial-security/security-requirements-contracting/personnel-security-screening/processes/security-clearance-request.html.\n\n\n\nIn this role, you will:\n\n -  Lead end-to-end deployment of North in private cloud and on-premises environments, including planning, configuration, testing, and rollout.\n\n - Partner with enterprise IT teams to assess infrastructure, security requirements, and data management practices.\n\n - Experiment at a high velocity and with a high level of quality to engage our customers and ultimately deliver solutions that exceed their expectations\n\n - Design and implement deployment strategies tailored to client needs, ensuring compliance with data privacy and security standards.\n\n - Troubleshoot and resolve deployment-related technical issues, providing timely solutions to minimize downtime.\n\nYou may be a good fit if:\n\n - You have experience with and enjoy working directly with customers\n\n - You have experience deploying enterprise software in private/hybrid cloud environments\n\n - You have proven experience administering production Kubernetes clusters and expertise with Helm\n\n - Familiarity with DevOps practices, CI/CD pipelines, and tools like Git for version control\n\n - You have strong expertise in cloud infrastructure (Azure, AWS, GCP), networking, and virtualization\n\n - You excel in fast-paced environments and can execute while priorities and objectives are a moving target\n\n - Familiarity with Canadian public sector security and compliance requirements (e.g., data sovereignty, access controls).\n\n\n\nCohere is committed to fair and transparent pay practices. The salary range listed for this role reflects the expected base compensation. Actual compensation offered will be determined by factors such as location, level, job-related knowledge, skills, education, and experience.\n\n - Canada:\n   \n   - For candidates in Canada, the Compensation Range is : $175,000 - $385,000 CAD\n\n\n\n\n\n\nFULL-TIME EMPLOYEES AT COHERE ENJOY THESE PERKS:\n\n - A weekly lunch stipend of $75/£75 or equivalent in your local currency for lunch.\n\n - Full health and dental benefits, including a separate budget for mental health.\n\n - RRSP matching, 401K, Pension Scheme.\n\n - 100% Parental Leave top-up for up to 6 months, for either parent.\n\n - Annual enrich","salary_min":175000,"salary_max":385000,"location":"Ottawa","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"senior","tags":["agents","payments","cloud","generative-ai","infrastructure"],"apply_url":"https://jobs.ashbyhq.com/cohere/52a2b83b-7537-4e88-af7b-e4e9630a96e0/application","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-28T20:35:23.101Z","expires_at":"2026-09-29T13:31:55.712365Z","created_at":"2026-04-30T05:46:56.273345Z","updated_at":"2026-08-30T13:31:55.852872Z","company_name":"Cohere","company_slug":"cohere","company_logo_url":"https://www.google.com/s2/favicons?domain=cohere.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/17669270-764c-4f80-a496-6cf3009fcf77"},{"id":"49815b01-da50-4ed6-905c-a071c44a3cab","company_id":"a0000000-0000-0000-0000-000000000009","title":"Engineering Manager, FDE Infrastructure (NORAM)","slug":"engineering-manager-fde-infrastructure-noram-5b3d8276","description":"Who are we?\n\nCohere is the leading security-first enterprise AI company.  We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.\n\nWe’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that.\n\nWe obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft.\n\nWe are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us!\n\nWe're looking for an Engineering Manager to lead our Deployment Engineering team. This isn't a typical management role — we need someone who leads from the front, gets their hands dirty, and drives impact. You'll manage a team of Forward Deployed Engineers who are on the front lines of deploying Cohere's North platform into customer environments. You should be ready to be a force to be reckoned with.\n\n\n\nLocation: North America (remote-first)\n\n\n\n\nWHAT YOU'LL DO\n\n - Lead and mentor a team of Forward Deployed Engineers \n\n - Drive end-to-end deployment of North in private cloud and on-premises environments\n\n - Take ownership of customer success from technical implementation through delivery\n\n - Collaborate closely with Product, Engineering, and Sales to shape how we deliver AI to enterprises\n\n - Mentor your team on cloud infrastructure, Kubernetes, and enterprise-grade deployments\n\n - Optimize performance for OpenSearch, databases, and other K8s services\n\n - Define scaling guidelines for GPU and CPU compute resources\n\n - Build processes \u0026 technology that scales — we're growing fast.\n\n\n\n\nWHAT WE'RE LOOKING FOR\n\n - 5+ years of experience in software engineering with demonstrated leadership\n\n - Hands-on experience deploying enterprise software at scale\n\n - Strong expertise in cloud infrastructure (Azure, AWS, GCP)\n\n - Experience with Kubernetes, Helm, and CI/CD pipelines\n\n - High agency — you don't wait for permission to solve problems\n\n - A doer mentality — you dig in and get stuff done\n\n - Experience managing engineers in a fast-paced, high-growth environment\n\n - Excellent communication skills in English.\n\n\n\n\nNICE TO HAVE'S\n\n - Fluency in additional European languages\n\n - Experience with AI/ML infrastructure\n\n - Background in enterprise security and compliance.\n\n\n\nCohere is committed to fair and transparent pay practices. The salary range listed for this role reflects the expected base compensation. Actual compensation offered will be determined by factors such as location, level, job-related knowledge, skills, education, and experience.\n\n - United States:\n   \n   - For candidates based in California, New York and Washington States, the compensation range is: $140,000 - $325,000 USD\n   \n   - For candidates based elsewhere in the US, the compensation range is: $120,000 – $275,000 USD\n\n - Canada:\n   \n   - For candidates in Canada, the Compensation Range is : $175,000 - $385,000 CAD\n\n\n\n\nFULL-TIME EMPLOYEES AT COHERE ENJOY THESE PERKS:\n\n - A weekly lunch stipend of $75/£75 or equivalent in your local currency for lunch.\n\n - Full health and dental benefits, including a separate budget for mental health.\n\n - RRSP matching, 401K, Pension Scheme.\n\n - 100% Parental Leave top-up for up to 6 months, for either parent.\n\n - Annual enrichment benefits:\n   \n   Arts \u0026 culture, fitness/wellness, quality time, and a workspace improvement credit.\n   \n   Education \u0026 learning stipend for conferences, courses, and coaching.\n\n - 6 weeks of paid vacation (30 working days!)\n\n - Budget for traveling to other offices if you are remote, plus an annual company offsite.\n\n\n\n\nHOW AND WHERE WE WORK:\n\n - Cohere is remote-friendly, but we also have offices in Toronto, London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul with more opening soon.\n\n - For those in the office: a daily lunch program, plenty of snacks, and regular community and social events.\n\n - For those not near an office: a co-working benefit so you can work alongside others in your city.\n\n - Everyone receives a $500 home office stipend to set up your workspace properly.\n   \n   \n\nIf any of the above doesn’t line up exactly with your experience, we still encourage you to apply. \n\n\nWe strive to create an inclusive work environment for all; we welcome applicants from all backgrounds and are committed to providing equal opportunities. Should you require any accommodations during the recruitment process, please submit an Accommodations Request Form https://docs.google.com/forms/d/12a6IrLdF3kI2nonKSr4tiFuz18rLQbaeYV-JM9L4o9Q/edit, and we will work together to meet your needs.\n\n\n\nWe may use AI-enabled tools to scr","salary_min":175000,"salary_max":385000,"location":"Canada","workplace":"remote","remote_scope":"restricted","job_type":"full-time","experience_level":"senior","tags":["payments","search","cloud","infrastructure"],"apply_url":"https://jobs.ashbyhq.com/cohere/6a6120d5-5e02-4811-99d9-6baf0b910e37/application","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-28T20:34:05.703Z","expires_at":"2026-09-29T13:31:56.992297Z","created_at":"2026-06-28T14:01:30.761519Z","updated_at":"2026-08-30T13:31:57.195425Z","company_name":"Cohere","company_slug":"cohere","company_logo_url":"https://www.google.com/s2/favicons?domain=cohere.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/49815b01-da50-4ed6-905c-a071c44a3cab"},{"id":"2c486dec-53f7-41d2-9720-2e9a10653b83","company_id":"e3915539-5a8f-4461-9f26-06366a918674","title":"Flight Software (FSW) Engineer - Platform - HITL Infrastructure","slug":"flight-software-fsw-engineer-platform-hitl-infrastructure-a61623fb","description":"Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the expertise, technology, and business model of the 21st century’s most innovative companies to the defense industry, Anduril is changing how military systems are designed, built and sold. Anduril’s family of systems is powered by Lattice OS, an AI-powered operating system that turns thousands of data streams into a realtime, 3D command and control center. As the world enters an era of strategic competition, Anduril is committed to bringing cutting-edge autonomy, AI, computer vision, sensor fusion, and networking technology to the military in months, not years.\n ABOUT THE JOB\n We are looking for a HITL and integration test focused software engineer to develop and proliferate software infrastructure to make bringing up, managing, and running tests against integrated testbeds easy. This role will be accountable to business problems around testbed configuration management, remote access and OPSEC segregation, integration test framework development, and third-party tool integration for requirements traceability and data review.  Experience with HITL operations, including real-time simulation, avionics integration, and closed-loop testing with physical hardware is highly valued. If you have a passion for flight and want to play a core role in enabling the Anduril AD\u0026S organization to continue scaling by orders of magnitude, this is the right role for you!\n  \n WHAT YOU'LL DO \n \n Develop a robust integration test framework capable of executing tests across SITL and HITL environments.\n Create scale-able infrastructure for defining and managing the configuration of testbeds.\n Enable shared access of hardware resources across OPSEC isolation boundaries.\n Proliferate common HITL and test infrastructure across programs ranging from group 5 aircraft, to autonomous boats, to space-based interceptors, and more.\n \n  \n REQUIRED QUALIFICATIONS \n \n Bachelor's Degree or higher in engineering, computer science, mathematics, physics, chemistry, or related technical discipline\n Advanced proficiency with Python\n Familiarity with C/C++\n Hands-on experience with HITL or SITL simulation operations, including test setup, execution, and debugging\n Must be eligible to obtain and maintain a U.S. TS clearance\n \n  \n PREFERRED QUALIFICATIONS \n \n Advanced proficiency with C/C++\n Experience with developing sub-system and system level models (Air vehicle, ground vehicles, ground based sensors and other types of complex systems)\n Flight test support experience, including on-site operations and real-time troubleshooting\n Experience with developing emulated/simulation interfaces for on-board avionics, including CAN, Ethernet, serial protocols, and custom hardware interfaces\n Demonstrated ability to own and operate HITL labs supporting multiple concurrent flight test programs \n US Salary Range\n $129,000 — $220,000 USD \n The salary range for this role is an estimate based on a wide range of compensation factors, inclusive of base salary only. Actual salary offer may vary based on (but not limited to) work experience, education and/or training, critical skills, and/or business considerations. Highly competitive equity grants are included in the majority of full time offers; and are considered part of Anduril's total compensation package. Additionally, Anduril offers top-tier benefits for full-time employees, including:   \n  \n Benefits \n At Anduril, we invest in our people. Our comprehensive, competitive benefits package (available at little to no cost to employees) ensures you’re supported in health, recovery, and whatever comes next.  For more information, Explore Our Benefits . \n  \n \n Protecting Yourself from Recruitment Scams \n Anduril is committed to maintaining the integrity of our Talent acquisition process and the security of our candidates. We've observed a rise in sophisticated phishing and fraudulent schemes where individuals impersonate Anduril representatives, luring job seekers with false interviews or job offers. These scammers often attempt to extract payment or sensitive personal information.\n \n To ensure your safety and help you navigate your job search with confidence, please keep the following critical points in mind:\n \n \n No Financial Requests:  Anduril will never solicit payment or demand personal financial details (such as banking information, credit card numbers, or social security numbers) at any stage of our hiring process. Our legitimate recruitment is entirely free for candidates.\n Please always verify communications: \n \n Direct from Anduril: If you receive an email from one of our recruiters, it will only come from an @anduril.com address.\n Via Agency Partner: If contacted by a recruiting agency for an Anduril role, their email will clearly identify their agency. If you suspect any suspicious activity, please verify the agency's authenticity by reachin","salary_min":129000,"salary_max":220000,"location":"Costa Mesa, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"mid","tags":["computer-vision","payments","cloud","infrastructure"],"apply_url":"https://boards.greenhouse.io/andurilindustries/jobs/5054523007?gh_jid=5054523007","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-28T16:39:50Z","expires_at":"2026-09-29T13:37:14.496802Z","created_at":"2026-08-29T13:37:43.589294Z","updated_at":"2026-08-30T13:37:14.631222Z","company_name":"Anduril","company_slug":"anduril","company_logo_url":"https://www.google.com/s2/favicons?domain=anduril.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/2c486dec-53f7-41d2-9720-2e9a10653b83"},{"id":"7aa92365-6b3e-4f51-a168-2674aa7d8a06","company_id":"76c63eb7-c307-4322-8c2b-c20216feec49","title":"Senior Data Infrastructure Engineer","slug":"senior-data-infrastructure-engineer-96f0f9f8","description":"Get to Know Us\n\nHorizon3 is a fast-growing, remote cybersecurity company dedicated to the mission of enabling organizations to proactively find and fix and verify exploitable attack vectors before criminals exploit them. Our flagship product, the NodeZeroTM platform, delivers production-safe autonomous pentests and other key assessment operations that scale across the largest internal, external, cloud, and hybrid cloud environments. NodeZero has been adopted by organizations of all sizes, from small educational institutions to government agencies and Global 100 enterprises. It is used by ITOps/SecOps teams, consulting pentesters, and MSSPs and MSPs. \n\nWe are a fusion of former U.S. Special Operations cyber operators, startup engineers, and formerly frustrated cybersecurity practitioners. We're committed to helping solve our common security problems: ineffective security tools, false positives resulting in alert fatigue, blind spots, \"checkbox” security culture, cybersecurity skills shortage, and the long lead time and expense of hiring outside consultants. Collectively, we are a team of learn it alls, committed to a culture of respect, collaboration, ownership, and results.\n\nWe are seeking a Senior Data Infrastructure Engineer to design, build, and own the data platform underneath our business-critical pipelines. Today the team supports account provisioning, product analytics, and the pipelines feeding our warehouse and Tableau reporting. You will help expand this into a globally deployed data platform spanning the warehouse, orchestration, and BI across multiple regions\n\nWhat You’ll Do\n\n - Design, build, and maintain business-critical data pipelines, including data ingestion and AWS Glue/Spark jobs feeding the warehouse and Tableau.\n\n - Build and operate core data infrastructure: the data warehouse, orchestration, and BI/reporting platform, deployed across multiple regions.\n\n - Help evaluate and shape our warehouse strategy, including lakehouse platforms such as Databricks.\n\n - Define and evolve event-driven contracts between producer and consumer teams, reasoning through batch vs. streaming, CDC, reprocessing, and multi-tenant/multi-region trade-offs.\n\n - Design and evolve warehouse schemas and data models.\n\n - Build observability into pipelines and platform components.\n\n - Drive work across producer and consumer teams, communicate clearly, and own outcomes end to end.\n\nWhat You’ll Bring\n\n - Strong software engineering craft: clean, correct, maintainable code with solid data-structure and problem-solving fundamentals.\n\n - Deep understanding of distributed systems and data architecture: pipeline and platform design, batch vs. streaming, CDC, event contracts, multi-tenant and multi-region trade-offs, reprocessing, and observability.\n\n - Fluent SQL and strong data-modeling skills, including designing schemas for a warehouse your team owns.\n\n - Experience owning production data infrastructure end to end.\n\n - Strong collaboration and ownership across dependent teams.\n\n - 7+ years of professional experience in data engineering, infrastructure engineering, or backend software engineering.\n\n - Bachelor’s degree in Computer Science or a related field, or equivalent practical experience.\n\n\n\nRequired Tech Stack Experience\n\n - AWS data stack, including Glue and Spark.\n\n - Cloud data warehouse or lakehouse platforms such as Amazon Redshift or Databricks.\n\n - Tableau or a comparable BI platform.\n\n - SQL and Python.\n\n - Workflow orchestration tooling such as Apache Airflow or Dagster\n\n \n\nPerks of Horizon3\n\n - Inclusive Team: We value diversity and promote an inclusive culture where everyone can thrive.\n\n - Growth Opportunities: Be part of a dynamic and growing team with numerous career development opportunities.\n\n - Innovative Culture: Work in a collaborative environment that encourages creativity and out-of-the-box thinking.\n\n - Hybrid \u0026 Remote Work: We embrace a mix of remote and hybrid work models depending on role and location, including our Chicago office, where some roles require regular in-office presence.\n\n - Competitive Compensation: We offer competitive salary, equity and benefits. Our benefits include health, vision \u0026 dental insurance for you and your family, a flexible vacation policy, and generous parental leave.\n\n\n\nCompensation and Values\n\nAt Horizon3, we believe that our people are our greatest asset, and our compensation philosophy reflects this core value. We are committed to fostering an environment where all employees feel valued, respected, and rewarded for their contributions. Our compensation structure is designed to be fair, competitive, and transparent, ensuring that every team member is recognized and compensated equitably across roles, levels, and locations.\n\nIn accordance with various State’s transparency regulations, we provide the following salary range information for this position:\n\n - Base salary range: $202,190 - $237,870 annually. The exact salary will be determined based on","salary_min":202190,"salary_max":237870,"location":"Remote (US)","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"senior","tags":["data-pipeline","cloud","distributed-systems","security","infrastructure"],"apply_url":"https://jobs.ashbyhq.com/horizon3ai/4651e6b3-a8b5-484f-ac87-5b6ca04581d1/application","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-27T16:58:31.284Z","expires_at":"2026-09-29T13:36:43.052847Z","created_at":"2026-08-29T13:37:07.314282Z","updated_at":"2026-08-30T13:36:43.188509Z","company_name":"Horizon3 AI","company_slug":"horizon3-ai","company_logo_url":"https://www.google.com/s2/favicons?domain=horizon3.ai\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/7aa92365-6b3e-4f51-a168-2674aa7d8a06"},{"id":"8d909b3d-3eac-45d3-8804-379bc82bc41b","company_id":"fa25a1f6-acd0-42b6-a229-f4d258ed5c3d","title":"Senior Software Engineer","slug":"senior-software-engineer-91b56607","description":"Employee Applicant Privacy Notice \n Who we are: \n \n Shape a brighter financial future with us.\n Together with our members, we’re changing the way people think about and interact with personal finance.\n We’re a next-generation financial services company and national bank using innovative, mobile-first technology to help our millions of members reach their goals. The industry is going through an unprecedented transformation, and we’re at the forefront. We’re proud to come to work every day knowing that what we do has a direct impact on people’s lives, with our core values guiding us every step of the way. Join us to invest in yourself, your career, and the financial world. \n Social Finance, LLC seeks Senior Software Engineer in San Francisco, CA: \n Job Duties: Lead the development and testing of system components/services, code and design reviews. Participate in shaping the technical architecture of the product. Help translate product requirements into user stories and technical solutions. Delivery highly available and scalable services in a production environment. Mentor other engineers, support the technical culture, and help grow the team. Generate ideas for new initiatives and technologies. Communicate with project leads, production managers and other software developers. Part-time telecommuting is an option. Hybrid work from Social Finance offices in San Francisco, CA. \n Requirements: Seven (7) years of experience in the job offered or in a related occupation.\n Special Skill Requirements:  \n \n 48 months of experience in: programming experience utilizing a modern technology stack using ReAct, graphQL, springboot, Airflow, Zookeeper, Terraform, Kafka and Flink. \n 48 months of experience in: Java, Kotlin, Python and JavaScript in-memory data structures, SQL and NoSQL database design. \n Service-Oriented Architecture (SOA) or microservice based application environment using Docker, Kubernetes and AWS Fargate. \n  AI terminologies, including Large Language Models (LLM) usage in agent development using AWS Bedrocks and Spring AI. \n  modern CI-CD and collaborative coding environments, including using Git for version control, code reviews, and terraform, AWS CDK for continuous development \n sophisticated testing systems for distributed systems at scale using JUnit. \n Operational excellence practices and using monitoring tools, including DataDog or AWS Cloudwatch. \n driving digital transformation initiatives through large-scale service migrations and tech stack modernization using feature gating mechanisms like Optimizely. \n \n Any suitable combination of education, training and/or experience is acceptable. Part-time telecommuting is an option. Hybrid work from Social Finance offices in San Francisco, CA. \n Salary: $221,187.00 - $243,305.00 per year.   \n Submit resume with references using the apply button on this posting or by email to: Req.# 1014.513.7  at: ATTN: HR, jobadverts@sofi.org .\n  \n  \n #LI-DNI\n Compensation and Benefits \n The base pay range for this role is listed below. Final base pay offer will be determined based on individual factors such as the candidate’s experience, skills, and location. \n  \n To view all of our comprehensive and competitive benefits, visit our  Benefits at SoFi   page!\n SoFi provides equal employment opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion (including religious dress and grooming practices), sex (including pregnancy, childbirth and related medical conditions, breastfeeding, and conditions related to breastfeeding), gender, gender identity, gender expression, national origin, ancestry, age (40 or over), physical or medical disability, medical condition, marital status, registered domestic partner status, sexual orientation, genetic information, military and/or veteran status, or any other basis prohibited by applicable state or federal law. \n The Company hires the best qualified candidate for the job, without regard to protected characteristics. \n Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. \n New York applicants: Notice of Employee Rights \n SoFi is committed to an inclusive culture. As part of this commitment, SoFi offers reasonable accommodations to candidates with physical or mental disabilities. If you need accommodations to participate in the job application or interview process, please let your recruiter know or email accommodations@sofi.com. \n Due to insurance coverage issues, we are unable to accommodate remote work from Hawaii or Alaska at this time. \n Internal Employees \n If you are a current employee, do not apply here - please navigate to our Internal Job Board in Greenhouse to apply to our open roles.","salary_min":221187,"salary_max":243305,"location":"Add ALL locations here","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"senior","tags":["microservices","distributed-systems","llm","api-design","cloud","infrastructure"],"apply_url":"https://sofi.com/careers/job/7978419003?gh_jid=7978419003","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-26T17:01:48Z","expires_at":"2026-09-29T13:48:49.042774Z","created_at":"2026-08-27T13:49:22.930166Z","updated_at":"2026-08-30T13:48:49.170491Z","company_name":"SoFi","company_slug":"sofi","company_logo_url":"https://www.google.com/s2/favicons?domain=sofi.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/8d909b3d-3eac-45d3-8804-379bc82bc41b"},{"id":"479ee1e5-bda5-4253-9664-0645a438ac86","company_id":"5d6de1f6-4d6c-463b-8a2b-a5caeadb97b4","title":"Senior Software Engineer - Airflow Infrastructure, NYC","slug":"senior-software-engineer-airflow-infrastructure-nyc-7ce717b7","description":"Astronomer empowers data teams to bring mission-critical software, analytics, and AI to life and is the company behind Astro, the industry-leading unified DataOps platform powered by Apache Airflow®. Astro accelerates building reliable data products that unlock insights, unleash AI value, and powers data-driven applications. Trusted by more than 800 of the world's leading enterprises, Astronomer lets businesses do more with their data. To learn more, visit  www.astronomer.io http://www.astronomer.io.\n\n\nABOUT THIS ROLE:\n\nAt Astronomer, we’re redefining how companies run Apache Airflow at scale. Our R\u0026D organization is home to some of the most innovative minds in cloud infrastructure and open-source software. \n\nWe’re looking for a Senior Software Engineer to join our Airflow Infra team, part of Astro, our flagship cloud platform. You’ll be building the critical layer that connects the open-source Airflow ecosystem to enterprise-grade, massively scalable cloud infrastructure. Your work will directly influence how global organizations orchestrate data pipelines at scale—making them faster, more reliable, and easier to manage.\n\nIf you’re driven by impact, excited by scale, and ready to work on the kind of infrastructure challenges that push the boundaries of what’s possible in cloud-native systems, this is the opportunity you’ve been waiting for.\n\n\n\nHybrid Work Model: For this role, you will embrace a flexible hybrid work model with at least 3 days per week in our New York City office.\n\n\n\n\nWHAT YOU GET TO DO:\n\n - Engineer backend services with high quality, maintainable and well tested code.\n\n - Partner with other engineers, product, customer reliability support, and leadership to achieve business goals and define how our systems should evolve.\n\n - Regularly engage in code reviews and provide constructive feedback.\n\n - Optimize the performance, reliability and scalability of existing backend services.\n\n - Investigate, prototype and propose ideas to improve user experience.\n\n - Create and maintain technical documentation for systems and processes, ensuring clarity and accessibility.\n\n - Participate in on-call rotation, troubleshoot and debug to solve incidents.\n\n\n\n\nWHAT YOU BRING TO THE ROLE:\n\n - 5+ years of experience building and delivering SaaS products.\n\n - Strong proficiency in Python or Golang.\n\n - Hands-on experience with Kubernetes.\n\n - Solid understanding of and experience with integrating with RESTful APIs and distributed systems.\n\n - Comfortable with testing frameworks, such as pytest.\n\n - Strong communication skills, both written and verbal, with experience in creating technical specifications.\n\n - A passion for reliability and operational excellence.\n\n - Ability to scope work and coordinate cross-functionally to address risks and ensure successful delivery.\n\n - Experience with software development best practices, such as code reviews, testing, CI/CD, version control, automation and debugging.\n\n - Ability to adjust to change and rapid pace of development.\n\n - Proactive approach to identifying and addressing issues, with a focus on ownership and accountability.\n\n\n\n\nBONUS POINTS IF YOU HAVE:\n\n - Experience with Apache Airflow\n\n\n\nThe estimated salary for this role ranges from $210,000 - $250,000 based on leveling and geography, along with an equity component and a comprehensive benefits package. This range is merely an estimate; actual compensation may deviate from this range based on skills, experience, and qualifications.\n\n\n\n#LI-Fulltime\n\n#LI-Hybrid\n\n\n\nAt Astronomer, we value diversity. We are an equal opportunity employer: we do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.","salary_min":210000,"salary_max":250000,"location":"New York, NY","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"senior","tags":["cloud","data-pipeline","distributed-systems","infrastructure"],"apply_url":"https://jobs.ashbyhq.com/astronomer/c02288d2-e50a-4ef0-8151-c5d3f6c333af/application","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-26T08:14:31.454Z","expires_at":"2026-09-29T13:47:15.068812Z","created_at":"2026-08-26T13:47:11.32077Z","updated_at":"2026-08-30T13:47:15.197828Z","company_name":"Astronomer","company_slug":"astronomer","company_logo_url":"https://www.google.com/s2/favicons?domain=astronomer.io\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/479ee1e5-bda5-4253-9664-0645a438ac86"},{"id":"a9fbcc14-8316-489c-8e20-4e4d06b4078b","company_id":"776e5e7d-beba-4889-a481-6d9d7c3af325","title":"Senior Software Engineer, Data Infrastructure","slug":"senior-software-engineer-data-infrastructure-e4d3cf7e","description":"About Decagon\n\nDecagon is the leading conversational AI platform empowering every brand to deliver concierge customer experiences.\n\nOur technology enables industry-defining enterprises like Avis Budget Group, Block’s Cash App and Square, Chime, Oura Health, and Hunter Douglas to deploy AI agents that power personalized, deeply satisfying interactions across voice, chat, email, SMS, and every other channel.\n\nWe’re building a future where customer experiences are being redefined from support tickets and hold music to faster resolutions, richer conversations, and deeper relationships. We’re proud to be backed by world-class investors who share that vision, including a16z, Accel, Bain Capital Ventures, Coatue, and Index Ventures, along with many others.\n\nWe’re an in-office company, driven by a shared commitment to excellence and velocity. Our values — Just Get It Done, Invent What Customers Want, Winner’s Mindset, and The Polymath Principle — shape how we work and grow as a team.\n\n\n\n\nABOUT THE TEAM\n\nThe Infrastructure team builds and operates the foundations that power Decagon: networking, data, ML serving, developer platform, and real‑time voice. We partner closely with product, data, and ML to deliver high‑scale, low‑latency systems with clear SLOs and great developer ergonomics.\n\nWe organize around four focus areas:\n\n - Core Infra: The foundational cloud stack—networking, compute, storage, security, and infrastructure‑as‑code—to ensure reliability, scale, and cost efficiency.\n\n - Data Infra: Streaming/batch data platforms powering analytics/BI and customer‑facing telemetry, including for customer‑managed and on‑prem environments.\n\n - ML Infra: GPU and model‑serving platforms for LLM inference with multi‑provider routing and support for on‑prem/air‑gapped deployments.\n\n - Platform (DevEx): CI/CD, paved paths, and core services that make shipping fast, safe, and consistent across teams.\n\nOur mission is to deliver magical support experiences — AI agents working alongside humans to resolve issues quickly and accurately.\n\n \n\nAbout the Role\nWe're hiring a Senior Data Infrastructure Engineer to design, build, and operate the data systems that power Decagon's AI products. You'll own critical data pipelines and storage layers end‑to‑end, improve reliability and performance, and create paved paths that let every Decagon engineer work confidently with data at scale.\n\nIn this role, you will\n\n - Design and implement high‑throughput data pipelines and streaming systems with strong SLOs, clear runbooks, and actionable telemetry.\n\n - Build and operate real‑time and batch ingestion infrastructure using tools like Kafka, Flink, and Airflow.\n\n - Own our analytical data layer — schema design, query performance, and cost optimization across ClickHouse, BigQuery, or similar.\n\n - Partner with research and product teams to architect data solutions, evaluate performance, and scale new features.\n\n - Tune pipeline and query latencies: optimize data paths, apply smart caching/partitioning, and hit tight p95/p99 targets.\n\n - Lead infrastructure‑as‑code (Terraform) and GitOps practices for data systems; reduce drift with reusable modules and policy‑as‑code.\n\n - Participate in on‑call and drive down toil through automation and elimination of recurring data issues.\n\nYour background looks something like this\n\n - 5+ years building and operating production data infrastructure at scale.\n\n - Hands-on experience with Tier 1 data technologies: ClickHouse, Kafka (or MSK/Pub‑Sub/RabbitMQ), and Flink or dbt.\n\n - Proven track record meeting high availability and low latency targets across streaming and batch workloads.\n\n - Excellent observability chops (OpenTelemetry, Prometheus/Grafana, Datadog) and strong incident response discipline.\n\n - Clear written communication and the ability to turn ambiguous data requirements into simple, reliable designs.\n\nEven better if you have\n\n - Experience with CDC tooling (Debezium) and orchestration frameworks (Airflow, Dagster, or Prefect)\n\n - Familiarity with Spark or Dask for large‑scale data processing\n\n - Experience with cloud data warehouses (Snowflake, BigQuery, Redshift, Databricks)\n\n - Experience being an early data/platform/infrastructure engineer at another company\n\n - Strong Kubernetes experience (GKE/EKS/AKS) and multi‑cloud exposure (GCP, AWS, Azure)\n\n - Experience with customer‑managed deployments\n\n \n\n\nCOMPENSATION\n\n$200K – $400K + Offers Equity\n\nThis range reflects the expected compensation for this role. Compensation within the range is determined based on experience, skills, and the scope of responsibilities, with flexibility for candidates who demonstrate exceptional impact. \n\nIn addition to base salary, we offer competitive equity. Final compensation may vary based on location within the United States.\n\n\n\nBenefits\n\nWe proudly offer the following benefits for our full-time employees:\n\n - Medical, Dental, and Vision benefit","salary_min":200000,"salary_max":400000,"location":"San Francisco, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"senior","tags":["agents","llm","data-pipeline","infrastructure"],"apply_url":"https://jobs.ashbyhq.com/decagon/97f14fcf-d8ff-41b2-9108-d35ddd9f595e/application","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-24T23:01:42.116Z","expires_at":"2026-09-29T13:37:44.094723Z","created_at":"2026-08-25T18:28:23.55152Z","updated_at":"2026-08-30T13:37:44.226906Z","company_name":"Decagon","company_slug":"decagon","company_logo_url":"https://www.google.com/s2/favicons?domain=decagon.ai\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/a9fbcc14-8316-489c-8e20-4e4d06b4078b"},{"id":"22350dcd-2163-4b68-95ae-b6b30ca73990","company_id":"a0000000-0000-0000-0000-000000000001","title":"Staff + Sr. Software Engineer, Scaling","slug":"staff-sr-software-engineer-scaling-cf5b270f","description":"About Anthropic \n Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.\n About the role\n Our Inference team is responsible for building and scaling the critical systems that serve Claude to millions of users worldwide. We bring Claude to life by serving our models via the industry’s largest compute-agnostic inference deployments. We are responsible for the entire stack from intelligent request routing to fleet-wide orchestration across diverse AI accelerators.\n The team has a dual mandate: maximizing compute efficiency to reliably serve our explosive customer growth, while enabling breakthrough research by giving our scientists the high-performance inference infrastructure they need to develop next-generation models. We tackle complex, distributed systems challenges across multiple accelerator families and emerging AI hardware running in multiple cloud platforms.\n Inference systems are highly performance sensitive distributed systems. Inference serves hundreds of thousands of customers every day, and the size \u0026 span of the inference fleet requires sophisticated routing, scaling, and networking systems.\n Key responsibilities\n \n Design, build, and maintain the distributed systems that serve Claude to millions of users worldwide\n Develop resilient, flexible systems that adapt in real time to real world events\n Develop intelligent request routing, load balancing, and traffic management systems across thousands of accelerators and multiple cloud providers\n Maximize compute efficiency and optimize cost across the fleet by autoscaling and orchestrating production, research, and experimental workloads across multiple cloud providers\n Build and operate production-grade deployment pipelines for releasing new models to users\n Provide high-performance inference infrastructure that enables researchers to develop next-generation models\n Integrate new AI accelerator platforms and support inference for new model architectures\n \n Minimum qualifications\n \n Significant software engineering experience, particularly with distributed systems\n Results-oriented, with a bias towards flexibility and impact\n Willingness to pick up slack, even if it goes outside your job description\n Desire to learn more about machine learning systems and infrastructure\n Thrive in environments where technical excellence directly drives both business results and research breakthroughs\n Care about the societal impacts of your work\n \n Preferred qualifications\n \n Experience with high-performance, large-scale distributed systems\n Experience implementing and deploying machine learning systems at scale\n Experience with load balancing, request routing, or traffic management systems\n Familiarity with LLM inference optimization, batching, and caching strategies\n Experience with Kubernetes and cloud infrastructure (AWS, GCP, Azure)\n Proficiency in Python or Rust\n \n Representative projects\n \n Designing intelligent routing algorithms that optimize request distribution across many accelerators in different environments\n Autoscaling our compute fleet to dynamically match supply with demand across production, research, and experimental workloads\n Building production-grade deployment pipelines for releasing new models to millions of users reliably\n Contributing to new inference features\n Supporting inference for new model architectures\n Analyzing observability data to tune performance based on real-world production workloads\n Managing multi-region deployments and geographic routing for global customers\n The annual compensation range for this role is listed below. \n For sales roles, the range provided is the role’s On Target Earnings (\"OTE\") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role.\n Annual Salary:\n $320,000 — $485,000 USD \n Logistics \n Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience\n Required field of study:  A field relevant to the role as demonstrated through coursework, training, or professional experience\n Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position\n Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices.\n Visa sponsorship:  We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this.\n We encourage you to apply even if you do not believe you meet every single qua","salary_min":320000,"salary_max":485000,"location":"San Francisco, CA","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["alignment","distributed-systems","llm","cloud","infrastructure"],"apply_url":"https://job-boards.greenhouse.io/anthropic/jobs/5400012008","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-24T19:24:05Z","expires_at":"2026-09-29T13:30:43.237516Z","created_at":"2026-08-25T18:26:22.790216Z","updated_at":"2026-08-30T13:30:43.378784Z","company_name":"Anthropic","company_slug":"anthropic","company_logo_url":"https://www.google.com/s2/favicons?domain=anthropic.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/22350dcd-2163-4b68-95ae-b6b30ca73990"},{"id":"433c380c-4e0a-4c7c-a64c-0f0d3bc9f75f","company_id":"e8c9f3a5-9310-43f5-9341-321fe6d93a92","title":"Engineering Manager, Cloud Infrastructure ","slug":"engineering-manager-cloud-infrastructure-927fa38a","description":"About us    \n Founded in 2017, Wayve is the leading developer of Embodied AI technology.  Our advanced AI software and foundation models enable vehicles to perceive, understand, and navigate any complex environment, enhancing the usability and safety of automated driving systems.\n Our vision is to create autonomy that propels the world forward.  Our intelligent, mapless, and hardware-agnostic AI products are designed for automakers, accelerating the transition from assisted to automated driving.  In our fast-paced environment big problems ignite us—we embrace uncertainty, leaning into complex challenges to unlock groundbreaking solutions. We aim high and stay humble in our pursuit of excellence, constantly learning and evolving as we pave the way for a smarter, safer future.\n At Wayve, your contributions matter.  We value diversity, embrace new perspectives, and foster an inclusive work environment; we back each other to deliver impact.  \n Make Wayve the experience that defines your career!  \n The role \n Lead a seasoned Platform Engineering team with visibility across the entire engineering organization. This team owns the multi-cloud foundation that every other engineering team builds on, so your decisions and your team's work are felt company-wide. You'll shape how we scale that foundation while keeping quality, cost, and security front and center as the business grows.\n Key responsibilities \n \n Manage a team of Platform Engineers, coaching them to drive strong impact.\n Drive the technical roadmap for cloud infrastructure with the tech lead.\n Own cost visibility and efficiency as cloud usage scales.\n Partner with Security on governance, access, and compliance.\n Balance resourcing against business needs and individual capacity.\n Keep prioritization sharp and processes lean.\n Partner with leadership to sustain a culture of collaboration, impact, and innovation.\n Hire and grow the team, anticipating business needs and pitching investment to leadership.\n Contribute to the day-to-day running of the Software team’s operations and larger collaborative efforts.\n \n About you \n Wayve is scaling fast, and we're looking for someone who's passionate about laying the cornerstone: building the cloud infrastructure foundations and best practices that will support the company for years to come. Here's what will help you thrive as our Engineering Manager, Cloud Infrastructure.\n Essential \n \n 3+ years of experience as an Engineering Manager in software or developer operations.\n Experience balancing project work with support and operational load.\n Passion for developing individual team members.\n Strong roadmap planning, stakeholder management, and cross-team alignment.\n Hands-on experience with a major cloud provider (Azure, AWS, or GCP) at scale.\n Knowledge of infrastructure-as-code (Terraform or similar).\n Familiarity with Linux and container orchestration (Kubernetes or similar).\n BS, MS, or PhD in Computer Science, Engineering, or equivalent experience.\n \n Desirable \n \n Experience operating a multi-cloud environment.\n Knowledge of build systems and modern CI/CD solutions.\n Experience developing in Python.\n \n This is a full-time role based in our office in Sunnyvale and the reasonably estimated salary for this role ranges from $ 276,100 to $ 311,400 , plus a competitive equity package. Actual compensation is based on the candidate's skills, qualifications, and experience.  \n At Wayve we want the best of all worlds so we operate a hybrid working policy that combines time together in our offices and workshops to fuel innovation, culture, relationships and learning, and time spent working from home. We operate core working hours so you can determine the schedule that works best for you and your team.  \n Wayve is committed to creating an inclusive interview experience. If you require any accommodations or adjustments to participate fully in our interview process, please let us know. \n We understand that everyone has a unique set of skills and experiences and that not everyone will meet all of the requirements listed above. If you’re passionate about self-driving cars and think you have what it takes to make a positive impact on the world, we encourage you to apply. At Wayve we're committed to creating a diverse, fair and respectful culture that is inclusive of everyone based on their unique skills and perspectives, and regardless of sex, race, religion or belief, ethnic or national origin, disability, age, citizenship, marital, domestic or civil partnership status, sexual orientation, gender identity, veteran status, pregnancy or related condition  (including breastfeeding) or any other basis as protected by applicable law.  \n For more information visit Careers at Wayve.  \n To learn more about what drives us, visit Values at Wayve  \n For US candidates only, please visit E-Verify Notice and Participation and Right to Work \n \n DISCLAIMER: We will not ask about marriage or pregnancy, care responsibilities ","salary_min":276100,"salary_max":311400,"location":"Sunnyvale, CA","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["autonomous-vehicles","generative-ai","cloud","infrastructure"],"apply_url":"https://wayve.firststage.co/jobs?gh_jid=8735504002","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-21T18:02:29Z","expires_at":"2026-09-29T13:43:23.01688Z","created_at":"2026-08-25T18:31:14.198023Z","updated_at":"2026-08-30T13:43:23.144308Z","company_name":"Wayve","company_slug":"wayve","company_logo_url":"https://www.google.com/s2/favicons?domain=wayve.ai\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/433c380c-4e0a-4c7c-a64c-0f0d3bc9f75f"},{"id":"48b0c99b-9952-40ef-a088-d36c43da6f1d","company_id":"31ae48bc-c938-4c26-a348-0bf3c089a446","title":"Senior Software Engineer - AI Infrastructure Performance Insights \u0026 Observability","slug":"senior-software-engineer-ai-infrastructure-performance-insights-observability-bb558088","description":"CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at  www.coreweave.com . \n About this role \n We're looking for a Senior Engineer to be a driving force on CoreWeave's Benchmarking \u0026 Performance team, with a focus on building the performance insights and observability systems that make our AI infrastructure legible at every layer, from individual GPUs and NVLink/InfiniBand fabrics up through distributed training and inference workloads. You will own how we detect, diagnose, and surface performance signals across every data center in our global infrastructure, turning billions of raw telemetry events into real-time insight that engineers, product teams, and executives can act on with confidence.\n This is not a straightforward data engineering or BI role. The role will focus on building the observability and insight tooling that lets us answer, in near real time, whether a GPU fleet, a fabric, or a training run is performing the way it should, and why it isn't when it's not. If you're energized by building the systems that turn raw infrastructure telemetry into trusted, actionable performance intelligence, and you want that work to sit closer to the hardware and the workload than to a dashboard, this role was built for you.\n What you'll do \n \n Performance Insights \u0026 Observability - Design and build the systems that continuously assess AI infrastructure health and performance: GPU utilization and efficiency, interconnect (NVLink, InfiniBand, RoCE) fabric behavior, distributed training and inference throughput, and hardware degradation signals. Build the detection and diagnosis logic that surfaces anomalies and regressions before they become incidents, not just dashboards that report on them after the fact.\n Time-Series \u0026 Metrics Infrastructure - Own and extend our time-series database (TSDB) layer as the backbone of real-time observability. Write and optimize PromQL/MetricsQL queries that power alerting, anomaly detection, and trend analysis across thousands of GPUs and hundreds of benchmark runs. Bridge streaming metrics and batch-analytical workloads so engineers get sub-second answers during live incidents and analysts get complete historical context for root cause work.\n Fabric \u0026 GPU Telemetry - Build and validate the pipelines and metrics that make network fabric and GPU-level behavior observable and comparable across racks, clusters, and hardware generations, including gray failure detection, congestion and error-rate signals, and health scoring that holds up under audit.\n Data Lake Architecture (in support of insight work) - Design and build the performance data lake that underpins the above: table formats (Apache Iceberg, Parquet, Avro), hot/cold tiering, and schema evolution for latency distributions, throughput metrics, GPU utilization, cost-per-token, and hardware health signals. This is foundational infrastructure, not the end product.\n Query Optimization \u0026 Performance - Profile and tune query engines against columnar and time-series stores so that the observability layer meets its own strict P99 latency and freshness SLAs. Benchmark the benchmarking infrastructure itself.\n BI \u0026 Reporting (secondary) - Where needed, build self-service views (Grafana, Looker, or similar) for engineers, product managers, and executives, but as a downstream output of the insight and observability work above, not the primary deliverable.\n \n Who you are \n \n 5+ years of experience building distributed systems, observability platforms, or performance engineering tooling, ideally for infrastructure or ML systems rather than general-purpose BI.\n Strong coding in Python or Go (C++ a plus) and deep familiarity with networked systems, GPU infrastructure, and performance analysis.\n Hands-on experience with Kubernetes at production scale, CI/CD, and observability stacks (Prometheus, Grafana, OpenTelemetry) used to monitor and diagnose infrastructure, not just report on it.\n Working knowledge of time-series databases and fluency in PromQL or MetricsQL for building real-time alerting and anomaly detection, not only historical dashboards.\n Familiarity with data lake architectures and modern table formats (Iceberg, Parquet, Avro) sufficient to support an insights platform, though this is not the primary skill this role is hiring for.\n Comfortable working close to hardware and workload behavior: GPU utilization patterns, interconnect fabric health, distributed training/inference performance characteristics.\n Strong communicator comfortable collaborating with cross-functional teams ","salary_min":182000,"salary_max":242000,"location":"Sunnyvale, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"senior","tags":["data-pipeline","pytorch","llm","gpu","distributed-systems","infrastructure"],"apply_url":"https://coreweave.com/careers/job?4702966006\u0026board=coreweave\u0026gh_jid=4702966006","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-13T20:47:10Z","expires_at":"2026-09-29T13:35:28.469364Z","created_at":"2026-08-25T18:27:34.745Z","updated_at":"2026-08-30T13:35:28.610554Z","company_name":"CoreWeave","company_slug":"coreweave","company_logo_url":"https://www.google.com/s2/favicons?domain=coreweave.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/48b0c99b-9952-40ef-a088-d36c43da6f1d"},{"id":"8cf9f724-5543-42e5-8ec7-6e485eeeb0a4","company_id":"72014eb6-e84d-48c2-af5c-5424ebec0b3c","title":"Senior Machine Learning Infrastructure Engineer, Embedding Platform","slug":"senior-machine-learning-infrastructure-engineer-embedding-platform-6b3a54da","description":"Reddit is a community of communities. It’s built on shared interests, passion, and trust, and is home to the most open and authentic conversations on the internet. Every day, Reddit users submit, vote, and comment on the topics they care most about. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet’s largest sources of information. For more information, visit www.redditinc.com .\n The LS Embedding Machine Learning Platform team is at the forefront of building highly expressive, machine learning models that power Reddit’s recommendation systems. We go beyond standard retrieval and ranking architectures, leveraging modern deep learning approaches and scalable model designs to enhance personalization across Reddit’s ecosystem. Our work impacts content discovery, user engagement, and platform growth at a massive scale.\n About the Role \n As a Senior Machine Learning Infrastructure Engineer , you will work across both model development and ML platform to build large-scale learning systems that improve recommendation and personalization on Reddit. At the senior level, you will own major technical components end to end: designing models, implementing training and evaluation pipelines, and driving production deployment in close partnership with ML platform, product, and cross-functional ML teams.\n Responsibilities \n \n Design, train, and improve large-scale machine learning platforms for recommendation or personalization systems.\n Own and deliver major ML systems components end to end, from problem framing through production rollout.\n Build and optimize end-to-end ML pipelines spanning data preparation, feature generation, training, evaluation, and deployment.\n Improve distributed training, model efficiency, and online inference performance.\n Apply modern modeling approaches including sequence modeling and related foundation-model techniques to Reddit use cases.\n Develop reliable serving and monitoring patterns for low-latency, high-throughput production ML systems.\n Work with cross-functional partners across product, relevance, ads, and core ML teams to deliver measurable improvements in user experience and business impact.\n Drive rigorous offline and online evaluation, including experimentation, model diagnostics, and feedback-loop improvement.\n Contribute to engineering quality through strong code, design reviews, documentation, and operational excellence.\n \n Qualifications \n \n 5+ years of experience in machine learning engineering, with a strong focus on large-scale ML infrastructure and recommendation or personalization systems.\n Expertise in modern deep learning architectures, including sequence models and foundational models.\n Experience building or scaling ML platform for large datasets and high-traffic production environments.\n Demonstrated ability to independently scope and execute ambiguous technical work, while owning high-quality implementation details.\n Solid understanding of distributed training and inference concepts, such as data parallelism, model parallelism, pipeline parallelism, or related optimization techniques.\n Proficiency in Python and experience with modern ML frameworks such as PyTorch, TensorFlow, or similar.\n Strong software engineering fundamentals, including system design, debugging, testing, and performance optimization.\n Experience with A/B testing, model evaluation frameworks, and real-time feedback loops in large-scale production systems.\n Excellent communication skills, with the ability to effectively present complex ML concepts to technical and non-technical stakeholders.\n \n Benefits: \n \n Comprehensive Healthcare Benefits and Income Replacement Programs\n 401k with Employer Match\n Global Benefit programs that fit your lifestyle, from workspace to professional development to caregiving support\n Family Planning Support\n Gender-Affirming Care\n Mental Health \u0026 Coaching Benefits\n Flexible Vacation \u0026 Paid Volunteer Time Off\n Generous Paid Parental Leave \n \n #LI-Remote\n Pay Transparency: \n This job posting may span more than one career level.\n In addition to base salary, this job is eligible to receive equity in the form of restricted stock units, and depending on the position offered, it may also be eligible to receive a commission. Additionally, Reddit offers a wide range of benefits to U.S.-based employees, including medical, dental, and vision insurance, 401(k) program with employer match, generous time off for vacation, and parental leave. To learn more, please visit https://www.redditinc.com/careers/ .\n To provide greater transparency to candidates, we share base salary ranges for all US-based job postings regardless of state. We set standard base pay ranges for all roles based on function, level, and country location, benchmarked against similar stage growth companies. Final offer amounts are determined by multiple factors including, skills, depth of work experience and relevant licenses/credentials, ","salary_min":190800,"salary_max":267100,"location":"Remote (US)","workplace":"remote","remote_scope":"restricted","job_type":"full-time","experience_level":"senior","tags":["healthcare","pytorch","deep-learning","distributed-systems","tensorflow","infrastructure","machine-learning"],"apply_url":"https://job-boards.greenhouse.io/reddit/jobs/8127022","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-12T23:02:49Z","expires_at":"2026-09-29T13:38:58.366394Z","created_at":"2026-08-25T18:28:56.636852Z","updated_at":"2026-08-30T13:38:58.502511Z","company_name":"Reddit","company_slug":"reddit","company_logo_url":"https://www.google.com/s2/favicons?domain=www.reddit.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/8cf9f724-5543-42e5-8ec7-6e485eeeb0a4"},{"id":"eaf53091-fb47-41f8-8b35-264e2d3c214d","company_id":"7f070aa1-7d20-4bbe-b6d2-68769923074e","title":"Senior Staff Software Engineer, DC Infrastructure","slug":"senior-staff-software-engineer-dc-infrastructure-b14db8aa","description":"Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.\n\n\n\nWe're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.\n\n\n\nWe're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.\n\n\n\nIf you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.\n\n\n\nCrusoe’s Data Center Infrastructure Engineering (DCIE) team is fundamental to our mission of providing AI hardware and infrastructure as a service. The team provides infrastructure for Crusoe’s fleet GPU’s and data center. The team sits at the nexus of high performance computing and AI infrastructure as a service.\n\nThe DCIE team owns, deployment maintenance, observability, critical environment, and automation. The team builds, and maintains GPU clusters, develops automation for logical and physical maintenance, provision systems, and observability tooling.\n\n\n\n\nABOUT THE ROLE:\n\nWe are seeking a highly skilled and motivated Software Engineer to join Crusoe’s Data Center Infrastructure Engineering team. This position is focused on the development of software for the management of a fleet of GPU servers as well as the data centers that house those systems. The role focuses on the developing and implementing  advanced diagnostic, observability, automation and repair tooling for high-performance GPU compute clusters.\n\nThe ideal new team member will be a hands-on problem solver who is comfortable working independently. The new team member will play a critical role in maintaining the health and scalability of Crusoe’s rapidly growing GPU fleet.\n\n\n\n\nWHAT YOU’LL BE DOING:\n\n - Developing and implementing deep-level diagnostics and troubleshooting of hardware faults within GPU racks and high-density compute systems.\n\n - Developing troubleshooting and automation tooling for GPU platforms including NVIDIA A100, H200, GB200, B200 and AMD 350X / 355X.\n\n - Developing automation and AI agents for executing component-level diagnosis and remediation for failed or degraded hardware.\n\n - In conjunction with data center operations develop innovative tooling and AI agents for managing the critical environment.\n\n - Developing tooling for post-repair validation and testing tools such as burn-in, Pytorch, and NVIDIA NCCL to ensure system stability and performance.\n\n - Own the deployment, monitoring, and operational support of developed tooling, ensuring solutions maximize GPU fleet availability and performance to drive customer success.\n\n - Developing automation and operational tooling for facilities management power as well as direct liquid cooling hardware systems\n\n\n\n\nWHAT YOU’LL BRING TO THE TEAM:\n\n - Software engineering experience.\n\n - The ability to identify a problem, rapidly develop a scalable solution and ship it.\n\n - Ability to lean in and assist team members working on critical or complex technical initiatives.\n\n - Ability to set the technical direction for a specific project and execute.\n\n - Expertise in distributed systems, reliability, and cloud platforms (Kubernetes, IaC, GCP etc.)\n\n - Strength in at least one programming language - Go, Python, Java, Rust.\n\n - Strong analytical and problem-solving skills.\n\n - Excellent communication and collaboration skills.\n\n - Ability to work independently and within a team\n\n\n\n\nNICE TO HAVE:\n\n - Experience with Temporal and Kubernetes.\n\n - Experience working directly with hardware vendors.\n\n - Background in large-scale GPU fleet operations or hyperscale data center environments.\n\n\n\n\nBENEFITS:\n\n - Industry competitive pay\n\n - Restricted Stock Units in a fast growing, well-funded technology company\n\n - Health insurance package options that include HDHP and PPO, vision, and dental for you and your dependents\n\n - Employer contributions to HSA accounts\n\n - Paid Parental Leave\n\n - Paid life insurance, short-term and long-term disability\n\n - Teladoc\n\n - 401(k) with a 100% match up to 4% of salary\n\n - Generous paid time off and holiday schedule\n\n - Cell phone reimbursement\n\n - Tuition reimbursement\n\n - Subscription to the Calm app\n\n - MetLife Legal\n\n - Company paid commuter benefit; $300 per month\n\n\n\nCompe","salary_min":250000,"salary_max":300000,"location":"San Francisco, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["distributed-systems","cloud","data-pipeline","pytorch","gpu","agents","infrastructure"],"apply_url":"https://jobs.ashbyhq.com/crusoe/d0e72f9f-37af-4391-98cd-4e187cd224ae/application","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-12T22:51:19.788Z","expires_at":"2026-09-29T13:35:58.784823Z","created_at":"2026-08-25T18:27:46.34593Z","updated_at":"2026-08-30T13:35:58.918777Z","company_name":"Crusoe","company_slug":"crusoe","company_logo_url":"https://www.google.com/s2/favicons?domain=crusoe.ai\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/eaf53091-fb47-41f8-8b35-264e2d3c214d"},{"id":"7f77c174-e767-4430-a27f-1c0da9eace2a","company_id":"7f070aa1-7d20-4bbe-b6d2-68769923074e","title":"Staff Software Engineer, DC Infrastructure","slug":"staff-software-engineer-7663b768","description":"Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.\n\n\n\nWe're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.\n\n\n\nWe're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.\n\n\n\nIf you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.\n\n\n\nCrusoe’s Data Center Infrastructure Engineering (DCIE) team is fundamental to our mission of providing AI hardware and infrastructure as a service. The team provides infrastructure for Crusoe’s fleet GPU’s and data center. The team sits at the nexus of high performance computing and AI infrastructure as a service.\n\nThe DCIE team owns, deployment maintenance, observability, critical environment, and automation. The team builds, and maintains GPU clusters, develops automation for logical and physical maintenance, provision systems, and observability tooling.\n\n\n\n\nABOUT THE ROLE:\n\nWe are seeking a highly skilled and motivated Software Engineer to join Crusoe’s Data Center Infrastructure Engineering team. This position is focused on the development of software for the management of a fleet of GPU servers as well as the data centers that house those systems. The role focuses on the developing and implementing  advanced diagnostic, observability, automation and repair tooling for high-performance GPU compute clusters.\n\nThe ideal new team member will be a hands-on problem solver who is comfortable working independently. The new team member will play a critical role in maintaining the health and scalability of Crusoe’s rapidly growing GPU fleet.\n\n\n\n\nWHAT YOU’LL BE DOING:\n\n - Developing and implementing deep-level diagnostics and troubleshooting of hardware faults within GPU racks and high-density compute systems.\n\n - Developing troubleshooting and automation tooling for GPU platforms including NVIDIA A100, H200, GB200, B200 and AMD 350X / 355X.\n\n - Developing automation and AI agents for executing component-level diagnosis and remediation for failed or degraded hardware.\n\n - In conjunction with data center operations develop innovative tooling and AI agents for managing the critical environment.\n\n - Developing tooling for post-repair validation and testing tools such as burn-in, Pytorch, and NVIDIA NCCL to ensure system stability and performance.\n\n - Own the deployment, monitoring, and operational support of developed tooling, ensuring solutions maximize GPU fleet availability and performance to drive customer success.\n\n - Developing automation and operational tooling for facilities management power as well as direct liquid cooling hardware systems\n\n\n\n\nWHAT YOU’LL BRING TO THE TEAM:\n\n - Software engineering experience.\n\n - The ability to identify a problem, rapidly develop a scalable solution and ship it.\n\n - Ability to lean in and assist team members working on critical or complex technical initiatives.\n\n - Ability to set the technical direction for a specific project and execute.\n\n - Expertise in distributed systems, reliability, and cloud platforms (Kubernetes, IaC, GCP etc.)\n\n - Strength in at least one programming language - Go, Python, Java, Rust.\n\n - Strong analytical and problem-solving skills.\n\n - Excellent communication and collaboration skills.\n\n - Ability to work independently and within a team\n\n\n\n\nNICE TO HAVE:\n\n - Experience with Temporal and Kubernetes.\n\n - Experience working directly with hardware vendors.\n\n - Background in large-scale GPU fleet operations or hyperscale data center environments.\n\n\n\n\nBENEFITS:\n\n - Industry competitive pay\n\n - Restricted Stock Units in a fast growing, well-funded technology company\n\n - Health insurance package options that include HDHP and PPO, vision, and dental for you and your dependents\n\n - Employer contributions to HSA accounts\n\n - Paid Parental Leave\n\n - Paid life insurance, short-term and long-term disability\n\n - Teladoc\n\n - 401(k) with a 100% match up to 4% of salary\n\n - Generous paid time off and holiday schedule\n\n - Cell phone reimbursement\n\n - Tuition reimbursement\n\n - Subscription to the Calm app\n\n - MetLife Legal\n\n - Company paid commuter benefit; $300 per month\n\n\n\nCompe","salary_min":215000,"salary_max":260000,"location":"San Francisco, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["agents","distributed-systems","data-pipeline","pytorch","gpu","cloud","infrastructure"],"apply_url":"https://jobs.ashbyhq.com/crusoe/9469308d-d063-4449-bd53-3043836b1a0e/application","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-12T22:50:13.339Z","expires_at":"2026-09-29T13:35:57.940694Z","created_at":"2026-04-13T09:41:26.548071Z","updated_at":"2026-08-30T13:35:58.079785Z","company_name":"Crusoe","company_slug":"crusoe","company_logo_url":"https://www.google.com/s2/favicons?domain=crusoe.ai\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/7f77c174-e767-4430-a27f-1c0da9eace2a"},{"id":"5034cc11-d680-45ce-88f8-98db61a33f70","company_id":"72014eb6-e84d-48c2-af5c-5424ebec0b3c","title":"Staff Machine Learning Infrastructure Engineer, Embedding Platform","slug":"staff-machine-learning-infrastructure-engineer-embedding-platform-4410257e","description":"Reddit is a community of communities. It’s built on shared interests, passion, and trust, and is home to the most open and authentic conversations on the internet. Every day, Reddit users submit, vote, and comment on the topics they care most about. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet’s largest sources of information. For more information, visit www.redditinc.com .\n The LS Embedding Machine Learning Platform team is at the forefront of building highly expressive machine learning models that power Reddit’s recommendation systems. We go beyond standard retrieval and ranking architectures, leveraging modern deep learning approaches and scalable model designs to enhance personalization across Reddit’s ecosystem. Our work impacts content discovery, user engagement, and platform growth at a massive scale.\n How You'll Have Impact \n As a Staff Machine Learning Infrastructure Engineer , you will own the technical direction for large-scale machine learning platform, guiding the development of advanced deep learning architectures and high-impact ML systems. You will partner with leadership to define ML roadmaps, drive innovation in scalable model design and training approaches, and ensure efficient, reliable deployment of ML models in production. This role offers an opportunity to influence key AI-driven systems across Reddit while mentoring and uplifting the team’s technical capabilities.\n What You’ll Do \n \n Architect and lead the development of next-generation, large-scale machine learning techniques.\n Define and execute the ML strategy, identifying opportunities to enhance personalization and recommendation quality across Reddit.\n Lead research initiatives on scalable machine learning systems and real-time model adaptation, bringing cutting-edge advancements into production.\n Partner with ML infrastructure teams to build high-performance, distributed training systems that efficiently scale across multiple GPUs and cloud environments.\n Establish and optimize real-time serving architectures for large-scale embeddings, ensuring low-latency inference and high throughput.\n Collaborate cross-functionally with teams in Feed Ranking, Ads, Content Understanding, and Core ML to integrate ML models into Reddit’s key AI-driven systems.\n Mentor and guide senior and mid-level ML engineers, fostering a culture of excellence, innovation, and knowledge sharing.\n Stay at the forefront of AI research, evaluating and introducing new modeling paradigms to keep Reddit’s ML ecosystem cutting-edge.\n Drive technical discussions, present findings to leadership, and contribute to long-term ML planning and decision-making.\n \n Who You Might Be: \n \n 8+ years of experience in machine learning engineering, with a strong focus on large-scale ML systems and recommendation or personalization systems.\n Expertise in modern deep learning architectures, including sequence models and foundational models.\n Deep understanding of complex multi-entity relationships in machine learning applications and how they are modeled in large-scale systems.\n Proven ability to design, implement, and optimize scalable ML architectures, from distributed training to real-time inference.\n Strong software engineering skills in Python, C++, or similar languages, with experience in ML infrastructure, high-performance computing, and cloud-based ML pipelines.\n Demonstrated leadership in driving ML strategy, mentoring engineers, and influencing cross-functional teams.\n Experience with A/B testing, model evaluation frameworks, and real-time feedback loops in large-scale production systems.\n Excellent communication skills, with the ability to effectively present complex ML concepts to technical and non-technical stakeholders. \n \n Benefits: \n \n Comprehensive Healthcare Benefits and Income Replacement Programs\n 401k with Employer Match\n Global Benefit programs that fit your lifestyle, from workspace to professional development to caregiving support\n Family Planning Support\n Gender-Affirming Care\n Mental Health \u0026 Coaching Benefits\n Flexible Vacation \u0026 Paid Volunteer Time Off\n Generous Paid Parental Leave \n \n #LI-Remote\n Pay Transparency: \n This job posting may span more than one career level.\n In addition to base salary, this job is eligible to receive equity in the form of restricted stock units, and depending on the position offered, it may also be eligible to receive a commission. Additionally, Reddit offers a wide range of benefits to U.S.-based employees, including medical, dental, and vision insurance, 401(k) program with employer match, generous time off for vacation, and parental leave. To learn more, please visit https://www.redditinc.com/careers/ .\n To provide greater transparency to candidates, we share base salary ranges for all US-based job postings regardless of state. We set standard base pay ranges for all roles based on function, level, and country location, benchma","salary_min":253300,"salary_max":354600,"location":"Remote (US)","workplace":"remote","remote_scope":"restricted","job_type":"full-time","experience_level":"lead","tags":["deep-learning","distributed-systems","healthcare","infrastructure","machine-learning"],"apply_url":"https://job-boards.greenhouse.io/reddit/jobs/8126982","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-12T22:33:37Z","expires_at":"2026-09-29T13:39:01.21974Z","created_at":"2026-08-25T18:28:56.767146Z","updated_at":"2026-08-30T13:39:01.355053Z","company_name":"Reddit","company_slug":"reddit","company_logo_url":"https://www.google.com/s2/favicons?domain=www.reddit.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/5034cc11-d680-45ce-88f8-98db61a33f70"},{"id":"7b9b8809-f3ac-4c1b-9808-82ec279c5fb4","company_id":"a0000000-0000-0000-0000-000000000001","title":"Software Engineer, Infrastructure, Interpretability","slug":"software-engineer-infrastructure-interpretability-3cb5ab0f","description":"About Anthropic \n Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.\n About the role:\n When you see what modern language models are capable of, do you wonder, \"How do these things work? How can we trust them?\"\n The Interpretability team at Anthropic works to understand what's actually happening inside trained models - and applies our best techniques to keep frontier AI safe as it rapidly improves.\n Think of us as doing \"neuroscience\" of neural networks using \"microscopes\" we build - or reverse-engineering neural networks like binary programs.\n More resources to learn about our work: \n \n \n Our Research blog - covering advances including Monosemantic Features and Circuits \n \n An Intro to Interpretability from our research lead, Chris Olah \n \n The Urgency of Interpretability from CEO Dario Amodei\n \n Engineering Challenges Scaling Interpretability - directly relevant to this role\n \n 60 Minutes segment - see a demo of tooling our team built\n \n New Yorker article - what it's like to work on one of AI's hardest open problems\n \n This role is an early hire on a new infrastructure effort within Interpretability: you'll help define its charter, not just execute it. \n Interpretability research requires deep access to frontier models while retaining a high degree of research flexibility. Your job is to build the paved path that makes that access secure by default, private by design, and low-friction for every researcher. The work spans four areas:\n \n \n Security : design the secure-by-default environments and access patterns that enable deep model access for an organization whose research requires it - done well, the same design improves both our security posture and research productivity.\n \n Privacy : build data-access patterns that ensure policy adherence as our research moves from theory into practical application\n \n Data \u0026 Compute Management : manage research data at petabyte scale and make efficient use of large accelerator fleets - storage lifecycle, capacity planning, and scheduling.\n \n Developer experience : agentic engineering, tooling and observability that keep researchers moving fast\n \n In this role, you’ll be deeply embedded alongside Interp Researchers to understand their workflows - building your understanding of the research as you go; at the same time you’ll bridge communication with Anthropic’s wider platform and security teams.. Every hour of researcher friction you remove is multiplied across the whole organization, and the infrastructure you build sets the pace at which interpretability results reach real safety decisions.\n Responsibilities:\n \n \n Design, build, and own shared infrastructure for Interpretability - research environments, data systems, and compute tooling that researchers rely on daily\n \n Lead cross-team efforts with our agentic engineering , security, compute, and storage platform teams, so that company-wide solutions serve research needs\n \n Discover and resolve major organization-wide developer experience issues\n \n Help take interpretability methods from research code to dependable audit pipelines\n \n You may be a good fit if you:\n \n \n Are highly proficient in at least one programming language (e.g., Python, Rust, Go, Java) and productive with Python\n \n Have significant experience building and operating secure and scalable software infrastructure - cloud systems, distributed systems, or developer tooling\n \n Have strong cross-functional communication skills - equally at home working with researchers and with platform and security teams\n \n Are extremely curious about unfamiliar domains\n \n Have a strong ability to prioritize the most impactful work and are comfortable operating with ambiguity and questioning assumptions\n \n Are curious about interpretability research and its role in AI safety (though no research experience is required!)\n \n Care about the societal impacts and ethics of your work\n \n Strong candidates may also have:\n \n \n Experience with cloud infrastructure (e.g. GCP or AWS), Kubernetes, networking and infrastructure-as-code\n \n Security engineering experience: identity / auth / access management, sandboxing, red teaming\n \n Experience with data warehousing, large-scale storage systems, and data lifecycle management - especially for research\n \n Experience with compute schedulers and accelerator fleet management\n \n Experience building developer productivity tooling and observability stacks\n \n Experience building tooling to accelerate research teams\n \n Representative Projects:\n \n \n Design and stand up a hardened research environment where researchers experiment directly on frontier model weights\n \n Build lifecycle management for petabytes of research data - visibility, retention, and cost efficiency\n \n B","salary_min":320000,"salary_max":485000,"location":"San Francisco, CA","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"principal","tags":["cloud","security","deep-learning","agents","distributed-systems","alignment","infrastructure","research"],"apply_url":"https://job-boards.greenhouse.io/anthropic/jobs/5388612008","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-12T22:29:15Z","expires_at":"2026-09-29T13:30:35.12076Z","created_at":"2026-08-25T18:26:19.069491Z","updated_at":"2026-08-30T13:30:35.264073Z","company_name":"Anthropic","company_slug":"anthropic","company_logo_url":"https://www.google.com/s2/favicons?domain=anthropic.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/7b9b8809-f3ac-4c1b-9808-82ec279c5fb4"},{"id":"99457528-4174-4282-b22e-cc762e23a5b2","company_id":"698abc6f-9497-4ea6-809f-f0f7c2788a46","title":"Senior Software Engineer, Agentic Infrastructure","slug":"senior-software-engineer-agentic-infrastructure-dbac5e34","description":"At Relativity Space, we’re building rockets to serve today’s needs and tomorrow’s breakthroughs. Our Terran R vehicle will deliver customer payloads to orbit, meeting the growing demand for launch capacity. But that’s just the start. Achieving commercial success with Terran R will unlock new opportunities to advance science, exploration, and innovation, pioneering progress that reaches beyond the known. \n Joining Relativity means becoming part of something where autonomy, ownership, and impact exist at every level. Here, you're not just executing tasks; you're solving problems that haven’t been solved before, helping develop a rocket, a factory, and a business from the ground up. Whether you’re in propulsion, manufacturing, software, avionics, or a corporate function, you’ll collaborate across teams, shape decisions, and see your work come to life in record time. Relativity is a place where creativity and technical rigor go hand in hand, and your voice will help define the stories we’re writing together. Now is a unique moment in time where it’s early enough to leave your mark on the product, the process, and the culture, but far enough along that Terran R is tangible and picking up momentum. The most meaningful work of your career is waiting. Join us. \n About the Team: \n Dark Matter Lab is a research group within Relativity Space focused on advanced aerospace systems, agentic engineering, and technologies outside the conventional roadmap. We build the infrastructure needed to turn new ideas into engineering capabilities people can actually depend on. \n About the Role: \n We run a stack of interconnected agentic systems — SybilClaw, yapCAD, Mechatron, Multigraph, local LLM inference, an inter-agent message bus, parametric CAD pipelines, a print farm, and the infrastructure connecting it all. We need an engineer who can own and evolve these systems as they move from prototypes into production engineering infrastructure. \n This role is for someone who has deployed an agentic harness (OpenClaw, Hermes, SybilClaw, or something comparable) in a real work environment. You understand what it takes to make these systems reliable when engineers depend on them every day. You’ll learn the stack, improve it, and work with forward-deployed engineers to package portions of it for deployment with internal and external customers. \n \n Own the reliability of our agentic infrastructure: agents, sessions, model routing, context pipelines, inter-agent communication, and supporting services \n Build and maintain infrastructure where good off-the-shelf solutions don't yet exist \n Deploy and operate local LLM inference across Mac Studio and GPU hardware \n Manage Linux/macOS systems, Proxmox VMs and containers, storage, backups, and recovery \n Maintain multi-site networking including L3 routing, VLANs, DNS, firewalls, and connectivity \n Build CLIs, dashboards, automation, and internal tools that make engineers faster \n Improve observability, debugging, and automated failure recovery \n Contribute upstream to open-source projects we rely on \n Help design practical security around local compute, data handling, and access controls \n Package and deploy portions of the stack into internal and customer environments \n Work directly with engineers to understand what they need and turn recurring problems into better infrastructure \n \n About You: \n \n 5+ years of experience building and operating complex software and compute infrastructure, with hands-on work across hardware, operating systems, networking, and automation \n Experience deploying an agentic harness such as OpenClaw, Hermes, SybilClaw, or a comparable system in a real work environment \n Strong Linux and macOS systems experience, including debugging services, processes, storage, permissions, and networking \n Experience operating physical infrastructure, VMs, or containers in production \n Strong networking fundamentals including routing, VLANs, DNS, and firewalls \n Ability to build software, scripts, CLIs, and integrations when existing tools aren't sufficient \n Meaningful experience contributing to or maintaining open-source software \n Strong judgment around reliability, security, performance, and simplicity \n \n Nice to haves but not required: \n \n Experience with ITAR-regulated, air-gapped, export-controlled, or similarly constrained environments \n Experience operating local LLM inference with Ollama, MLX, vLLM, llama.cpp, or similar \n Experience with model routing, context management, tool execution, agent state, or multi-agent systems \n Experience building internal developer infrastructure used daily by other engineers \n Experience deploying systems into customer or forward-deployed environments \n A GitHub, Gitea, or other body of work where we can see what you've built \n \n This role requires in-office presence at least three days per week, with flexibility to work remotely when the work allows. Much of this infrastructure is physical, local-first, an","salary_min":154000,"salary_max":230000,"location":"Long Beach, CA","workplace":"remote","remote_scope":"unknown","job_type":"full-time","experience_level":"senior","tags":["llm","fine-tuning","agents","infrastructure"],"apply_url":"https://boards.greenhouse.io/relativity/jobs/8702834002?gh_jid=8702834002","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-12T22:02:57Z","expires_at":"2026-09-29T13:49:00.355275Z","created_at":"2026-08-25T18:33:30.918833Z","updated_at":"2026-08-30T13:49:00.482057Z","company_name":"Relativity","company_slug":"relativity","company_logo_url":"https://www.google.com/s2/favicons?domain=relativity.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/99457528-4174-4282-b22e-cc762e23a5b2"},{"id":"636e7c4b-ab44-4780-a4db-87055c91bcc8","company_id":"698abc6f-9497-4ea6-809f-f0f7c2788a46","title":"Staff Software Engineer, Agentic Infrastructure","slug":"staff-software-engineer-agentic-infrastructure-bf0e25b8","description":"At Relativity Space, we’re building rockets to serve today’s needs and tomorrow’s breakthroughs. Our Terran R vehicle will deliver customer payloads to orbit, meeting the growing demand for launch capacity. But that’s just the start. Achieving commercial success with Terran R will unlock new opportunities to advance science, exploration, and innovation, pioneering progress that reaches beyond the known. \n Joining Relativity means becoming part of something where autonomy, ownership, and impact exist at every level. Here, you're not just executing tasks; you're solving problems that haven’t been solved before, helping develop a rocket, a factory, and a business from the ground up. Whether you’re in propulsion, manufacturing, software, avionics, or a corporate function, you’ll collaborate across teams, shape decisions, and see your work come to life in record time. Relativity is a place where creativity and technical rigor go hand in hand, and your voice will help define the stories we’re writing together. Now is a unique moment in time where it’s early enough to leave your mark on the product, the process, and the culture, but far enough along that Terran R is tangible and picking up momentum. The most meaningful work of your career is waiting. Join us. \n About the Team: \n Dark Matter Lab is a research group within Relativity Space focused on advanced aerospace systems, agentic engineering, and technologies outside the conventional roadmap. We build the infrastructure needed to turn new ideas into engineering capabilities people can actually depend on. \n About the Role: \n We run a stack of interconnected agentic systems — SybilClaw, yapCAD, Mechatron, Multigraph, local LLM inference, an inter-agent message bus, parametric CAD pipelines, a print farm, and the infrastructure connecting it all. We need an engineer who can own and evolve these systems as they move from prototypes into production engineering infrastructure. \n This role is for someone who has deployed an agentic harness (OpenClaw, Hermes, SybilClaw, or something comparable) in a real work environment. You understand what it takes to make these systems reliable when engineers depend on them every day. You’ll learn the stack, improve it, and work with forward-deployed engineers to package portions of it for deployment with internal and external customers. \n \n Own the reliability of our agentic infrastructure: agents, sessions, model routing, context pipelines, inter-agent communication, and supporting services \n Build and maintain infrastructure where good off-the-shelf solutions don't yet exist \n Deploy and operate local LLM inference across Mac Studio and GPU hardware \n Manage Linux/macOS systems, Proxmox VMs and containers, storage, backups, and recovery \n Maintain multi-site networking including L3 routing, VLANs, DNS, firewalls, and connectivity \n Build CLIs, dashboards, automation, and internal tools that make engineers faster \n Improve observability, debugging, and automated failure recovery \n Contribute upstream to open-source projects we rely on \n Help design practical security around local compute, data handling, and access controls \n Package and deploy portions of the stack into internal and customer environments \n Work directly with engineers to understand what they need and turn recurring problems into better infrastructure \n \n About You: \n \n 7+ years of experience building and operating complex software and compute infrastructure, with hands-on work across hardware, operating systems, networking, and automation \n Experience deploying an agentic harness such as OpenClaw, Hermes, SybilClaw, or a comparable system in a real work environment \n Strong Linux and macOS systems experience, including debugging services, processes, storage, permissions, and networking \n Experience operating physical infrastructure, VMs, or containers in production \n Strong networking fundamentals including routing, VLANs, DNS, and firewalls \n Ability to build software, scripts, CLIs, and integrations when existing tools aren't sufficient \n Meaningful experience contributing to or maintaining open-source software \n Strong judgment around reliability, security, performance, and simplicity \n \n Nice to haves but not required: \n \n Experience with ITAR-regulated, air-gapped, export-controlled, or similarly constrained environments \n Experience operating local LLM inference with Ollama, MLX, vLLM, llama.cpp, or similar \n Experience with model routing, context management, tool execution, agent state, or multi-agent systems \n Experience building internal developer infrastructure used daily by other engineers \n Experience deploying systems into customer or forward-deployed environments \n A GitHub, Gitea, or other body of work where we can see what you've built \n \n This role requires in-office presence at least three days per week, with flexibility to work remotely when the work allows. Much of this infrastructure is physical, local-first, an","salary_min":181000,"salary_max":271000,"location":"Long Beach, CA","workplace":"remote","remote_scope":"unknown","job_type":"full-time","experience_level":"lead","tags":["fine-tuning","agents","llm","infrastructure"],"apply_url":"https://boards.greenhouse.io/relativity/jobs/8694874002?gh_jid=8694874002","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-12T22:02:56Z","expires_at":"2026-09-29T13:49:00.447553Z","created_at":"2026-08-25T18:33:30.924053Z","updated_at":"2026-08-30T13:49:00.574243Z","company_name":"Relativity","company_slug":"relativity","company_logo_url":"https://www.google.com/s2/favicons?domain=relativity.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/636e7c4b-ab44-4780-a4db-87055c91bcc8"},{"id":"991f8296-af04-49ea-b15e-7d9b5fc51b29","company_id":"a0000000-0000-0000-0000-000000000001","title":"AI Infrastructure Operations, Demand Planning","slug":"ai-infrastructure-operations-demand-planning-3cf6fb61","description":"About Anthropic \n Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.\n About the Role\n Anthropic runs one of the largest and fastest-growing infrastructure fleets in the industry, across multiple accelerator families, CPU families, clouds, neoclouds, and on-prem sites. Capacity Engineering owns the data, tooling, and systems that let Anthropic plan, measure, and maximize utilization of that fleet: we partner on supply deals, wire telemetry from day zero, own the canonical capacity data layer, and build the planning and enforcement tools every research and product team relies on. This role sits in the Planning pillar, on the Demand Planning team, and works daily with research engineering, pretraining, inference, compute supply, finance, and external vendors.\n You own the tranches. The job has two halves that feed each other. Upstream, you take the Demand Planning forecast and turn it into per-tranche requirements — shape, interconnect, region, supporting resources, date — and carry those into sourcing negotiations and data center build reviews so we contract for capacity we can actually use when we need it. Downstream, you own the integrated schedule and system of record for every tranche in flight — from contracted through reserved, ingested, in-cluster, healthy, and occupied — and you drive the owners of each hop to their dates. Every slip you see downstream becomes a contract-language fix, an automation, or a correction fed back to the forecast.\n What you'll do\n \n Turn the forecast into per-tranche requirements. Take the Demand Planning forecast plus direct input from research, pretraining, and inference planners, and convert it into concrete accelerator, interconnect, region, supporting-resource, and date requirements for each tranche. Represent those in sourcing negotiations and data center build reviews, including which contractual terms actually move delivery dates.\n Qualify tranches for deliverability before signature. The Capacity Planner signs fit-to-forecast; you sign whether the shape can land schedulable, healthy, and instrumented in that region on that date, with storage, egress, identity in place.\n Close the delivery loop. Track forecast-versus-delivered on shape, region, and timing for every tranche; publish the variance; and feed it back to Demand Planning and into the next contract.\n Own the bring-up system of record. Define the canonical contract-to-occupied state machine with explicit entry and exit criteria per stage, and make it a first-class object in the capacity data layer so every downstream tool sees in-flight capacity, not only what has landed.\n Run a portfolio of bring-ups in parallel — new cloud regions, on-prem sites, neocloud blocks — with one integrated schedule spanning provider milestones, cluster creation, network turn-up, storage readiness, health burn-in, and first-workload landing. \n Drive readiness automation: All capacity systems are fully integrated for all new capacity, from contracted through ingested, automated and scaled.\n Instrument and publish the numbers that matter — time-to-occupied and paid-idle dollars per tranche — with executive-level reporting on status, tradeoffs, and risk across the portfolio.\n \n What you bring\n \n Significant experience delivering large-scale infrastructure — cloud regions, accelerator clusters, HPC systems, or bare-metal fleets — at multi-region scale or ≥10k accelerators (or CPU/storage equivalent).\n Technical range from through cluster orchestration and node health, up to the telemetry and planning tables on top — enough to debug where they disagree rather than route it.\n SQL and enough Python to answer your own questions and build your own reporting.\n A degree in a technical field or an equivalent engineering track record.\n \n Preferred\n \n Reserved-capacity onboarding, private offers, or capacity commitments with cloud or neocloud providers.\n Enough demand-planning exposure to challenge a forecast, translate it into per-tranche requirements, and feed delivery variance back into it.\n Data center or colocation delivery: power and space planning, network turn-up, site acceptance, vendor management.\n Accelerator health and burn-in, collective-communications sanity testing, or fleet-health SLOs — and a rigorous definition of \"healthy.\"\n Systems of record or lifecycle services for infrastructure assets.\n Onboarding a new hardware generation into an existing scheduler and observability stack.\n The annual compensation range for this role is listed below. \n For sales roles, the range provided is the role’s On Target Earnings (\"OTE\") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for","salary_min":320000,"salary_max":405000,"location":"San Francisco, CA","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["alignment","pre-training","search","infrastructure"],"apply_url":"https://job-boards.greenhouse.io/anthropic/jobs/5382750008","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-11T12:14:01Z","expires_at":"2026-09-29T13:30:10.669189Z","created_at":"2026-08-25T18:26:10.627699Z","updated_at":"2026-08-30T13:30:10.854518Z","company_name":"Anthropic","company_slug":"anthropic","company_logo_url":"https://www.google.com/s2/favicons?domain=anthropic.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/991f8296-af04-49ea-b15e-7d9b5fc51b29"},{"id":"e2ee6535-e8ad-47aa-80bf-88a13e3a97fc","company_id":"a0000000-0000-0000-0000-000000000001","title":"Staff+ Site Reliability Engineer, Safeguards ML Infra","slug":"staff-site-reliability-engineer-safeguards-ml-infra-ebc3e8d7","description":"About Anthropic \n Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.\n About the role: \n The Safeguards ML Infra team designs, builds, and operates the production infrastructure that powers Claude's safety systems. We own the critical backend services that ensure safety on the token generation path, and we own the operational work of getting those systems safely into production: standing up safeguards for every new model launch, and deploying new safety classifiers as they ship. Every frontier model release runs through this team – we configure, verify, and roll out safeguards across every platform Claude runs on (1P, AWS Bedrock, GCP Vertex, etc.), and we lead incident response when issues arise.\n This role sits at the center of that operational work. You'll ensure safeguards are properly configured and deployed for model launches and own the off-cycle deployment of new safety classifiers — canarying changes, verifying that the right safeguards are provably live on the right models, and holding rollback authority when something looks wrong. Every launch should also shrink the checklist, and the manual verifications should evolve into a system that runs itself. You'll turn launch runbooks into tooling, hand-built checks into continuous validation, and one-off deploys into a repeatable pipeline.\n We're looking for engineers with deep experience in production change management at scale — people who have owned deploy pipelines, config management systems, rollout safety, or launch readiness for systems under real production pressure. Familiarity with ML research or transformer architectures is not required — you will learn that on the job. What we prioritize is production judgment: a track record of shipping changes to critical systems safely, and of automating yourself out of the work you did last quarter.\n What you'll do: \n \n Launch captain model releases: stand up, configure, and verify safeguards for every new model, and serve as the safeguards point of contact in the launch room during release windows.\n Own the off-cycle deployment of new safety classifiers as they ship from research — canarying rollouts, running post-deploy validations, and investigating discrepancies when something looks wrong.\n Verify that the right safeguards are provably live on the right models across every deployment platform (1P, AWS Bedrock, GCP Vertex, etc.), and detect and eliminate configuration drift between them.\n Automate yourself out of last quarter's work: turn launch runbooks into tooling, hand-built checks into continuous validation, and one-off deploys into a repeatable pipeline.\n \n Plan to use Claude aggressively to do this! And be a trailblazer that paves the path for safe agentic operations of safety-critical systems.\n \n Build and maintain a safeguards registry with full provenance — what is running in production, on which model, on which platform, and when and by whom it was deployed.\n Participate in on-call and operational-duty rotations covering service incidents, model provisioning, and time-sensitive research and safety launches.\n \n You may be a good fit if you: \n \n Have owned production change management at scale — deploy pipelines, config management systems, canary analysis — and have strong opinions about what \"verified\" means.\n Have run high-stakes releases: served as a launch captain, incident commander, or release owner for systems where a bad deploy has real consequences, and are energized rather than drained by being in the critical path.\n Have meaningful on-call experience for production systems, including incident response and postmortem-driven improvements — and a track record of turning (and fixing!) postmortem action items into process and tooling changes.\n Have a desire to close the gap where nobody has yet raised their hand, even if it requires manually hand-holding processes until automation and tooling can be built.\n Have hands-on experience deploying and operating on cloud platforms (AWS, GCP) at scale.\n Are proficient in Python; experience with Rust is a plus but not required.\n \n Strong candidates may also have: \n \n 8+ years of industry software engineering or site reliability engineering experience.\n A demonstrated history of reducing operational toil through automation, including transitioning teams from manual deployment processes to self-serve pipelines.\n Experience running launch or production-readiness review processes across multiple teams.\n Familiarity with LLM inference systems and the operational characteristics of transformer-based models.\n The annual compensation range for this role is listed below. \n For sales roles, the range provided is the role’s On Target Earnings (\"OTE\") range, meaning th","salary_min":405000,"salary_max":485000,"location":"San Francisco, CA","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["llm","agents","alignment","cloud","rust","infrastructure","devops"],"apply_url":"https://job-boards.greenhouse.io/anthropic/jobs/5230394008","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-11T00:28:36Z","expires_at":"2026-09-29T13:30:36.924785Z","created_at":"2026-08-25T18:26:19.565568Z","updated_at":"2026-08-30T13:30:37.072954Z","company_name":"Anthropic","company_slug":"anthropic","company_logo_url":"https://www.google.com/s2/favicons?domain=anthropic.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/e2ee6535-e8ad-47aa-80bf-88a13e3a97fc"}],"page":1,"per_page":20,"total":570,"total_is_exact":true,"total_pages":29}
