{"access":{"catalog_url":"https://aidevboard.com/api/v1/catalog","description":"Public read endpoints are open and free. API keys are optional for stable agent identity and keyed hourly throttling.","docs_url":"https://aidevboard.com/docs","employer_pilot_url":"https://aidevboard.com/verified-interview-pilot","mode":"open","register_url":"https://aidevboard.com/api/v1/register"},"candidate_resume_action":{"application_authorized":false,"candidate_charge":0,"endpoint":"https://aidevboard.com/api/v1/candidate/resume-preview","job_id_json_path":"jobs[].id","method":"POST","preview_requires_identity":false,"required_body_fields":["job_id","evidence_bullets"],"requires_explicit_human_review":true,"saved_artifact_protocol":"mcp","saved_artifact_requires_verified_human":true,"saved_artifact_tool":"compile_job_specific_resume","search_requires_identity":false,"status":"available_after_candidate_selects_job","submission_performed":false,"uses_candidate_verified_evidence":true},"degraded":false,"estimated":false,"has_next":true,"jobs":[{"id":"30922887-a8eb-4d71-b922-a837aef04ace","company_id":"4c0fefc3-173a-4227-a823-4d67d3e70ff0","title":"Senior Research Engineer","slug":"senior-research-engineer-9f57f845","description":"Persons in these roles are expected to work from our offices in Seattle. On-site requirements vary based on position and team. If you have questions about on-site work arrangements for this role, please ask your recruiter.\n Our base salary range is $174,240 - $261,360, and in addition we have generous bonus plans to provide a competitive compensation package. \n Who You Are: \n OlmoEarth is growing — more partners, more use cases, and a platform that is evolving quickly. We are looking for a Senior Research Engineer who can collaborate with our partners to tailor the OlmoEarth models to a wide range of specific applications across multiple domains. \n Who We Are:  \n OlmoEarth is an open, end-to-end platform built around a family of foundation models for Earth observation. The platform enables users to create custom fine-tuned models to detect and classify novel geospatial features, handling the full loop: imagery acquisition, annotation, distributed model training and inference, and visualization.\n Our partners span some of the most respected institutions working on wildfire risk, crop mapping, mangrove conservation, and forest protection. OlmoEarth sits within the AI for the Planet group at the Allen Institute for AI, a small, mission-driven team working on conservation, food security, disaster resilience, and climate solutions.\n Learn more: https://allenai.org/olmoearth \n What We Believe \n \n The mission is the point. OlmoEarth exists to put powerful Earth observation tools into the hands of people working on conservation, food security, and climate. Every partner engagement this role supports is connected to that goal. If it matters to you that your day-to-day work adds up to something larger, you are in the right place.\n Good operations are invisible and indispensable. When coordination, documentation, and follow-through are working well, the whole team moves faster and partners have a better experience. This role is the engine behind that.\n Our partners are the signal. We learn what to build and how to improve by staying close to the people using the platform. The feedback, patterns, and friction you surface in this role directly shapes what the team works on next.\n In-person matters. A lot of the best work on this team happens in quick, unplanned conversations between engineering, research, and partnerships. We are mostly in the office because that is where this kind of collaboration happens naturally.\n Say what you think. We make better decisions when people share what they are actually seeing — whether that is a process that is not working, a partner need we are missing, or an idea for doing something differently. Everyone here is still learning, and we like it that way.\n \n Your Next Challenge: \n You will work with partners to deploy OlmoEarth for their use cases. This will require you to move fluidly across the entire OlmoEarth team, working with partners, engineers and researchers. You will make meaningful contributions to all the components of OlmoEarth’s infrastructure (from the finetuning code to model pretraining to our rslearn backend).\n Use case enablement \n \n Collaborate closely with partners to deploy OlmoEarth models in challenging contexts. This will prioritize contexts and partners for which we don’t have immediate solutions or there’s an opportunity to standardize a high quality approach for common use cases.. \n Explore novel use cases for the OlmoEarth models (e.g. post-hoc addition of new modalities, effectively leveraging embeddings in different contexts) which can unlock new use cases and partners.\n \n Partner Communications \u0026 Coordination \n \n As part of model development, maintain communications with key partners to ensure their success using the OlmoEarth platform.\n Communicate  internally  so that partner needs are clearly understood by the OlmoEarth machine learning research, engineering and partnership teams. \n \n Product Improvement and Research \n \n Work closely with the engineering and partnerships  team to feed lessons you learn when deploying models into our infrastructure. This includes improvements to rslearn, OlmoEarth Studio. \n Collaborate with the research team to identify and fix issues with the OlmoEarth models preventing their deployment in specific important applications. \n Continually update our model adaptation approaches to improve model performance for all our partners. This includes updating our fine-tuning approaches, developing recipes for new applications and improving the UI so that modelling trade-offs can be better understood by users. \n Support agent evaluations and development.\n \n What You’ll Need: \n Required\n \n 2+ years of experience deploying machine learning solutions. This covers the full stack of machine learning, including understanding the business case and requirements, training models, and deploying them at scale.\n Technical experience using machine learning tools. This includes fluency in PyTorch, experience debugging traini","salary_min":174240,"salary_max":261360,"location":"Seattle, WA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"senior","tags":["search","fine-tuning","pytorch","pre-training","robotics","generative-ai","research"],"apply_url":"https://job-boards.greenhouse.io/thealleninstitute/jobs/8140098","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-20T18:47:01Z","expires_at":"2026-09-28T13:48:53.582394Z","created_at":"2026-08-25T18:32:53.66843Z","updated_at":"2026-08-29T13:48:53.738873Z","company_name":"Allen Institute for AI","company_slug":"allen-institute-for-ai","company_logo_url":"https://www.google.com/s2/favicons?domain=allenai.org\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/30922887-a8eb-4d71-b922-a837aef04ace"},{"id":"7ed7422c-421c-4fee-b478-a18ac174105a","company_id":"e8c9f3a5-9310-43f5-9341-321fe6d93a92","title":"Principal Machine Learning Engineer, Geometric Vision","slug":"principal-machine-learning-engineer-geometric-vision-81629411","description":"About us    \n Founded in 2017, Wayve is the leading developer of Embodied AI technology.  Our advanced AI software and foundation models enable vehicles to perceive, understand, and navigate any complex environment, enhancing the usability and safety of automated driving systems.\n Our vision is to create autonomy that propels the world forward.  Our intelligent, mapless, and hardware-agnostic AI products are designed for automakers, accelerating the transition from assisted to automated driving.  In our fast-paced environment big problems ignite us—we embrace uncertainty, leaning into complex challenges to unlock groundbreaking solutions. We aim high and stay humble in our pursuit of excellence, constantly learning and evolving as we pave the way for a smarter, safer future.\n At Wayve, your contributions matter.  We value diversity, embrace new perspectives, and foster an inclusive work environment; we back each other to deliver impact.  \n Make Wayve the experience that defines your career!  \n The role  \n As a Principal Engineer on the Model Foundations team you will  build the geometric vision and 3D foundation models that underpin our autonomous driving systems.You will work at the intersection of large-scale deep learning, geometric computer vision, and real-world robotics, developing models that learn 3D structure and dynamics from fleet-scale sensor data.\n You will be a hands-on technical leader. You will set direction for geometric vision, prototype and train new model architectures, build the data and supervision needed to scale them, and take successful ideas through to deployment on real vehicles.\n Key responsibilities \n \n Design and train 3D foundation models and world models using large-scale driving data.\n Develop model architectures for 3D perception, geometric reasoning, reconstruction, and world modeling across space and time.\n Build scalable data generation and auto-labeling pipelines that produce high-quality geometric supervision from large volumes of sensor data.\n Develop and scale offline SLAM and 3D reconstruction systems and pipelines, using large-scale sensor data to recover accurate trajectories, scene geometry, calibration signals, and geometric supervision for model training and evaluation.\n Develop and apply techniques in multi-view geometry, neural rendering, NeRFs, Gaussian Splatting, implicit 3D representations, and feedforward 3D modeling.\n Explore geometry-aware tokenization and representation learning, including efficient ways to encode and fuse information across cameras, viewpoints, time, and sensing modalities.\n Develop foundation vision models that make effective use of camera, radar, LiDAR, and other sensor data for learning rich representations of the physical world.\n Explore video and generative modeling approaches for learning scene structure, dynamics, and future evolution from driving data.\n Train and evaluate models at scale on distributed compute, rapidly iterating on architectures, objectives, data, and training recipes.\n Develop automated evaluation and ground-truth systems for measuring geometric consistency, reconstruction quality, 3D understanding, and downstream driving performance.\n Optimize and deploy models into production autonomous-driving systems, working across model architecture, inference, and onboard constraints.\n Set technical direction for geometric vision at Wayve and work closely with researchers and engineers across foundation models, perception, simulation, data, sensing, and deployment.\n \n About you   \n In order to set you up for success as a Principal Machine Learning Engineer, Geometric Vision at Wayve, we’re looking for the following skills and experience.  \n Essential  \n \n Deep expertise in 3D computer vision, geometric vision, or 3D machine learning, with experience in areas such as multi-view geometry, neural rendering, reconstruction, implicit representations, or world modeling.\n Strong experience designing, training, and evaluating modern deep-learning models at scale, using PyTorch or a comparable framework.\n Strong mathematical and technical foundations in geometry, linear algebra, probability, optimization, and 3D transformations, combined with excellent software engineering skills in Python and C++\n A track record of taking difficult research problems from idea to working system, including building large-scale data, training, evaluation, or deployment pipelines.\n Principal-level technical leadership: the ability to identify high-leverage problems, set research and engineering direction, make strong architectural decisions, and raise the technical bar across teams.\n \n Desirable  \n \n 3D and geometric vision: multi-view geometry, dense 3D reconstruction, neural fields, NeRFs, Gaussian Splatting, or feedforward 3D models.\n Foundation and world models: large-scale vision pre-training, self-supervised learning, video models, generative models, or learned scene dynamics.\n Geometric data engines: offline SLAM, stru","salary_min":407330,"salary_max":460020,"location":"Sunnyvale, CA","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"principal","tags":["distributed-systems","pytorch","deep-learning","computer-vision","autonomous-vehicles","robotics","computer-graphics","pre-training"],"apply_url":"https://wayve.firststage.co/jobs?gh_jid=8724862002","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-18T19:30:55Z","expires_at":"2026-09-28T13:44:29.482689Z","created_at":"2026-08-25T18:31:14.335359Z","updated_at":"2026-08-29T13:44:29.63944Z","company_name":"Wayve","company_slug":"wayve","company_logo_url":"https://www.google.com/s2/favicons?domain=wayve.ai\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/7ed7422c-421c-4fee-b478-a18ac174105a"},{"id":"c329b188-7612-4641-ac6d-796f7502eb3d","company_id":"c587b06c-b6f0-4d1d-b694-6fb6abc2a6bb","title":"Senior Research Engineer, LLM Training \u0026 Post-Training","slug":"senior-research-engineer-llm-training-post-training-b05c8bf9","description":"Who We Are \n Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with less friction.\n Through our merger with Voltage Park, a neocloud and AI Factory, Lightning AI combines developer-first software with cost-efficient, large-scale compute. Teams get the tools they need for experimentation, training, and production inference, with security, observability, and control built in.\n We serve solo researchers, startups, and large enterprises. Lightning AI operates globally with offices in New York City, San Francisco, Seattle, and London, and is backed by Coatue, Index Ventures, Bain Capital Ventures, and Firstminute.\n The Way We Work\n The people who thrive here are builders who move fast, communicate openly, take ownership, and continuously improve themselves, their teams, and our company. Here's what that looks like in practice:\n \n Move with Urgency: We move quickly, make thoughtful decisions, and keep momentum. We value action over perfection and learn by shipping.\n Take Ownership: We own outcomes, not just our individual work. We make decisions that move the company forward and follow through.\n Communicate Openly: We communicate directly, seek to understand, and create clarity for others. Honest conversations help us move faster together.\n Build Great Teams: We lead by example, empower others, and create healthy teams where people can do their best work.\n Raise the Bar: We're always improving ourselves. We learn from feedback, consistently challenge ourselves to grow, and focus on the work that matters most.\n Think Long-Term: We design for what's next. We create scalable systems, simplify complexity, and use AI and automation to amplify our impact.\n \n  \n What We're Looking For\n We're are looking for an experienced Senior Research Engineer who has built, trained, and optimized modern transformer-based language models to join our Research Engineering function at Lightning.\n This role will focus on advancing how large language models are trained, fine-tuned, evaluated, and deployed across Lightning AI's platform and real-world customer workloads. It will work across model training, post-training, PyTorch, distributed systems, and AI systems engineering to improve model quality, training efficiency, and developer productivity while collaborating closely with researchers, infrastructure engineers, and customers.\n We're looking for someone who enjoys turning cutting-edge research into production systems. You have deep experience training and improving transformer-based language models, strong software engineering fundamentals, and a passion for solving difficult problems across model training, evaluation, and AI systems. Rather than building applications on top of existing models, you're motivated by improving the models themselves and the systems that power them. Our work spans models that power the Lightning AI platform, customer-specific model workloads, and research that translates into reusable training and platform capabilities.\n This role is hybrid with a minimum of 2 in-office days per week in San Francisco, Seattle, NYC, or London, with fully remote work considered for candidates outside of our office hub locations. All employees participate in occasional team and company offsites. \n  \n What You'll Do \n \n Design, build, and optimize training and post-training pipelines for large language models.\n Improve model quality through supervised fine-tuning, continued pretraining, preference optimization, reinforcement learning, evaluation, and experimentation.\n Build and improve PyTorch-based training infrastructure, tooling, and developer workflows.\n Optimize distributed training across multi-GPU environments by improving throughput, memory efficiency, scalability, and GPU utilization.\n Investigate model training issues, including convergence, instability, communication overhead, and performance bottlenecks.\n Design evaluation methodologies, benchmark models, analyze failure modes, and acheive model improvements through experimentation.\n Collaborate directly with customers to understand real-world workloads and translate those learnings into improvements across Lightning AI's research platform.\n Partner closely with research, infrastructure, and platform engineering teams to build production-ready AI systems.\n Contribute to open-source projects through new features, tooling improvements, documentation, and community engagement\n \n  \n What You’ll Need \n Required Qualifications \n \n Significant experience training, fine-tuning, evaluating, and/or optimizing transformer-based language models using PyTorch.\n Experience with modern LLM training and post-training techniques such as continued pretraining, SFT, RLHF, preference optimization (DPO, PPO, GRPO), reward modeling, or similar approaches.\n Strong understanding of distributed training and multi-node systems, ","salary_min":165000,"salary_max":310000,"location":"New York, NY","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"senior","tags":["reinforcement-learning","pre-training","pytorch","gpu","fine-tuning","llm","distributed-systems","search"],"apply_url":"https://job-boards.greenhouse.io/lightningai/jobs/7860628003","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-12T19:54:51Z","expires_at":"2026-09-28T13:34:06.427018Z","created_at":"2026-08-25T18:27:03.622997Z","updated_at":"2026-08-29T13:34:06.632514Z","company_name":"Lightning AI","company_slug":"lightning-ai","company_logo_url":"https://www.google.com/s2/favicons?domain=lightning.ai\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/c329b188-7612-4641-ac6d-796f7502eb3d"},{"id":"783b7feb-6610-43ee-9559-b9c3f01b699b","company_id":"c0136eba-1fff-477a-8968-c5435a645cd3","title":"Technical Lead, Multimodal Transformers","slug":"technical-lead-multimodal-transformers-9e9f79ce","description":"Kodiak Robotics, Inc. was founded in 2018 and has become a leader in autonomous ground transportation committed to a safer and more efficient future for all. The company has developed an artificial intelligence (AI) powered technology stack purpose-built for commercial trucking and the public sector. The company delivers freight daily for its customers across the southern United States using its autonomous technology. In 2024, Kodiak became the first known company to publicly announce delivering a driverless semi-truck to a customer. Kodiak is also leveraging its commercial self-driving software to develop, test and deploy autonomous capabilities for the U.S. Department of Defense.\n Kodiak's autonomy stack is built on AI that fuses diverse sensor streams into a unified, actionable understanding of the world. We are developing GigaFusionNet, a large-scale multimodal transformer that learns rich, joint representations across camera, LiDAR, and radar through attention-based fusion. We are looking for a technical leader to own the architecture direction of this effort and grow the engineers building it. This is a senior individual contributor role with significant scope. You will set technical direction for multimodal fusion at Kodiak, lead the workstream executing against it, and be accountable for the results landing on trucks. In this role, you will: \n \n Own the architecture roadmap for multimodal transformers that fuse camera, LiDAR, and radar into unified representations \n Lead the project end to end: problem framing, experiment design, implementation, and production deployment \n Drive research direction on cross-modal attention, token fusion strategies, and efficient multi-stream tokenization, and make the calls on what gets built \n Set the technical bar through design reviews, code reviews, and architectural decisions on scalable training pipelines \n Mentor junior and mid-level engineers, and raise the level of research execution across the org \n Define the pretraining strategy, including self-supervised and contrastive objectives that learn transferable multimodal representations \n Partner with perception, planning, and infrastructure leads to align model design with system-level latency and compute budgets \n \n What you’ll bring: \n \n PhD with 5+ years of industry experience, or MS/BS with 8+ years, in AI, Computer Science, or a related field \n Track record of leading multi-engineer technical efforts from research through production \n Deep expertise in transformer architectures, particularly in multimodal or multi-stream settings \n Strong command of cross-attention, token fusion, and modality alignment techniques \n Experience mentoring engineers and growing technical talent \n Expert proficiency in Python and PyTorch, including large-scale distributed training and mixed-precision optimization \n Ability to influence technical direction across teams without formal authority \n Passion for building AI that reasons over the full breadth of sensory input to operate safely in the real world \n \n What we offer: \n \n Competitive compensation package including equity and annual bonuses \n Excellent Medical, Dental, and Vision plans through Kaiser Permanente, Cigna, and  MetLife (including a medical plan with infertility benefits) \n MetLife Legal Services, Identity \u0026 Fraud Protection, Hospital Indemnity Insurance, Accident Insurance, \u0026 Critical Illness Insurance \n Flexible PTO, 10 paid holidays, and generous parental leave policies \n Our office is centrally located in Mountain View, CA \n Office perks: dog-friendly, free catered lunch, a fully stocked kitchen, and free EV charging \n Long Term Disability, Short Term Disability, Life Insurance \n Wellbeing Benefits - Headspace through Cigna, Calm through Kaiser, One Medical, Gympass, Spring Health through Cigna, Rula (mental health navigation)  \n Fidelity 401(k) \n Commuter, FSA, Dependent Care FSA, HSA \n Various incentive programs (referral bonuses, patent bonuses, etc.) \n The pay range listed below reflects the base salary  in our SF/Silicon Valley location,  across several internal levels. Actual starting pay will be based on job-related factors including: work location, experience, relevant training, education, skill level and performance during interview. Total compensation at Kodiak includes base pay, equity, bonus and a competitive benefits package\n California Pay Range\n $230,000 — $300,000 USD \n  \n At Kodiak, we strive to build a diverse community working towards our common company goals in a safe and collaborative environment where harassment of any kind is strictly prohibited. Kodiak is committed to equal opportunity employment regardless of race, ethnicity, religion, gender identity, sexual orientation, age, disability, or veteran status, or any other basis protected by applicable law.\n  \n In alignment with its business operations, Kodiak adheres to all relevant statutes, regulations, and administrative prerequisites. Accordingly, roles that carry more ","salary_min":230000,"salary_max":300000,"location":"Mountain View, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["data-pipeline","pre-training","robotics","distributed-systems","pytorch","autonomous-vehicles","transformers","machine-learning"],"apply_url":"https://job-boards.greenhouse.io/kodiak/jobs/4363440009","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-11T16:42:20Z","expires_at":"2026-09-28T13:39:18.374821Z","created_at":"2026-08-25T18:28:50.595861Z","updated_at":"2026-08-29T13:39:18.58826Z","company_name":"Kodiak Robotics","company_slug":"kodiak-robotics","company_logo_url":"https://www.google.com/s2/favicons?domain=kodiak.ai\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/783b7feb-6610-43ee-9559-b9c3f01b699b"},{"id":"991f8296-af04-49ea-b15e-7d9b5fc51b29","company_id":"a0000000-0000-0000-0000-000000000001","title":"AI Infrastructure Operations, Demand Planning","slug":"ai-infrastructure-operations-demand-planning-3cf6fb61","description":"About Anthropic \n Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.\n About the Role\n Anthropic runs one of the largest and fastest-growing infrastructure fleets in the industry, across multiple accelerator families, CPU families, clouds, neoclouds, and on-prem sites. Capacity Engineering owns the data, tooling, and systems that let Anthropic plan, measure, and maximize utilization of that fleet: we partner on supply deals, wire telemetry from day zero, own the canonical capacity data layer, and build the planning and enforcement tools every research and product team relies on. This role sits in the Planning pillar, on the Demand Planning team, and works daily with research engineering, pretraining, inference, compute supply, finance, and external vendors.\n You own the tranches. The job has two halves that feed each other. Upstream, you take the Demand Planning forecast and turn it into per-tranche requirements — shape, interconnect, region, supporting resources, date — and carry those into sourcing negotiations and data center build reviews so we contract for capacity we can actually use when we need it. Downstream, you own the integrated schedule and system of record for every tranche in flight — from contracted through reserved, ingested, in-cluster, healthy, and occupied — and you drive the owners of each hop to their dates. Every slip you see downstream becomes a contract-language fix, an automation, or a correction fed back to the forecast.\n What you'll do\n \n Turn the forecast into per-tranche requirements. Take the Demand Planning forecast plus direct input from research, pretraining, and inference planners, and convert it into concrete accelerator, interconnect, region, supporting-resource, and date requirements for each tranche. Represent those in sourcing negotiations and data center build reviews, including which contractual terms actually move delivery dates.\n Qualify tranches for deliverability before signature. The Capacity Planner signs fit-to-forecast; you sign whether the shape can land schedulable, healthy, and instrumented in that region on that date, with storage, egress, identity in place.\n Close the delivery loop. Track forecast-versus-delivered on shape, region, and timing for every tranche; publish the variance; and feed it back to Demand Planning and into the next contract.\n Own the bring-up system of record. Define the canonical contract-to-occupied state machine with explicit entry and exit criteria per stage, and make it a first-class object in the capacity data layer so every downstream tool sees in-flight capacity, not only what has landed.\n Run a portfolio of bring-ups in parallel — new cloud regions, on-prem sites, neocloud blocks — with one integrated schedule spanning provider milestones, cluster creation, network turn-up, storage readiness, health burn-in, and first-workload landing. \n Drive readiness automation: All capacity systems are fully integrated for all new capacity, from contracted through ingested, automated and scaled.\n Instrument and publish the numbers that matter — time-to-occupied and paid-idle dollars per tranche — with executive-level reporting on status, tradeoffs, and risk across the portfolio.\n \n What you bring\n \n Significant experience delivering large-scale infrastructure — cloud regions, accelerator clusters, HPC systems, or bare-metal fleets — at multi-region scale or ≥10k accelerators (or CPU/storage equivalent).\n Technical range from through cluster orchestration and node health, up to the telemetry and planning tables on top — enough to debug where they disagree rather than route it.\n SQL and enough Python to answer your own questions and build your own reporting.\n A degree in a technical field or an equivalent engineering track record.\n \n Preferred\n \n Reserved-capacity onboarding, private offers, or capacity commitments with cloud or neocloud providers.\n Enough demand-planning exposure to challenge a forecast, translate it into per-tranche requirements, and feed delivery variance back into it.\n Data center or colocation delivery: power and space planning, network turn-up, site acceptance, vendor management.\n Accelerator health and burn-in, collective-communications sanity testing, or fleet-health SLOs — and a rigorous definition of \"healthy.\"\n Systems of record or lifecycle services for infrastructure assets.\n Onboarding a new hardware generation into an existing scheduler and observability stack.\n The annual compensation range for this role is listed below. \n For sales roles, the range provided is the role’s On Target Earnings (\"OTE\") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for","salary_min":320000,"salary_max":405000,"location":"San Francisco, CA","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["alignment","search","pre-training","infrastructure"],"apply_url":"https://job-boards.greenhouse.io/anthropic/jobs/5382750008","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-11T12:14:01Z","expires_at":"2026-09-28T13:30:11.563469Z","created_at":"2026-08-25T18:26:10.627699Z","updated_at":"2026-08-29T13:30:11.72368Z","company_name":"Anthropic","company_slug":"anthropic","company_logo_url":"https://www.google.com/s2/favicons?domain=anthropic.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/991f8296-af04-49ea-b15e-7d9b5fc51b29"},{"id":"0629b5fc-5e11-4813-932b-9d15bb730500","company_id":"d49c7f16-1314-459a-acab-7b3d38ee01a9","title":"Member of Technical Staff, Security Engineer","slug":"member-of-technical-staff-security-engineer-5c946307","description":"Magic’s mission is to build safe AGI that accelerates humanity’s progress on the world’s most important problems. We believe the most promising path to safe AGI lies in automating research and code generation to improve models and solve alignment more reliably than humans can alone. Our approach combines frontier-scale pre-training, domain-specific RL, ultra-long context, and inference-time compute to achieve this goal.\n\n\n\n\nABOUT THE ROLE:\n\nAs a Security Engineer at Magic, you will build the software, automation, and infrastructure that protect our models, systems, and engineering organization from both traditional and frontier AI cyber threats. You will respond to security threats, help us build good practices around security incident/threat response.\n\nThe role sits at the intersection of software engineering, infrastructure, and security. You will work closely with researchers, infrastructure engineers, and platform teams to reduce risk while considering the productivity impact. You will help us make the secure path the easy, productive path.\n\nThe threat landscape is changing, supply chain attacks are a major concern, prompt injections and other LLM-native attack vectors are worrying. As frontier AI systems become increasingly capable, security is evolving beyond defending infrastructure alone. Modern models can autonomously discover vulnerabilities, chain exploits across multiple systems, manipulate agents through prompt injection, and escape “secure” containment.\n\nAI safety is extremely important to us at Magic. You will collaborate with our alignment team on whatever they might need to ensure that Magic's models are a safe net-benefit to the world.\n\n\n\n\nWHAT YOU MIGHT WORK ON: \n\n - Design, build, and maintain security tooling and automation used by Magic's engineering teams\n\n - Secure developer infrastructure, CI/CD pipelines, build systems, and production services and make the platform tools easy to use\n\n - Support our sandboxes team with building secure sandboxing and isolation for untrusted code, agent and model execution\n\n - Build guardrails against emerging AI attack vectors, including prompt injection, tool abuse, exploit chaining, and autonomous model behavior\n\n - Protect model weights, research artifacts, datasets, and production infrastructure from unauthorized access, lateral movement, and exfiltration\n\n - Respond to security incidents and live threats\n   \n   \n\n\nWHAT WE’RE LOOKING FOR: \n\n - Software engineering skills with experience building production systems in Rust, Go, Python, C++, Kotlin, TypeScript, or similar languages\n\n - A very good understanding of modern MITRE att\u0026ck techniques\n\n - Actively engaged in the role of LLMs in securing and attacking software systems\n\n - On-call readiness 24/7, assisted by our team\n\n - Experience securing distributed systems, cloud infrastructure, or large-scale production environments\n\n - Ability to develop high-complexity cloud Linux-based exploits\n\n - Deep understanding of UNIX or other operating systems, networking, authentication, authorization, and modern security architecture\n\n - Experience building automation and tooling to secure fast paced organizations\n\n - Track record of exceptional personal integrity, accountability and trustworthiness in high autonomy environments\n\n\n\n\nOUR CULTURE\n\n - Integrity. Words and actions should be aligned\n\n - Hands-on. At Magic, everyone is building\n\n - Teamwork. We move as one team, not N individuals\n\n - Focus. Safely deploy AGI. Everything else is noise\n\n - Quality. Magic should feel like magic\n\nMagic strives to be the place where high-potential individuals can do their best work. We value quick learning and grit just as much as skill and experience.\n\n\n\n\nCOMPENSATION, BENEFITS, AND PERKS (US):\n\n - Annual salary ranges between $225K - $550K based on experience\n\n - Equity is a significant part of total compensation, in addition to salary\n\n - 401(k) plan with 6% salary matching\n\n - Generous health, dental and vision insurance for you and your dependents\n\n - Unlimited paid time off\n\n - Visa sponsorship and relocation stipend to bring you to SF, if possible\n\n - A small, fast-paced, highly focused team","salary_min":225000,"salary_max":550000,"location":"San Francisco, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["distributed-systems","code-generation","cloud","llm","security","alignment","pre-training"],"apply_url":"https://jobs.ashbyhq.com/magic.dev/f9b3e872-cffa-400a-b9e0-621149c5f566/application","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-07T02:15:51.453Z","expires_at":"2026-09-28T13:35:55.954191Z","created_at":"2026-08-25T18:27:38.225391Z","updated_at":"2026-08-29T13:35:56.107874Z","company_name":"Magic","company_slug":"magic","company_logo_url":"https://www.google.com/s2/favicons?domain=magic.dev\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/0629b5fc-5e11-4813-932b-9d15bb730500"},{"id":"0c6dbcba-9806-47c5-96d9-c315911a45b7","company_id":"e8dfc4ee-9649-4fd0-9c16-90d38a1954e1","title":"Member of Technical Staff, Lead Researcher","slug":"member-of-technical-staff-lead-researcher-43af18fd","description":"About the Role \n DoorDash is building an AI Research org from the ground up, and we're hiring our founding researchers. This is not a role inside an existing team — it's a role that defines what the team becomes. You'll have an outsized influence on the research agenda, hiring, culture, infrastructure choices, and how research connects to the rest of DoorDash.\n DoorDash sits on a uniquely valuable substrate for AI research: a real-world, multi-sided marketplace operating at massive scale, with millions of consumers, merchants, and Dashers generating data that no academic lab and few companies can access. We want to build a research org that takes that seriously — one that produces work the broader field cares about, and that fundamentally reshapes how local commerce works.\n You should apply if you want to do ambitious, publishable research in an environment with the data, compute, and operational reach to actually deploy what you build.\n What You'll Do \n \n Set the research agenda for one or more areas of DoorDash AI Research, in close collaboration with the founding team and leadership\n Lead high-impact research projects end-to-end — from problem framing through publication and, where appropriate, production deployment\n Help build the team — interview, recruit, and mentor researchers, engineers, and fellows joining the org\n Shape the org's culture and operating model — how we publish, how we collaborate with product teams, how we balance open research with proprietary work\n Partner across DoorDash with ML platform, product, and operations teams to identify the highest-leverage research bets and translate findings into real-world impact\n \n What You'll Have Access To \n \n Novel proprietary data at marketplace scale — logistics traces, merchant operations, consumer behavior, real-time supply and demand signals, and longitudinal data unavailable anywhere else\n Scalable data collection — ability to design and run structured data collection, leveraging DoorDash’s world-class operational scale, from in-the-wild image and video capture to operational task demonstrations and human-in-the-loop annotation, at a scale and physical-world coverage no other org can match\n High compute budgets for training and inference, sized to support frontier-scale experimentation including large-model pre-training and post-training, RL training runs, and large-scale evaluation sweeps\n Full research infrastructure — DoorDash's internal RL stack, RL environments built on real operational systems, training and evaluation pipelines, and agent evaluation harnesses, with engineering support to extend them as your research demands\n Direct access to leadership — a seat at the table for the decisions that shape the research org, with the autonomy to operate as a principal-level researcher\n Publication freedom — we expect and support publication at top venues (NeurIPS, ICML, ICLR, RSS, CoRL, KDD, etc.) with a fast, supportive internal review process\n Compute and data for external collaborators — budget to bring in academic collaborators, fellows, and visiting researchers as your agenda requires\n \n Research Areas \n We are broadly interested in researchers across the following areas, though the right candidate may reshape this list:\n \n Agentic systems for logistics and local commerce — long-horizon planning, tool use, multi-agent coordination, and evaluation methodologies for agents operating in physical-world marketplaces\n Memory and personalization — transfer RL, continual learning, harness-based improvements, and systems that adapt to individual consumers, merchants, and Dashers over time without catastrophic forgetting or unsafe drift\n Foundation models for marketplace dynamics — forecasting, pricing, matching, and personalization at marketplace scale, including domain-specific pre-training and post-training\n Evaluation and measurement — new benchmarks, eval harnesses, and methodologies for ML systems deployed in messy, real-world operational settings\n Multimodal understanding — vision, speech, and language applied to merchant catalogs, in-store and on-the-road imagery, and consumer interfaces\n Robotics and embodied AI for last-mile delivery — perception, planning, and learning systems for the physical edge of the marketplace\n \n Who We're Looking For \n \n A strong research track record:  first-author publications at top ML venues, or equivalent demonstrated output (widely-used systems, influential open-source work, or research artifacts adopted at scale)\n Experience leading ambitious research projects end-to-end, from problem framing through to results that other people built on\n Taste:  the ability to identify which problems are worth working on, when an approach is exhausted, and when a result is real\n A builder's instinct:  comfortable working close to real systems and real data, not just on benchmarks\n Excitement about the founding-team aspects of the role: hiring, agenda-setting, culture-building, and","salary_min":203500,"salary_max":299300,"location":"San Francisco, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["pre-training","cloud","healthcare","generative-ai","fine-tuning","agents","search","robotics"],"apply_url":"https://job-boards.greenhouse.io/doordashusa/jobs/8105902","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-04T16:00:31Z","expires_at":"2026-09-28T13:50:49.871244Z","created_at":"2026-08-25T18:33:45.011426Z","updated_at":"2026-08-29T13:50:50.034882Z","company_name":"DoorDash","company_slug":"doordash","company_logo_url":"https://www.google.com/s2/favicons?domain=doordash.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/0c6dbcba-9806-47c5-96d9-c315911a45b7"},{"id":"2229113f-ed51-4e63-bb05-2344fae3bf5c","company_id":"6ce2d21e-b00f-4343-9bd0-5ac62ff81431","title":"Staff Research Scientist, Foundation Models Recipes","slug":"staff-research-scientist-foundation-models-recipes-4b8515df","description":"Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver—The World's Most Experienced Driver™—to improve access to mobility while saving thousands of lives now lost to traffic crashes. The Waymo Driver powers Waymo’s fully autonomous ride-hail service and can also be applied to a range of vehicle platforms and product use cases. The Waymo Driver has provided over ten million rider-only trips, enabled by its experience autonomously driving over 100 million miles on public roads and tens of billions in simulation across 15+ U.S. states.\n The mission of the Waymo AI Foundations team is to develop machine learning solutions addressing open problems in autonomous driving, towards the goal of safely operating Waymo vehicles in dozens of cities and under all driving conditions. As part of our work, we also initiate and foster collaborations with other research teams in Alphabet. AI Foundations areas that we are currently focusing on include reinforcement learning, learning from demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust evaluation.\n In this hybrid role, you will report to a Senior Director of AI Foundations \n You will: \n \n Own the data recipe for Waymo’s Foundation Model pre-training and post-training\n Lead and drive science on best practices around clustering, filtering, de-duplication, long-tail data mining, memorization, etc. \n Tech-lead a team of research engineers to build the data flywheel to enable the above from Waymo’s massive driving data\n Integrate emerging research from the broader community to do rigorous ablations and promising data research techniques. This includes scaling ladders for pre-training, data mix optimizations and RL recipes (preferences) for post-training. \n Engage with the wider research community on best practices around data centric evaluation creation and refinement \n Partner with engineering and research teams across Waymo to share recipes, techniques, and post-training best practices to accelerate our collective know-how\n \n You have: \n \n PhD or Masters in Computer Science, Machine Learning, Robotics, or a similar technical field; with 3+ years of industry or post-doc research experience in Data-centric AI, Reinforcement Learning or Foundation Models\n Demonstration of original contributions to the field through high-impact publications (ArXiv, peer-reviewed conferences like NeurIPS/ICLR/CVPR), technical blog posts, or significant open-source contributions\n Proficiency and in-depth knowledge of the inner workings of an ML framework (e.g. Pytorch, JAX, Tensorflow)\n \n We prefer: \n \n Extensive experience working with data and recipes for large scale foundation models\n ML infra experience: training, evaluating and deploying ML models at scale\n Deep learning experience, especially with generative models, e.g., LLMs/VLMs, and/or reinforcement learning\n \n In accordance with Washington state law, we are highlighting our comprehensive benefits package, which is available to all eligible US based employees. Benefits for this role include:\n \n Health, dental, vision, life, disability insurance\n Retirement Benefits: 401(k) with company match\n Paid Time Off: 20 days of vacation per year, accruing at a rate of 6.15 hours per pay period for the first five years of employment\n Sick Time: 40 hours/year (statutory, where applicable); 5 days/event (discretionary)\n Maternity Leave (Short-Term Disability + Baby Bonding): 28-30 weeks\n Baby Bonding Leave: 18 weeks\n Holidays: 13 paid days per year\n \n Please note that Waymo may not be able to employ remotely in all locations. Please speak with your recruiter about your preferred location for remote work when you begin the interview process\n The expected base salary range for this full-time position across US locations is listed below. Actual starting pay will be based on job-related factors, including exact work location, experience, relevant training and education, and skill level. Your recruiter can share more about the specific salary range for the role location or, if the role can be performed remote, the specific salary range for your preferred location, during the hiring process.  \n Waymo employees are also eligible to participate in Waymo’s discretionary annual bonus program, equity incentive plan, and generous Company benefits program, subject to eligibility requirements.  \n Salary Range\n $251,000 — $310,000 USD","salary_min":251000,"salary_max":310000,"location":"Mountain View, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["autonomous-vehicles","search","tensorflow","robotics","pytorch","deep-learning","pre-training","generative-ai"],"apply_url":"https://careers.withwaymo.com/jobs?gh_jid=8075587","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-03T23:19:22Z","expires_at":"2026-09-28T13:35:26.990236Z","created_at":"2026-08-25T18:27:26.009368Z","updated_at":"2026-08-29T13:35:27.180296Z","company_name":"Waymo","company_slug":"waymo","company_logo_url":"https://www.google.com/s2/favicons?domain=waymo.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/2229113f-ed51-4e63-bb05-2344fae3bf5c"},{"id":"145cd846-4fea-4df7-a373-0f0eba950f29","company_id":"a0000000-0000-0000-0000-000000000001","title":"Research Scientist, Takeoff Intel","slug":"research-scientist-takeoff-intel-cd718d9d","description":"About Anthropic \n Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.\n About the role\n We're looking for a Research Scientist who has done hands-on research on large models (pretraining, fine-tuning, RL, evals, or agents scaffolds) and wants to focus on measuring and understanding recursive-self-improvement. You know what the model-development loop looks like from the inside: which signals matter and where the real bottlenecks are. On this team you'll use that judgment to decide what's worth measuring, design the evaluations and models that measure it, and interpret what the results mean for how fast this is moving.\n  \n We're hiring at both junior and senior levels. Senior researchers should be comfortable doing hands-on technical work alongside setting research direction.\n Responsibilities\n \n \n Identify the signals that track AI R\u0026D acceleration and design the evaluations that measure them\n \n Build quantitative models of capability growth and self-improvement dynamics, grounded in evaluation and telemetry data\n \n Run experiments and evals to test hypotheses about automation and capability\n \n Make opinionated research bets and own the outcome\n \n Write graded assessments of what our measurements show, for internal decision-makers and public reporting\n \n Collaborate with pretraining, RL, economic research, and policy teams\n \n You may be a good fit if you\n \n \n Have done hands-on research on large language models: pretraining, fine-tuning, RL, evals, or agent systems\n \n Have strong quantitative instincts, are comfortable with quantitative modeling and reasoning\n \n Have experience in forecasting, may have published AI forecasting scenarios\n \n Can design an evaluation from a vague question and defend the methodology\n \n Write clearly and calibrate: state confidence, name what would change your conclusion\n \n Are motivated by impact: comfortable with work whose output is graded assessments and system-card sections more often than papers\n \n Care about AI safety and think carefully about where rapid capability growth leads\n \n Strong candidates may also have\n \n \n Trained or RL'd frontier models hands-on\n \n Experience with scaling laws, capability forecasting, or emergent-capability studies\n \n A physics, applied-math, or similarly quantitative background that moved into ML\n \n Written a system card section, capability report, or methodology document that others cite\n \n Experience supervising and correcting AI-written code\n \n Some examples of our work\n \n \n Anthropic ECI:  our adaptation of Epoch Capabilities Index published in all recent system cards to measure capability acceleration\n \n AI R\u0026D capability assessments in the Claude system cards\n \n When AI Builds Itself : all data in the article comes from our team\n The annual compensation range for this role is listed below. \n For sales roles, the range provided is the role’s On Target Earnings (\"OTE\") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role.\n Annual Salary:\n $350,000 — $850,000 USD \n Logistics \n Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience\n Required field of study:  A field relevant to the role as demonstrated through coursework, training, or professional experience\n Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position\n Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices.\n Visa sponsorship:  We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this.\n We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed.  Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work. We think AI systems like the ones we're building have enormous social and ethical implications. We think this makes representation even more important, and we strive to include a range of diverse perspectives on our team. Your safety matters to us. To protect yourself from potential scams, remember that Anthropic recruiters only contact you from @anthropic.com email addresses","salary_min":350000,"salary_max":850000,"location":"San Francisco, CA","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"principal","tags":["llm","fine-tuning","alignment","pre-training","research"],"apply_url":"https://job-boards.greenhouse.io/anthropic/jobs/5370669008","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-07-24T23:31:13Z","expires_at":"2026-09-28T13:30:32.362125Z","created_at":"2026-07-25T14:00:36.359264Z","updated_at":"2026-08-29T13:30:32.509792Z","company_name":"Anthropic","company_slug":"anthropic","company_logo_url":"https://www.google.com/s2/favicons?domain=anthropic.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/145cd846-4fea-4df7-a373-0f0eba950f29"},{"id":"2cbdb142-2bc4-4a36-a0eb-2c8b1efac653","company_id":"2721f049-2cf2-4e3e-82d0-8d8df89c8f90","title":"Senior Machine Learning Engineer, Model Training and Reinforcement Learning","slug":"ml-systems-engineer-large-scale-model-training-rl-infrastructure-a5ac7fa0","description":"About Nebius: \n Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.\n Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.\n Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R\u0026D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R\u0026D.\n The role   \n Nebius Token Factory is building an AI training and model post-training capability for frontier model improvement. This role owns the infrastructure that makes large-scale training and RL experiments possible, reliable, reproducible, and efficient. The work sits at the intersection of distributed systems, GPU performance, model training frameworks, RL pipelines, and production engineering. \n A Senior Machine Learning Engineer owns substantial ML work end to end. They can translate an ambiguous capability goal into concrete experiments, implement and debug training and RL recipes, build the supporting data and systems, and deliver measurable improvements in model quality, experiment throughput, and reliability. They are deeply hands-on and can independently debug both model-behavior failures and distributed training failures.\n Your responsibilities :   \n \n \n Design and run model-training and post-training experiments, including SFT , continued pretraining, preference optimization ( DPO /IPO/ KTO ), and RL methods such as RLHF / RLAIF , PPO , and GRPO .\n \n Build reward functions, judge models, verifiers, task environments, and evaluation sets for reasoning, coding, tool use, and agentic workflows.\n \n Create synthetic data and data pipelines, including teacher-student generation, self-play, rejection sampling, filtering, and quality scoring.\n \n Analyze model-behavior failures and turn them into targeted data, reward, or algorithm improvements.\n \n Build and maintain distributed training and RL infrastructure using frameworks such as Megatron- LM , DeepSpeed, PyTorch FSDP /DTensor, Ray, verl, slime, AReaL, or OpenRLHF.\n \n Implement and debug parallelism strategies (tensor, pipeline, sequence/context, expert, and data parallelism) and build reliable rollout, reward-serving, checkpointing, and experiment-orchestration components.\n \n Profile and improve GPU utilization, memory usage, communication efficiency, training throughput, and inference/serving performance.\n \n Design rigorous evaluations and ablations for capability, instruction following, reasoning, tool use, safety, and regression risk.\n \n Write clear experiment plans, design docs, benchmark reports, and runbooks, and partner across research and platform teams.\n \n Must-haves :   \n \n \n Strong Python and PyTorch engineering skills, with the ability to move quickly from idea to experiment to working system.\n \n Hands-on experience across at least two of: model training, post-training/ RL , applied modeling, data pipelines, or large-scale ML systems.\n \n Ability to design rigorous experiments with baselines, ablations, metrics, and failure analysis.\n \n Practical understanding of modern LLM behavior, instruction tuning, preference optimization, and evaluation challenges.\n \n Practical understanding of transformer training bottlenecks, memory pressure, communication overhead, and checkpointing.\n \n Ability to reason quantitatively about model quality, throughput, utilization, reliability, cost, and research velocity.\n \n Strong communication skills and ability to collaborate with researchers, engineers, and leadership.\n \n Nice - to - have s :   \n \n \n Experience with LLM post-training, RL , agents, reward modeling, synthetic data, or model evaluation.\n \n Experience with RL frameworks or pipelines such as verl, slime, AReaL, OpenRLHF, TRL , or custom PPO / GRPO / RLHF systems.\n \n Experience with Megatron- LM , DeepSpeed, PyTorch FSDP /DTensor, Ray, Slurm, or Kubernetes on large GPU clusters.\n \n Familiarity with NCCL , CUDA , Triton, Nsight, InfiniBand/ RDMA , and H100/H200/B200 clusters, or with model serving and inference optimization.\n \n Publications, open-source contributions, or production impact in LLM post-training, RL , reasoning, coding models, synthetic data, distributed training, or evaluation.\n \n Experience designing agent environments, tool-use tasks, or verifier-based rewards.\n \n Key employee benefits in the US: \n \n \n Health insurance:  100% company-paid medical, dental, and vision coverage for employees and families.\n \n 401(k) plan:  Up to 4% company match with immediate vesting.\n \n Parental leave:  20 weeks paid for primary caregivers, 12 weeks for secondary caregivers.\n \n Remote work reimbu","salary_min":195200,"salary_max":262200,"location":"Palo Alto, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"senior","tags":["pytorch","mlops","gpu","pre-training","data-pipeline","reinforcement-learning","agents","llm"],"apply_url":"https://careers.nebius.com/?gh_jid=4926274101","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-07-22T22:52:25Z","expires_at":"2026-09-28T13:46:37.503072Z","created_at":"2026-07-24T14:15:03.931596Z","updated_at":"2026-08-29T13:46:37.662671Z","company_name":"Nebius","company_slug":"nebius","company_logo_url":"https://www.google.com/s2/favicons?domain=nebius.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/2cbdb142-2bc4-4a36-a0eb-2c8b1efac653"},{"id":"8732df68-77f6-421f-a2f4-b540a3a85837","company_id":"a0000000-0000-0000-0000-000000000001","title":"Pre-training Distributed Systems Tech Lead / Manager","slug":"evals-infrastructure-tech-lead-manager-d2163e2a","description":"About Anthropic \n Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.\n About the Role\n Anthropic is at the forefront of AI research, dedicated to developing safe, ethical, and powerful artificial intelligence. Our mission is to ensure that transformative AI systems are aligned with human interests. We're looking for an experienced tech lead to join our Evals Infrastructure team, building the systems that let us measure what our models can actually do. Evaluation is how we know whether a model is safe to ship — you'd own the infrastructure that makes those measurements fast, reliable, and trustworthy at scale. In this role you'll work at the intersection of inference, research and infrastructure engineering: managing the large scale distributed systems that orchestrate evals for our frontier models, building and scaling the harnesses researchers use to design and run evals, making results reproducible and interpretable, and ensuring eval signal is available where decisions get made. Your work directly shapes what we build and what we don't.\n Responsibilities \n \n Lead the team building the distributed systems that schedule, orchestrate, and execute evals for our frontier model training\n Own eval throughput and cost: compute allocation across suites, queueing against constrained accelerator pools, caching and reuse of eval work\n Build and scale the harnesses researchers use to define, run, and iterate on evals\n Make eval results trustworthy — determinism, reproducibility, and honest uncertainty quantification on reported metrics\n Ensure eval signal reaches the dashboards and reviews where launch decisions actually get made\n Contribute directly as an engineer while managing and growing the team, prioritizing its work, and coaching your reports\n \n You may be a good fit if you \n \n Have led technical projects end-to-end on large-scale distributed systems, and have 1+ years managing engineers (or tech-lead-with-reports experience)\n Are strong in Python and Rust\n Have built high-throughput, fault-tolerant systems on cloud or on-prem accelerator fleets\n Care about measurement quality, not just pipeline uptime — you'd notice if a metric moved for the wrong reason\n Communicate well with researchers and can translate research needs into infrastructure\n Are deeply interested in the transformative effects of advanced AI and committed to safe development\n \n Strong candidates may have \n \n Worked on LLM inference or training infrastructure\n Experience with eval or benchmarking systems, especially agentic evals requiring sandboxed execution\n Working statistical literacy — variance, confidence intervals, sample-size sufficiency for noisy metrics\n Experience with observability and regression detection over time-series metrics\n \n Sample Projects \n \n Rebuilding the eval orchestration layer to cut wall-clock time on the pre-train eval suite\n Designing compute allocation and scheduling so eval suites fit inside a fixed fraction of a production run's chip-hours\n Adding rigorous uncertainty estimates to top-line dashboard metrics so checkpoint-to-checkpoint comparisons are actually decision-grade\n Building sandboxed execution infrastructure for agentic evals\n The annual compensation range for this role is listed below. \n For sales roles, the range provided is the role’s On Target Earnings (\"OTE\") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role.\n Annual Salary:\n $500,000 — $850,000 USD \n Logistics \n Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience\n Required field of study:  A field relevant to the role as demonstrated through coursework, training, or professional experience\n Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position\n Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices.\n Visa sponsorship:  We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this.\n We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed.  Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematur","salary_min":500000,"salary_max":850000,"location":"San Francisco, CA","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["llm","agents","pre-training","alignment","distributed-systems","research"],"apply_url":"https://job-boards.greenhouse.io/anthropic/jobs/5367417008","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-07-22T20:04:09Z","expires_at":"2026-09-28T13:30:26.350902Z","created_at":"2026-07-24T14:00:20.388339Z","updated_at":"2026-08-29T13:30:26.499245Z","company_name":"Anthropic","company_slug":"anthropic","company_logo_url":"https://www.google.com/s2/favicons?domain=anthropic.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/8732df68-77f6-421f-a2f4-b540a3a85837"},{"id":"5135f8c4-2ddd-450e-aeb2-2df77b89bc40","company_id":"6ce2d21e-b00f-4343-9bd0-5ac62ff81431","title":"Director, Foundation Model Data Recipes ","slug":"director-foundation-model-data-recipes-063ef7fa","description":"Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver—The World's Most Experienced Driver™—to improve access to mobility while saving thousands of lives now lost to traffic crashes. The Waymo Driver powers Waymo’s fully autonomous ride-hail service and can also be applied to a range of vehicle platforms and product use cases. The Waymo Driver has provided over ten million rider-only trips, enabled by its experience autonomously driving over 100 million miles on public roads and tens of billions in simulation across 15+ U.S. states.\n The mission of the Waymo AI Foundations team is to develop machine learning solutions addressing open problems in autonomous driving, towards the goal of safely operating Waymo vehicles in dozens of cities and under all driving conditions. As part of our work, we also initiate and foster collaborations with other research teams in Alphabet. AI Foundations areas that we are currently focusing on include reinforcement learning, learning from demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust evaluation.\n In this hybrid role, you will report to the Senior Director, AI Foundations.\n  \n You will: \n \n Participate in Waymo’s Foundation World Model pre-training and post-training by owning the data recipe\n Invest in data best practices - best practices around clustering, filtering, de-duplication, long-tail data mining, memorization, etc. Also help lead a team of research engineers to build the data flywheel to enable the above from Waymo’s massive driving data\n Integrate emerging research from the broader community to do rigorous ablations and promising data research techniques. Focus on managing a team of scientists that would work on scaling ladders for pre-training, data mix optimizations and RL recipes (preferences) for post-training. \n Engage with the wider research community on best practices around data centric evaluation creation and refinement \n Partner with engineering and research teams across Waymo to share recipes, techniques, and post-training best practices to accelerate our collective know-how\n \n  \n You have: \n \n PhD or Masters in Computer Science, Machine Learning, Robotics, or a similar technical field; with 5+ years of industry or post-doc research experience in Reinforcement Learning or Foundation Models\n Demonstration of original contributions to the field through high-impact publications (ArXiv, peer-reviewed conferences like NeurIPS/ICLR/CVPR), technical blog posts, or significant open-source contributions\n Experience leading teams of research scientists and engineers to deploy real world impact in AI - including showing real world impact. Ability to be hands-on when needed to adapt models as and when needed.\n \n  \n We prefer: \n \n Extensive experience working with data and recipes for large scale foundation models\n PhD in Computer Science, Machine Learning, or Robotics, with a research focus on Data centric AI, Reinforcement Learning, Foundation Models, or Multimodal learning\n \n  \n The expected base salary range for this full-time position across US locations is listed below. Actual starting pay will be based on job-related factors, including exact work location, experience, relevant training and education, and skill level. Your recruiter can share more about the specific salary range for the role location or, if the role can be performed remote, the specific salary range for your preferred location, during the hiring process.  \n Waymo employees are also eligible to participate in Waymo’s discretionary annual bonus program, equity incentive plan, and generous Company benefits program, subject to eligibility requirements.  \n Salary Range\n $349,000 — $431,000 USD","salary_min":349000,"salary_max":431000,"location":"Mountain View, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["pre-training","robotics","search","generative-ai","reinforcement-learning","autonomous-vehicles"],"apply_url":"https://careers.withwaymo.com/jobs?gh_jid=8075690","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-07-20T22:39:19Z","expires_at":"2026-09-28T13:35:19.277885Z","created_at":"2026-07-22T14:04:58.20554Z","updated_at":"2026-08-29T13:35:19.43594Z","company_name":"Waymo","company_slug":"waymo","company_logo_url":"https://www.google.com/s2/favicons?domain=waymo.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/5135f8c4-2ddd-450e-aeb2-2df77b89bc40"},{"id":"e57cf48d-3756-4016-8e50-400a76bbaa5d","company_id":"714f360f-a244-487d-b3f0-0c43518a9e66","title":"Staff Machine Learning Engineer, Visual AI","slug":"staff-machine-learning-engineer-computer-vision-147d8a7f","description":"About Pinterest: \n Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product.\n Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the  flexibility to do your best work. Creating a career you love? It’s Possible.\n At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI.\n Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here .\n About the Team: \n Hundreds of millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love.\n Within Pinterest, the Pinterest Labs organization focuses on applied ML research and development to power the platform. Labs works across a broad variety of AI/ML initiatives, including LLMs/VLM, agent design, core computer vision, multimodal representation learning, visual generative modeling, recommender systems, graph learning, and more. This is the group that develops the foundation AI models that fully leverage the hundreds of billions of Pins and the associated knowledge graphs, and ships new product capabilities to fully utilize these technologies.\n We are currently hiring for the Visual team in Labs, which develops Pinterest's foundation visual models. In this role, you'll work with Pinterest's rich visual-text dataset to train large-scale VLMs, encoders, and diffusion models from scratch that are continuously shipped to production to power visualization and search capabilities. The team is subdivided into two pods, which share pretraining datasets, training and RL infrastructure, and generally co-develop our visual modeling ecosystem. The visual understanding pod builds the core visual embeddings and token models such as PinCLIP and deeply integrates them with the VLM/LLM/agent ecosystem at Pinterest. The visual generative pod builds Pinterest Canvas , our production image editing and generation model used across our assistant, visualization, and monetization products. Labs is staffed to be a highly collaborative environment where research scientists, ML engineers, product builders, and infrastructure engineers all work directly together, so that the team can plug into any technical effort at the company to support fast productionization.\n What you’ll do: \n \n Build state-of-the-art visual encoders, VLMs, and diffusion models that power Pinterest's visual AI capabilities\n Experiment with billion-scale image datasets, backed by large-scale GPU computing.\n Build flexible visual reasoning tools such as composed image retrieval, promptable image feature computation, instruction-tuned embedding and generative editing models, and more.\n Read research papers, participate in group discussions, and directly participate in brainstorming the company's overall AI strategy.\n Help construct data agents to build training data that can be shared across multimodal representation, composed image retrieval, image-editing generation, and visual language modeling.\n Collaborate directly with product engineers and infrastructure engineers to ship new capabilities in the core product.\n Publish and share your work through conferences like CVPR and KDD, paper submissions, and blog posts.\n Mentor junior researchers and research interns within the Pinterest Labs organization.\n Collaborate across a team situated across San Francisco, Seattle, NYC, and remote roles.\n \n What we’re looking for: \n \n Research engineers and scientists with experience building and training large scale vision models of all categories.\n Experience with multimodal representations and visual language modeling is strongly preferred.\n A track record of research contributions (e.g., publications, open-source work) and/or shipping ML models to production.\n Hands-on experience with large-scale model training and modern deep learning frameworks (e.g., PyTorch).\n Strong collaboration skills and a demonstrated ability to work effectively in a small, fast-moving team.\n M.S. or PhD in Machine Learning or related academic areas, or equivalent work experience.\n Experience using AI-accelerated research tooling akin to auto-research, data agents, etc","salary_min":189308,"salary_max":389753,"location":"San Francisco, CA","workplace":"remote","remote_scope":"unknown","job_type":"full-time","experience_level":"lead","tags":["computer-vision","diffusion-models","search","pytorch","llm","deep-learning","pre-training","machine-learning"],"apply_url":"https://www.pinterestcareers.com/jobs/?gh_jid=8015537","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-07-13T17:51:37Z","expires_at":"2026-09-28T13:39:31.960691Z","created_at":"2026-07-15T14:10:33.975738Z","updated_at":"2026-08-29T13:39:32.116557Z","company_name":"Pinterest","company_slug":"pinterest","company_logo_url":"https://www.google.com/s2/favicons?domain=www.pinterest.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/e57cf48d-3756-4016-8e50-400a76bbaa5d"},{"id":"8c402485-1400-4e3b-aacf-eaa1ab3b5dfb","company_id":"3da82454-107f-427f-88e7-01f315ef93fb","title":"Research Engineer - Distributed Training","slug":"research-engineer-distributed-training-19cda6e4","description":"OWN YOUR INTELLIGENCE\n\n\n\nPrime Intellect is building the open superintelligence stack: the infrastructure frontier AI labs build internally, made available to every ambitious AI team. \n\n\n\nOur platform, Lab, unifies compute, environments, evaluations, secure sandboxes, high-performance training, and deployment into one full-stack system for post-training at frontier scale - from SFT and RL to tool use, agent workflows, and continuously improving production models. We are building open frontier AI: open-source models trained end to end for long-horizon tasks like autonomous research, and the full-stack platform our own research team uses to build them. The next generation of AI companies, enterprises, and research teams do not just need more GPUs. They need the ability to turn their own workflows, tools, data, and feedback loops into superintelligence they own.\n\nWe train open frontier models and ship the same stack to our customers. Its spans the full stack of training, deploying and continuously improving models — compute, large-scale RL, environments, sandboxes, evals, and deployment.\n\n\n\nPrime Intellect has raised $150M in total funding from Founders Fund, Radical Ventures, NVIDIA, and exceptional AI, infrastructure, and enterprise operators — including Andrej Karpathy, Dwarkesh Patel, and leaders and founders from Ramp, Perplexity, Harvey, Mercor, Zapier, Datadog, Semianalysis, Cognition, OpenAI, Thinking Machines, Together AI, SemiAnalysis, LangChain, Browserbase, Cloudflare, Sierra, Databricks, Airbnb, OpenRouter, Standard Intelligence, Fleet, Core Auto, and more. We are looking for people who want to build at the intersection of frontier research, real infrastructure, and go-to-market for a category that does not fully exist yet.\n\n\n\n\nWHAT YOU’LL WORK ON\n\n - Build and optimize the distributed training infrastructure behind our pre-training and large-scale RL training workloads by contributing to our prime-rl https://github.com/PrimeIntellect-ai/prime-rl framework.\n\n - Improve end-to-end training efficiency across compute, memory, networking, and scheduling layers.\n\n - Design and implement low-level performance optimizations, including kernels, communication paths, and runtime improvements.\n\n - Work on distributed training systems spanning data, tensor, and pipeline parallel workloads.\n\n - Help shape the architecture of our RL training stack, including async rollout and post-training systems.\n\n - Contribute to open-source libraries and internal infrastructure used for frontier-scale model training.\n\n - Collaborate closely with researchers and infrastructure engineers to translate bottlenecks into concrete systems improvements.\n\n - Stay at the frontier of training systems, inference systems, compiler/runtime tooling, and hardware-aware optimization techniques.\n\n\n\n\n\nYOU MAY BE A FIT IF YOU HAVE\n\n - Strong systems engineering experience in AI/ML infrastructure, especially around large-scale model training or inference.\n\n - Deep familiarity with PyTorch and distributed training frameworks such as PyTorch Distributed, DeepSpeed, FSDP, Megatron, vLLM, Ray, or related tooling.\n\n - Experience optimizing training performance across kernels, memory movement, communication overhead, or parallelization strategy.\n\n - Hands-on experience with large-scale training techniques including data parallelism, tensor parallelism, and pipeline parallelism.\n\n - Strong understanding of GPU architecture, profiling, and performance debugging.\n\n - Ability to identify bottlenecks across the stack and drive improvements from first principles.\n\n - Comfort working in a fast-moving environment with ambiguous problems and high ownership.\n\n\n\n\nESPECIALLY EXCITING\n\n - Experience writing or optimizing CUDA / Triton kernels.\n\n - Experience with compiler or runtime optimization for ML systems.\n\n - Experience working on RL training infrastructure, rollout systems, or asynchronous training pipelines.\n\n - Experience with multi-node GPU clusters and high-performance networking.\n\n - Contributions to open-source ML systems or infrastructure projects.\n\n - Interest in publishing technical work or sharing insights through engineering blogs and technical writing.\n\n\n\n\n\n\n\nBENEFITS \u0026 PERKS\n\n - Cash Compensation Range of $150-350k, plus equity incentives, aligning your success with the growth and impact of Prime Intellect.\n\n - Flexible work arrangements, with the option to work remotely or in-person at our offices in San Francisco.\n\n - Visa sponsorship and relocation assistance for international candidates.\n\n - Quarterly team off-sites, hackathons, conferences and learning opportunities.\n\n - Opportunity to work with a talented, hard-working and mission-driven team, united by a shared passion for leveraging technology to accelerate science and AI.\n\nIf you’re excited about building the systems foundation for frontier-scale training and open superintelligence, we’d love to hear from you.","salary_min":150000,"salary_max":350000,"location":"San Francisco, CA","workplace":"remote","remote_scope":"unknown","job_type":"full-time","experience_level":"senior","tags":["pre-training","search","distributed-systems","gpu","pytorch","llm","agents","research"],"apply_url":"https://jobs.ashbyhq.com/PrimeIntellect/8bd52610-175c-42a7-a7cd-b29c45f9d305/application","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-07-08T18:43:34.749Z","expires_at":"2026-09-28T13:41:04.992526Z","created_at":"2026-04-13T15:01:32.550978Z","updated_at":"2026-08-29T13:41:05.142284Z","company_name":"Prime Intellect","company_slug":"PrimeIntellect","company_logo_url":"https://www.google.com/s2/favicons?domain=primeintellect.ai\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/8c402485-1400-4e3b-aacf-eaa1ab3b5dfb"},{"id":"5b4c2841-d819-4ab1-87e4-988c9bff0235","company_id":"a0000000-0000-0000-0000-000000000003","title":"Senior Software Engineer, Identity","slug":"senior-software-engineer-identity-542360e2","description":"Software is eating the world, but AI is eating software. We live in unprecedented times – AI has the potential to exponentially augment human intelligence. Every person will have a personal tutor, coach, assistant, personal shopper, travel guide, and therapist throughout life. As the world adjusts to this new reality, leading platform companies are scrambling to build LLMs at billion scale, while large enterprises figure out how to add it to their products. To make them safe, aligned and actually useful, these models need human eval and reinforcement learning through human feedback (RLHF) during pre-training, fine-tuning, and production evaluations. This is the main innovation that’s enabled ChatGPT to get such a large headstart among competition.\n At Scale, our products include the Generative AI Data Engine, SGP, Donovan, and others that power the most advanced LLMs and generative models in the world through world-class RLHF, human data generation, model evaluation, safety, and alignment. The data we are producing is some of the most important work for how humanity will interact with AI.\n At the foundation of these products is the Identity  Engineering team.  In this role, you will help support the design and development of core software systems specifically focused on identity, access management, authorization, and authentication.  You’ll also get widespread exposure to the forefront of the AI race as Scale sees it in enterprises, startups, governments, and large tech companies.\n You will:\n \n Drive the design, and implementation of our identity infrastructure to ensure secure authentication and authorization across enterprise systems.\n Build software for authentication mechanisms such as Single Sign-On (SSO), Multi-Factor Authentication (MFA), and federated identity solutions (SAML, OAuth, OpenID Connect).\n Build software for authorization mechanisms such as Relation-based access control (ReBAC), Attribute-based access control (ABAC), Role-based access control (RBAC).\n Build software-defined identity governance policies to ensure compliance with security policies, industry regulations (e.g., NIST, SOC2, ISO 27001), and organizational standards.\n Present technical information to teams and stakeholders, providing guidance and insight on identity management and best practices.\n \n Ideally you’d have:\n \n 5+ years of full-time engineering experience, post-graduation with specialities in infrastructure and identity systems.\n Infrastructure expertise – IAM controls, Infrastructure as Code (Terraform, Pulumi), microservice deployment best practices.\n Hands-on experience working with OpenFGA, Authzed, Cedar, Topaz, or similar authorization frameworks at scale.\n Strong understanding of Zanzibar-based ReBAC models, relationship tuples, and access control evaluation.\n Strong knowledge of authentication standards such as OAuth 2.0, OIDC, SAML, and JWT, as well as industry standard IdP solutions like EntraID, Okta, etc.\n Extensive experience in software development and a deep understanding of distributed systems and public cloud platforms (AWS preferred).\n Show a track record of independent ownership of successful engineering projects.\n Possess excellent communication and collaboration skills, and the ability to translate complex technical concepts to non-technical stakeholders.\n \n Nice to haves:\n \n Experience securing API access and implementing access control mechanisms at the application level.\n Multi-cloud infrastructure experience – AWS, Azure, GCP, and more.\n Proficiency in integrating IAM solutions with applications built using frameworks such as Java, Python, Node.js, or .NET.\n Mentorship/leadership experience supporting junior engineers\n Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position and may be inclusive of several career levels at Scale; it will be determined during the interview process based on work location and additional factors, including job-related skills, experience, qualifications, interview performance, and relevant education or training. Scale employees in eligible roles are also granted equity based compensation, subject to Board of Director approval. Your recruiter can share more about the specific salary range for your preferred location during the hiring process, and confirm whether the hired role will be eligible for equity grant. You'll also receive benefits including, but not limited to: comprehensive health, dental and vision coverage, retirement benefits, a learning and development stipend, and generous PTO. Additionally, this role may be eligible for additional benefits such as a commuter stipend. \n Please reference the job posting's subtitle for where this position will be located. For pay transparency purposes, the base salary range for this full-time position in the locations of San Francisco, New York, Seat","salary_min":216000,"salary_max":270000,"location":"San Francisco, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"senior","tags":["fine-tuning","distributed-systems","cloud","reinforcement-learning","generative-ai","llm","microservices","pre-training"],"apply_url":"https://job-boards.greenhouse.io/scaleai/jobs/4711898005","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-07-08T18:08:31Z","expires_at":"2026-09-28T13:31:43.946184Z","created_at":"2026-07-09T14:01:29.84661Z","updated_at":"2026-08-29T13:31:44.101607Z","company_name":"Scale AI","company_slug":"scale-ai","company_logo_url":"https://www.google.com/s2/favicons?domain=scale.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/5b4c2841-d819-4ab1-87e4-988c9bff0235"},{"id":"f0473c5a-729d-405b-813d-b6c60cb2c44a","company_id":"a0000000-0000-0000-0000-000000000001","title":"Research Engineer, Life Sciences","slug":"research-engineer-life-sciences-531dd7ae","description":"About Anthropic \n Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.\n About the Role \n We're seeking an exceptional Research Engineer to join our Life Sciences team at Anthropic. Our team is organized around the north star goal of accelerating progress in the life sciences, from early discovery through translation, by an order of magnitude. Our team likes to think across the whole model stack. In this role, you'll combine your deep expertise in machine learning engineering to develop novel evaluation frameworks and training strategies that push the frontier of what AI can achieve in biology.\n You'll work at the intersection of cutting-edge AI and the biological sciences, developing rigorous methods to measure and improve model performance on complex scientific tasks. You'll collaborate closely with world-class researchers and engineers to build AI systems that can engage in all phases of research and development, while maintaining our commitment to safety and beneficial impact.\n Previous experience in life sciences is welcome, but not required for this role.\n Minimum Qualifications \n \n Demonstrated experience training and evaluating large language models\n Proficiency in Python and familiarity with modern ML development practices\n Experience building and managing data pipelines for large-scale datasets\n Comfortable navigating ambiguity and developing solutions in rapidly evolving research environments\n Strong written and verbal communication skills, with the ability to work independently while collaborating effectively across cross-functional teams\n \n Preferred Qualifications \n \n 8+ years of machine learning experience\n Prior work experience in AI and biology, including graduate studies (molecular biology, biochemistry, computational biology, or related fields)\n Experience working with large-scale biological datasets\n Published research or practical experience in scientific AI applications or long-horizon reasoning\n Background in reinforcement learning and/or pretraining\n Knowledge of containerization technologies (e.g., Docker, Kubernetes) and cloud deployment at scale\n Demonstrated ability to work across multiple domains, such as language modeling, systems engineering, and scientific computing\n Contributions to open-source scientific software or databases\n The annual compensation range for this role is listed below. \n For sales roles, the range provided is the role’s On Target Earnings (\"OTE\") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role.\n Annual Salary:\n $350,000 — $500,000 USD \n Logistics \n Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience\n Required field of study:  A field relevant to the role as demonstrated through coursework, training, or professional experience\n Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position\n Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices.\n Visa sponsorship:  We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this.\n We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed.  Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work. We think AI systems like the ones we're building have enormous social and ethical implications. We think this makes representation even more important, and we strive to include a range of diverse perspectives on our team. Your safety matters to us. To protect yourself from potential scams, remember that Anthropic recruiters only contact you from @anthropic.com email addresses. In some cases, we may partner with vetted recruiting agencies who will identify themselves as working on behalf of Anthropic. Be cautious of emails from other domains. Legitimate Anthropic recruiters will never ask for money, fees, or banking information before your first day. If you're ever unsure about a communication, don't click any links—visit  anthropic.com/careers  directly for confirmed position openings.\n How we're different \n We bel","salary_min":350000,"salary_max":500000,"location":"San Francisco, CA","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["pre-training","llm","reinforcement-learning","search","data-pipeline","alignment","research"],"apply_url":"https://job-boards.greenhouse.io/anthropic/jobs/5265365008","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-07-03T14:01:13Z","expires_at":"2026-09-28T13:30:29.801342Z","created_at":"2026-07-04T14:00:23.291795Z","updated_at":"2026-08-29T13:30:29.948469Z","company_name":"Anthropic","company_slug":"anthropic","company_logo_url":"https://www.google.com/s2/favicons?domain=anthropic.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/f0473c5a-729d-405b-813d-b6c60cb2c44a"},{"id":"7076d892-7466-4fa7-947a-419a9fd28340","company_id":"a0000000-0000-0000-0000-000000000003","title":"Software Engineer, Identity","slug":"software-engineer-identity-c818ccb1","description":"Software is eating the world, but AI is eating software. We live in unprecedented times – AI has the potential to exponentially augment human intelligence. Every person will have a personal tutor, coach, assistant, personal shopper, travel guide, and therapist throughout life. As the world adjusts to this new reality, leading platform companies are scrambling to build LLMs at billion scale, while large enterprises figure out how to add it to their products. To make them safe, aligned and actually useful, these models need human eval and reinforcement learning through human feedback (RLHF) during pre-training, fine-tuning, and production evaluations. This is the main innovation that’s enabled ChatGPT to get such a large headstart among competition.\n At Scale, our products include the Generative AI Data Engine, SGP, Donovan, and others that power the most advanced LLMs and generative models in the world through world-class RLHF, human data generation, model evaluation, safety, and alignment. The data we are producing is some of the most important work for how humanity will interact with AI.\n At the foundation of these products is the Identity  Engineering team.  In this role, you will help support the design and development of core software systems specifically focused on identity, access management, authorization, and authentication.  You’ll also get widespread exposure to the forefront of the AI race as Scale sees it in enterprises, startups, governments, and large tech companies.\n You will:\n \n Drive the design, and implementation of our identity infrastructure to ensure secure authentication and authorization across enterprise systems.\n Build software for authentication mechanisms such as Single Sign-On (SSO), Multi-Factor Authentication (MFA), and federated identity solutions (SAML, OAuth, OpenID Connect).\n Build software for authorization mechanisms such as Relation-based access control (ReBAC), Attribute-based access control (ABAC), Role-based access control (RBAC).\n Build software-defined identity governance policies to ensure compliance with security policies, industry regulations (e.g., NIST, SOC2, ISO 27001), and organizational standards.\n Present technical information to teams and stakeholders, providing guidance and insight on identity management and best practices.\n \n Ideally you’d have:\n \n 3+ years of full-time engineering experience, post-graduation with specialities in infrastructure and identity systems.\n Infrastructure expertise – IAM controls, Infrastructure as Code (Terraform, Pulumi), microservice deployment best practices.\n Hands-on experience working with OpenFGA, Authzed, Cedar, Topaz, or similar authorization frameworks at scale.\n Strong understanding of Zanzibar-based ReBAC models, relationship tuples, and access control evaluation.\n Strong knowledge of authentication standards such as OAuth 2.0, OIDC, SAML, and JWT, as well as industry standard IdP solutions like EntraID, Okta, etc.\n Extensive experience in software development and a deep understanding of distributed systems and public cloud platforms (AWS preferred).\n Show a track record of independent ownership of successful engineering projects.\n Possess excellent communication and collaboration skills, and the ability to translate complex technical concepts to non-technical stakeholders.\n \n Nice to haves:\n \n Experience securing API access and implementing access control mechanisms at the application level.\n Multi-cloud infrastructure experience – AWS, Azure, GCP, and more.\n Proficiency in integrating IAM solutions with applications built using frameworks such as Java, Python, Node.js, or .NET.\n Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position and may be inclusive of several career levels at Scale; it will be determined during the interview process based on work location and additional factors, including job-related skills, experience, qualifications, interview performance, and relevant education or training. Scale employees in eligible roles are also granted equity based compensation, subject to Board of Director approval. Your recruiter can share more about the specific salary range for your preferred location during the hiring process, and confirm whether the hired role will be eligible for equity grant. You'll also receive benefits including, but not limited to: comprehensive health, dental and vision coverage, retirement benefits, a learning and development stipend, and generous PTO. Additionally, this role may be eligible for additional benefits such as a commuter stipend. \n Please reference the job posting's subtitle for where this position will be located. For pay transparency purposes, the base salary range for this full-time position in the locations of San Francisco, New York, Seattle is:\n $180,000 — $225,000 USD \n PLEASE NOTE:  Our policy","salary_min":180000,"salary_max":225000,"location":"San Francisco, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"senior","tags":["fine-tuning","cloud","llm","distributed-systems","pre-training","generative-ai","reinforcement-learning","microservices"],"apply_url":"https://job-boards.greenhouse.io/scaleai/jobs/4710484005","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-06-30T20:51:01Z","expires_at":"2026-09-28T13:31:45.205294Z","created_at":"2026-07-01T14:01:23.481984Z","updated_at":"2026-08-29T13:31:45.357166Z","company_name":"Scale AI","company_slug":"scale-ai","company_logo_url":"https://www.google.com/s2/favicons?domain=scale.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/7076d892-7466-4fa7-947a-419a9fd28340"},{"id":"3ff0bc99-7ec1-4627-96ef-86e258203813","company_id":"332b7698-676b-4a3e-8b02-81b1195c5af6","title":"Staff Software Engineer, AI Runtime","slug":"staff-software-engineer-ai-runtime-2e3bacea","description":"P-1930 At Databricks, we are passionate about enabling data teams to solve the world's toughest problems — from making the next mode of transportation a reality to accelerating the development of medical breakthroughs. We do this by building and running the world's best data and AI infrastructure platform so our customers can use deep data insights to improve their business.\n Training and customizing state-of-the-art AI models is one of the most demanding workloads in computing, and it sits at the heart of Databricks' Mosaic AI mission. AI Runtime (AIR) is our managed platform for large-scale GPU training and fine-tuning. It gives customers on-demand access to fleets of the latest accelerators and a serverless experience that hides the complexity of provisioning, scheduling, and orchestrating multi-node jobs, with the resilience to keep training running for days or weeks across thousands of GPUs. AIR powers the full spectrum of custom training, from fine-tuning open models to pre-training frontier-scale foundation models, for some of the most sophisticated AI teams in the world.\n As a Staff Software Engineer for AI Runtime, you will play a critical role in building and scaling the systems that make large-scale training fast, reliable, and effortless. You will drive the architecture and evolution of the managed GPU training stack, spanning scheduling and capacity, distributed training performance, fault tolerance, and the developer experience of launching and operating jobs at scale. Beyond hands-on contributions to core systems, you will help define the long-term technical vision for AIR, mentor senior engineers, partner across product, research, and platform teams, and lead the initiatives that expand the technical and business impact of custom training at Databricks.\n The impact you will have: \n \n Drive the architecture and evolution of AIR's managed GPU training platform, delivering scalable, high-throughput, and resilient training across fleets that span thousands of accelerators.\n Solve the hardest problems in large-scale training, including multi-node orchestration, distributed parallelism strategies, GPU scheduling and dynamic routing, high-throughput data loading, and checkpoint and restore for very long-running jobs.\n Push GPU efficiency and training performance, raising utilization (such as model FLOPs utilization and end-to-end throughput) and lowering cost per training run across diverse model architectures and hardware generations.\n Build the resilience and observability foundations that keep multi-node jobs healthy, detecting and recovering from hardware and software failures with minimal disruption to customers.\n Partner with product, research, and platform teams to shape the APIs, CLI, and developer experience that make it easy to launch, monitor, and debug production training jobs.\n Lead end-to-end engineering efforts, from design through production rollout, holding a high bar for performance, correctness, and reliability.\n Make direct, high-impact contributions to the core systems behind AIR, and help bring up support for the latest accelerators and new regions as the fleet grows.\n Champion engineering excellence, mentor other engineers through design reviews and technical discussions, and help shape Databricks' long-term technical direction in AI training infrastructure.\n \n  \n What we look for: \n \n 10+ years of experience building and operating large-scale distributed systems, with significant depth in GPU training infrastructure, high-performance computing, or ML systems.\n Hands-on experience with distributed training frameworks (such as PyTorch, FSDP, DeepSpeed, or Megatron) and the parallelism strategies (data, tensor, pipeline, and sequence parallelism) used to train large models.\n Strong understanding of training resilience patterns, including checkpointing, failure detection, and automatic recovery for long-running, multi-node jobs.\n Solid grasp of GPU performance fundamentals, including accelerator architecture, high-speed interconnects (such as NVLink and InfiniBand or RoCE), collective communication, and the bottlenecks that govern training throughput and utilization.\n Experience building and operating managed, multi-tenant platform products in the cloud, with clear SLAs and SLOs for availability, performance, and reliability.\n Strong foundation in algorithms, data structures, and system design as applied to performance-sensitive, large-scale distributed systems.\n Proven ability to deliver technically complex, high-impact initiatives that create clear customer or business value.\n Strong communication skills and the ability to collaborate across product, research, and infrastructure teams in a fast-moving environment.\n Strategic, product-oriented mindset with the ability to align technical execution to a long-term vision, and a passion for mentoring engineers and fostering technical excellence.\n BS in Computer Science or a related field (MS or PhD preferred).\n \n  \n  \n Pay Ran","salary_min":190000,"salary_max":265000,"location":"Mountain View, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["distributed-systems","pytorch","fine-tuning","pre-training","generative-ai"],"apply_url":"https://databricks.com/company/careers/open-positions/job?gh_jid=8582271002","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-06-08T21:32:45Z","expires_at":"2026-09-28T13:32:31.098853Z","created_at":"2026-06-28T14:02:15.882818Z","updated_at":"2026-08-29T13:32:31.252004Z","company_name":"Databricks","company_slug":"databricks","company_logo_url":"https://www.google.com/s2/favicons?domain=databricks.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/3ff0bc99-7ec1-4627-96ef-86e258203813"},{"id":"9726ecea-9913-45ac-9639-e77847699045","company_id":"332b7698-676b-4a3e-8b02-81b1195c5af6","title":"Senior Software Engineer, AI Runtime","slug":"senior-software-engineer-ai-runtime-216c21ab","description":"P-1428\n At Databricks, we are passionate about enabling data teams to solve the world's toughest problems — from making the next mode of transportation a reality to accelerating the development of medical breakthroughs. We do this by building and running the world's best data and AI infrastructure platform so our customers can use deep data insights to improve their business.\n Training and customizing state-of-the-art AI models is one of the most demanding workloads in computing, and it sits at the heart of Databricks' Mosaic AI mission. AI Runtime (AIR) is our managed platform for large-scale GPU training and fine-tuning. It gives customers on-demand access to fleets of the latest accelerators and a serverless experience that hides the complexity of provisioning, scheduling, and orchestrating multi-node jobs, with the resilience to keep training running for days or weeks across thousands of GPUs. AIR powers the full spectrum of custom training, from fine-tuning open models to pre-training frontier-scale foundation models, for some of the most sophisticated AI teams in the world.\n As a Senior Software Engineer for AI Runtime, you will play a critical role in building and scaling the systems that make large-scale training fast, reliable, and effortless. You will drive the architecture and evolution of the managed GPU training stack, spanning scheduling and capacity, distributed training performance, fault tolerance, and the developer experience of launching and operating jobs at scale. Beyond hands-on contributions to core systems, you will help shape the technical direction for AIR, mentor other engineers, partner across product, research, and platform teams, and contribute to the initiatives that expand the technical and business impact of custom training at Databricks.\n The impact you will have: \n \n Drive the architecture and evolution of AIR's managed GPU training platform, delivering scalable, high-throughput, and resilient training across fleets that span thousands of accelerators.\n Solve the hardest problems in large-scale training, including multi-node orchestration, distributed parallelism strategies, GPU scheduling and dynamic routing, high-throughput data loading, and checkpoint and restore for very long-running jobs.\n Push GPU efficiency and training performance, raising utilization (such as model FLOPs utilization and end-to-end throughput) and lowering cost per training run across diverse model architectures and hardware generations.\n Build the resilience and observability foundations that keep multi-node jobs healthy, detecting and recovering from hardware and software failures with minimal disruption to customers.\n Partner with product, research, and platform teams to shape the APIs, CLI, and developer experience that make it easy to launch, monitor, and debug production training jobs.\n Lead end-to-end engineering efforts, from design through production rollout, holding a high bar for performance, correctness, and reliability.\n Make direct, high-impact contributions to the core systems behind AIR, and help bring up support for the latest accelerators and new regions as the fleet grows.\n Champion engineering excellence, mentor other engineers through design reviews and technical discussions, and contribute to Databricks' technical direction in AI training infrastructure.\n \n  \n What we look for: \n \n 5+ years of experience building and operating large-scale distributed systems, with experience in GPU training infrastructure, high-performance computing, or ML systems.\n Experience with distributed training frameworks (such as PyTorch, FSDP, DeepSpeed, or Megatron) and the parallelism strategies (data, tensor, pipeline, and sequence parallelism) used to train large models.\n Strong understanding of training resilience patterns, including checkpointing, failure detection, and automatic recovery for long-running, multi-node jobs.\n Solid grasp of GPU performance fundamentals, including accelerator architecture, high-speed interconnects (such as NVLink and InfiniBand or RoCE), collective communication, and the bottlenecks that govern training throughput and utilization.\n Experience building and operating managed, multi-tenant platform products in the cloud, with clear SLAs and SLOs for availability, performance, and reliability.\n Strong foundation in algorithms, data structures, and system design as applied to performance-sensitive, large-scale distributed systems.\n Proven ability to deliver technically complex, high-impact initiatives that create clear customer or business value.\n Strong communication skills and the ability to collaborate across product, research, and infrastructure teams in a fast-moving environment.\n Customer-focused mindset with the ability to align implementation details with product goals, and a passion for mentoring engineers and fostering technical excellence.\n BS in Computer Science or a related field (MS or PhD preferred).\n  \n Pay Range Transparency \n Databricks is committ","salary_min":160000,"salary_max":225000,"location":"Mountain View, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"senior","tags":["pytorch","fine-tuning","pre-training","generative-ai","distributed-systems"],"apply_url":"https://databricks.com/company/careers/open-positions/job?gh_jid=8582276002","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-06-08T21:32:38Z","expires_at":"2026-09-28T13:32:26.818912Z","created_at":"2026-06-28T14:02:12.287649Z","updated_at":"2026-08-29T13:32:26.977876Z","company_name":"Databricks","company_slug":"databricks","company_logo_url":"https://www.google.com/s2/favicons?domain=databricks.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/9726ecea-9913-45ac-9639-e77847699045"},{"id":"4a0ffe7d-807a-4b2f-adf3-b421a54fbf3a","company_id":"3c528ec7-088e-499c-b3ea-9926667c7188","title":"Senior Research Scientist, Machine Learning (BioFM)","slug":"senior-research-scientist-machine-learning-biofm-c9f00afe","description":"About Us\nDeep Genomics is at the forefront of using artificial intelligence to transform drug discovery. Our proprietary AI platform decodes the complexity of RNA biology to identify novel drug targets, mechanisms, and therapeutics inaccessible through traditional methods. With expertise spanning machine learning, bioinformatics, data science, engineering, and drug development, our multidisciplinary team in Toronto and Cambridge, MA is revolutionizing how new medicines are created.\nOpportunity\nWe are seeking an exceptional and creative Senior/Staff Machine Learning Scientist to lead and innovate within our core AI research team, specifically focusing on the creative building of Biological Foundation Models (BioFMs). You will pioneer novel deep learning architectures and pre-training paradigms that learn the fundamental language of the genome and cellular biology. Rather than just applying out-of-the-box ML to biological datasets, you will design the next generation of BioFMs from tackling complex -omics data at scale. If you are a first-principles thinker excited to bridge advanced ML with genome biology to solve high-impact, frontier problems in human health and drug discovery, this is a unique opportunity.\n\n\n","salary_min":175000,"salary_max":200000,"location":"Toronto, Canada","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"senior","tags":["pre-training","generative-ai","deep-learning","machine-learning","research"],"apply_url":"https://jobs.lever.co/deepgenomics/74978439-123f-4000-ae76-e70730c6cfa0/apply","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-06-08T20:39:58.761Z","expires_at":"2026-09-28T13:42:59.475885Z","created_at":"2026-06-28T14:11:15.343855Z","updated_at":"2026-08-29T13:42:59.635893Z","company_name":"Deep Genomics","company_slug":"deep-genomics","company_logo_url":"https://www.google.com/s2/favicons?domain=deepgenomics.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/4a0ffe7d-807a-4b2f-adf3-b421a54fbf3a"}],"page":1,"per_page":20,"total":142,"total_is_exact":true,"total_pages":8}
