{"access":{"catalog_url":"https://aidevboard.com/api/v1/catalog","description":"Public read endpoints are open and free. API keys are optional for stable agent identity and keyed hourly throttling.","docs_url":"https://aidevboard.com/docs","employer_pilot_url":"https://aidevboard.com/verified-interview-pilot","mode":"open","register_url":"https://aidevboard.com/api/v1/register"},"candidate_resume_action":{"application_authorized":false,"candidate_charge":0,"endpoint":"https://aidevboard.com/api/v1/candidate/resume-preview","job_id_json_path":"jobs[].id","method":"POST","preview_requires_identity":false,"required_body_fields":["job_id","evidence_bullets"],"requires_explicit_human_review":true,"saved_artifact_protocol":"mcp","saved_artifact_requires_verified_human":true,"saved_artifact_tool":"compile_job_specific_resume","search_requires_identity":false,"status":"available_after_candidate_selects_job","submission_performed":false,"uses_candidate_verified_evidence":true},"degraded":false,"estimated":false,"has_next":true,"jobs":[{"id":"55f8d3f4-58f6-4ca2-9b49-1e84deeaec13","company_id":"e8c9f3a5-9310-43f5-9341-321fe6d93a92","title":"Triage Automation Engineer","slug":"triage-automation-engineer-90734211","description":"About us    \n Founded in 2017, Wayve is the leading developer of Embodied AI technology.  Our advanced AI software and foundation models enable vehicles to perceive, understand, and navigate any complex environment, enhancing the usability and safety of automated driving systems.\n Our vision is to create autonomy that propels the world forward.  Our intelligent, mapless, and hardware-agnostic AI products are designed for automakers, accelerating the transition from assisted to automated driving.  In our fast-paced environment big problems ignite us—we embrace uncertainty, leaning into complex challenges to unlock groundbreaking solutions. We aim high and stay humble in our pursuit of excellence, constantly learning and evolving as we pave the way for a smarter, safer future.\n At Wayve, your contributions matter.  We value diversity, embrace new perspectives, and foster an inclusive work environment; we back each other to deliver impact.  \n Make Wayve the experience that defines your career!  \n The role  \n As a Triage Automation Engineer at Wayve , you'll play a key role in scaling our Triage and Detectives workflow through process improvement, tooling, and automation. You'll partner with Triage Specialists and Detective Engineers to turn repeated, manual analysis into reliable, productised automation, including triage bots, behavioural classifiers, and workflow tooling. That reduces manual load and speeds up how quickly we identify, understand, and resolve issues across our test and on-road fleets.\n This is a highly collaborative, detail-oriented role with a direct impact on the safety and efficiency of Wayve's development pipeline.\n Key responsibilities:\n \n Document and measure existing Triage and Detective workflows to identify opportunities for automation and process improvement\n Prioritise improvements based on cost, benefit, impact, and feasibility\n Partner with Detective Engineers to integrate their scripts and tools into scalable, productised workflows\n Work with development, ML, and AI teams to deliver the tooling improvements Triage needs\n Build automation pipelines — including triage bots and behavioural classifiers — that reduce manual load on Triage Specialists\n Validate automation changes and outputs, including the accuracy and reliability of classifications and suggested root causes\n Document newly implemented automations: what they do, how they work, how to use them, known limitations, and expected outputs\n Measure and report on triage quality and throughput to track the impact of automation\n Collaborate with senior management and cross-functional stakeholders to shape the roadmap for business-critical automation\n Willingness to travel domestically and internationally (including trips to our London office)\n \n About you   \n In order to set you up for success as a Triage Automation Engineer at Wayve, we’re looking for the following skills and experience.  \n Essential \n \n 3+ years of experience working with complex systems, ideally within robotics or autonomous vehicles\n Strong scripting and analytical skills (e.g. Python, SQL, Bash/Shell)\n Hands-on experience operating in a remote Linux environment\n Experience building or maintaining data pipelines or notebooks (e.g. Databricks, Jupyter)\n Great communication skills, able to explain complex technical problems to both technical and non-technical stakeholders\n Expertise using issue tracking and configuration management tools such as Jira, Confluence, and Bitbucket/GitLab\n Comfort with ambiguity — able to measure an existing workflow, identify where automation adds value, and scope a sensible solution\n \n Desirable \n \n Experience with web development languages (e.g. HTML, CSS, React, Java) for building internal tooling\n Practical experience with machine learning or classification models (e.g. PyTorch)\n Experience with cloud services (ideally Microsoft Azure)\n Passion for taking research ideas to production\n Track record of promoting statistical rigour and experimental best practice\n Experience working in a fast-moving tech company or startup\n \n This is a full-time role based in our office in Sunnyvale.  At Wayve we want the best of all worlds so we operate a hybrid working policy that combines time together in our offices and workshops to fuel innovation, culture, relationships and learning, and time spent working from home. The reasonably estimated salary for this role ranges from $144,500–$183,200, plus a competitive equity package. Actual compensation is based on the candidate's skills, qualifications, and experience.\n  \n Wayve is committed to creating an inclusive interview experience. If you require any accommodations or adjustments to participate fully in our interview process, please let us know. \n We understand that everyone has a unique set of skills and experiences and that not everyone will meet all of the requirements listed above. If you’re passionate about self-driving cars and think you have what it t","salary_min":144500,"salary_max":183200,"location":"Sunnyvale, CA","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"mid","tags":["autonomous-vehicles","generative-ai","robotics","data-pipeline","pytorch","evaluation"],"apply_url":"https://wayve.firststage.co/jobs?gh_jid=8756182002","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-28T17:56:55Z","expires_at":"2026-09-29T13:43:31.521698Z","created_at":"2026-08-29T13:44:35.170073Z","updated_at":"2026-08-30T13:43:31.651707Z","company_name":"Wayve","company_slug":"wayve","company_logo_url":"https://www.google.com/s2/favicons?domain=wayve.ai\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/55f8d3f4-58f6-4ca2-9b49-1e84deeaec13"},{"id":"a60887bd-18b6-4819-b8ac-a8ce688f7d3f","company_id":"a0000000-0000-0000-0000-000000000003","title":"Machine Learning Research Scientist, Evaluations","slug":"machine-learning-research-scientist-evaluations-47ca5c35","description":"Scale works with the industry's leading AI labs to provide high quality data and accelerate progress in GenAI research. We are looking for Research Scientists and Research Engineers with expertise in LLM post-training (SFT, RLHF, reward modeling) and evaluation. This role is on the evaluation pod within the GenAI Research Organization and will focus on building benchmarks and diagnosing model failure modes in both text and multimodal modalities.\n In this role, you will develop rigorous evaluations and diagnostic methods that reveal where frontier models fail and why. You will collaborate with researchers and engineers to define best practices in evaluation-driven AI development. You will also partner with top foundation model labs to translate failure analysis into technical and strategic input on the next generation of generative AI models.\n You will: \n \n Analyze model behavior to identify, characterize, and diagnose failure modes in frontier LLMs and Agents.  You’ll identify everything from capability gaps and reasoning errors to robustness and alignment issues, all focusing on RCA.\n Design and build benchmarks and evaluation methods that measure LLM capabilities in both text and multimodal modalities.\n Apply post-training expertise (SFT, RLHF, reward modeling) to connect observed failures to the data and training interventions that address them.\n Publish research findings in top-tier AI conferences.\n \n Ideally you’d have: \n \n Ph.D. or Master's degree in Computer Science, Machine Learning, AI, or a related field.\n Deep understanding of deep learning, reinforcement learning, and large-scale model fine-tuning.\n Experience with post-training techniques such as RLHF, preference modeling, or instruction tuning, and with LLM evaluation or benchmark development.\n Excellent written and verbal communication skills.\n Published research in areas of machine learning at major conferences (NeurIPS, ICML, ICLR, ACL, EMNLP, CVPR, etc.) and/or journals.\n Previous experience in a customer facing role.\n Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position and may be inclusive of several career levels at Scale; it will be determined during the interview process based on work location and additional factors, including job-related skills, experience, qualifications, interview performance, and relevant education or training. Scale employees in eligible roles are also granted equity based compensation, subject to Board of Director approval. Your recruiter can share more about the specific salary range for your preferred location during the hiring process, and confirm whether the hired role will be eligible for equity grant. You'll also receive benefits including, but not limited to: comprehensive health, dental and vision coverage, retirement benefits, a learning and development stipend, and generous PTO. Additionally, this role may be eligible for additional benefits such as a commuter stipend. \n Please reference the job posting's subtitle for where this position will be located. For pay transparency purposes, the base salary range for this full-time position in the locations of San Francisco, New York, Seattle is:\n $180,600 — $225,750 USD \n PLEASE NOTE:  Our policy requires a 90-day waiting period before reconsidering candidates for the same role. This allows us to ensure a fair and thorough evaluation of all applicants. \n About Us: \n At Scale, our mission is to develop reliable AI systems for the world's most important decisions. Our products provide the high-quality data and full-stack technologies that power the world's leading models, and help enterprises and governments build, deploy, and oversee AI applications that deliver real impact. We work closely with industry leaders like Meta, Ernst \u0026 Young, Mayo Clinic, Time Inc., the Government of Qatar, and U.S. government agencies including the Army and Air Force. We are expanding our team to accelerate the development of AI applications. \n We believe that everyone should be able to bring their whole selves to work, which is why we are proud to be an inclusive and equal opportunity workplace. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability status, gender identity or Veteran status.  \n We are committed to working with and providing reasonable accommodations to applicants with physical and mental disabilities. If you need assistance and/or a reasonable accommodation in the application or recruiting process due to a disability, please contact us at accommodations@scale.com. Please see the United States Department of Labor's Know Your Rights poster for additional information. \n We comply with the United States Department of Labor's Pay Transparency provision .  \n PLEASE NOTE: We co","salary_min":180600,"salary_max":225750,"location":"San Francisco, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"senior","tags":["deep-learning","fine-tuning","generative-ai","reinforcement-learning","search","nlp","llm","evaluation"],"apply_url":"https://job-boards.greenhouse.io/scaleai/jobs/4728014005","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-26T18:51:42Z","expires_at":"2026-09-29T13:31:40.355249Z","created_at":"2026-08-27T13:31:38.7307Z","updated_at":"2026-08-30T13:31:40.501539Z","company_name":"Scale AI","company_slug":"scale-ai","company_logo_url":"https://www.google.com/s2/favicons?domain=scale.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/a60887bd-18b6-4819-b8ac-a8ce688f7d3f"},{"id":"2af68795-5861-40b1-95ce-04d100978f08","company_id":"e8c9f3a5-9310-43f5-9341-321fe6d93a92","title":"Senior Data Scientist","slug":"senior-data-scientist-e0977b92","description":"About us    \n Founded in 2017, Wayve is the leading developer of Embodied AI technology.  Our advanced AI software and foundation models enable vehicles to perceive, understand, and navigate any complex environment, enhancing the usability and safety of automated driving systems.\n Our vision is to create autonomy that propels the world forward.  Our intelligent, mapless, and hardware-agnostic AI products are designed for automakers, accelerating the transition from assisted to automated driving.  In our fast-paced environment big problems ignite us—we embrace uncertainty, leaning into complex challenges to unlock groundbreaking solutions. We aim high and stay humble in our pursuit of excellence, constantly learning and evolving as we pave the way for a smarter, safer future.\n At Wayve, your contributions matter.  We value diversity, embrace new perspectives, and foster an inclusive work environment; we back each other to deliver impact.  \n Make Wayve the experience that defines your career!  \n The Role\n As a Data Scientist supporting AI engineers, you will partner with one or more engineering teams, developing actionable insights that guide improvements to the Wayve AI Driver. Using experimental and observational analyses of real and simulated driving, you will help teams advance the functionality, safety, and performance of the Wayve AI Driver, helping to advance Wayve as the leader in end-to-end AI for autonomous mobility.\n This means you might:\n \n Formulate and iterate upon the performance metrics that organize our engineering efforts and guide progress toward commercial success\n Design experiments and targeted off-road measurements to ensure that we deliver product requirements to customers while maintaining safety and performance\n Investigate factors in model training and inference leading to bottlenecks in functionality and performance, identifying and validating hypotheses for unlocking improvements\n \n About you \n Essential:\n \n 3+ years experience working in a Data Science role.\n Fluent in querying and building large datasets, writing production-level SQL for use in data-transformation pipelines.\n Prior experience designing robust real-world experiments (e.g. A/B) and critically evaluating test-statistics\n Foundations in the fundamentals behind statistics: testing appropriate distributions, testing the assumptions behind frequentist stats\n Proficient in using a statistical scripting language and data science/ML packages (e.g. python such as pandas, sklearn, statsmodels, scipy or R such as dplyr, caret, stats)\n Well-versed in summarising, visualising and communicating findings in an accessible and compelling way\n Track record of influencing team direction through your findings\n A bias towards deriving actionable insight that can be used to drive prioritisation and strategy for others.\n Comfortable working asynchronously across time zones with cross-functional partners\n You are deeply curious about building something new and relish the idea of helping to define AV2.0 and how we build it.\n \n Desirable:\n \n Practical experience with machine learning (e.g. PyTorch). Passion to take research ideas to production.\n Track record of promoting statistical rigour and experimental best practices in your prior roles.\n Prior experience using causal inference/econometric techniques and bayesian methodologies for hypothesis testing.\n Prior experience using large datasets with distributed computing (e.g. spark, hadoop or other map-reduce tech)\n Experience working in a fast-moving tech company or startup.\n \n This role is a full-time role based in Sunnyvale, CA (hybrid) and the reasonably estimated salary for this role ranges from $209,700 to 266,800, plus a competitive equity package. Actual compensation is based on the candidate's skills, qualifications, and experience. At Wayve we want the best of all worlds so we operate a hybrid working policy that combines time together in our offices and workshops to fuel innovation, culture, relationships and learning, and time spent working from home.   We operate core working hours so you can determine the schedule that works best for you and your team. \n Wayve is committed to creating an inclusive interview experience. If you require any accommodations or adjustments to participate fully in our interview process, please let us know. \n We understand that everyone has a unique set of skills and experiences and that not everyone will meet all of the requirements listed above. If you’re passionate about self-driving cars and think you have what it takes to make a positive impact on the world, we encourage you to apply. At Wayve we're committed to creating a diverse, fair and respectful culture that is inclusive of everyone based on their unique skills and perspectives, and regardless of sex, race, religion or belief, ethnic or national origin, disability, age, citizenship, marital, domestic or civil partnership status, sexual orientation, gender identity, veteran ","salary_min":209700,"salary_max":266800,"location":"Sunnyvale, CA","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"senior","tags":["autonomous-vehicles","distributed-systems","pytorch","generative-ai","data-science","evaluation"],"apply_url":"https://wayve.firststage.co/jobs?gh_jid=8728411002","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-20T20:06:09Z","expires_at":"2026-09-29T13:43:27.258396Z","created_at":"2026-08-25T18:31:14.400233Z","updated_at":"2026-08-30T13:43:27.392293Z","company_name":"Wayve","company_slug":"wayve","company_logo_url":"https://www.google.com/s2/favicons?domain=wayve.ai\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/2af68795-5861-40b1-95ce-04d100978f08"},{"id":"fbd953c7-087c-41ea-bca5-24b67b33e933","company_id":"e8c9f3a5-9310-43f5-9341-321fe6d93a92","title":"Full Stack Software Engineer, Evaluation Tools","slug":"full-stack-software-engineer-evaluation-tools-055de53a","description":"About us    \n Founded in 2017, Wayve is the leading developer of Embodied AI technology.  Our advanced AI software and foundation models enable vehicles to perceive, understand, and navigate any complex environment, enhancing the usability and safety of automated driving systems.\n Our vision is to create autonomy that propels the world forward.  Our intelligent, mapless, and hardware-agnostic AI products are designed for automakers, accelerating the transition from assisted to automated driving.  In our fast-paced environment big problems ignite us—we embrace uncertainty, leaning into complex challenges to unlock groundbreaking solutions. We aim high and stay humble in our pursuit of excellence, constantly learning and evolving as we pave the way for a smarter, safer future.\n At Wayve, your contributions matter.  We value diversity, embrace new perspectives, and foster an inclusive work environment; we back each other to deliver impact.  \n Make Wayve the experience that defines your career!  \n The role\n The Evaluation Tools group builds the internal products that accelerate the full AI Driver development loop, from defining a test to debugging model behaviour. Model developers, researchers and QA engineers across Wayve depend on our tools to understand driving performance and scale evaluation to millions of scenarios, and every major model release runs through them.\n You'll join the Search \u0026 Agents squad in Sunnyvale. We build the search, scenario mining and agentic tooling that lets anyone find the right driving scenarios and turn them into tests without writing SQL. Our work is the entry point to a fully agentic development loop: identify an issue → mine for scenarios → build and run a test suite → root-cause the failure → retrain → repeat.\n You'll work across the stack and own features end-to-end: talking to users, shaping ideas, building robust software, and validating impact. Evaluation is an evolving challenge, so you'll also have plenty of opportunity to define new projects as user needs emerge.\n Why this role matters \n \n Make natural language scenario search reliable and self-service, so finding test data no longer depends on scarce SQL and mining expertise\n Take scenario mining from proof-of-concept to production scale, unlocking certification-grade test creation for teams like Validation\n Extend the evaluation MCP so agents can run search-to-test workflows end to end\n Build trust in search and mining results, so teams can make fast, confident decisions grounded in clear evidence\n \n Challenges you will own\n \n Build fullstack features across search, scenario mining, agentic tooling, and the web applications our users work in to curate scenarios and build test suites\n Work directly with users to understand pain points, then define and measure what success looks like (adoption, time saved, reliability) before you build\n Ship high-quality software with strong attention to reliability, performance and maintainability\n Contribute to architectural decisions that scale across the team’s products\n Use AI tools to improve how we build, from speeding up development to refining team workflows and product quality\n \n About you\n In order to set you up for success as a Fullstack Software Engineer at Wayve, we're looking for the following skills and experience.\n Essential \n \n Strong development skills in Python, TypeScript, and JavaScript, with experience using React or similar front-end frameworks\n Strong SQL skills with a solid understanding of database design and optimisation\n Experience designing and building production-level fullstack systems and reliable, high-performance APIs\n Track record of building robust, maintainable systems and applying engineering best practices\n Excellent communication and collaboration skills, including working effectively across time zones\n Comfortable working independently in a fast-paced, high-context, ambiguous environment\n You are a power user of AI development tools and have a mindset for building AI-augmented workflows\n \n Desirable \n \n Experience with search, embeddings or retrieval systems\n Experience with the Databricks platform, and large-scale data processing with Spark or similar\n Experience with job orchestration frameworks such as Flyte\n Experience building agentic or LLM-powered developer tools (e.g. MCP servers)\n Experience building tools for technical users (ML, robotics, infrastructure)\n Experience with scientific data visualisation or complex data interfaces This role is a full-time role based in Sunnyvale, CA (hybrid) and the reasonably estimated salary for this role ranges from $209,000 to 266,000, plus a competitive equity package. Actual compensation is based on the candidate's skills, qualifications, and experience. At Wayve we want the best of all worlds so we operate a hybrid working policy that combines time together in our offices and workshops to fuel innovation, culture, relationships and learning, and time spent working from ho","salary_min":209000,"salary_max":266000,"location":"Sunnyvale, CA","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"senior","tags":["llm","robotics","agents","generative-ai","autonomous-vehicles","evaluation","fullstack"],"apply_url":"https://wayve.firststage.co/jobs?gh_jid=8733546002","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-20T16:28:42Z","expires_at":"2026-09-29T13:43:23.857644Z","created_at":"2026-08-25T18:31:14.236069Z","updated_at":"2026-08-30T13:43:23.988134Z","company_name":"Wayve","company_slug":"wayve","company_logo_url":"https://www.google.com/s2/favicons?domain=wayve.ai\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/fbd953c7-087c-41ea-bca5-24b67b33e933"},{"id":"eaacbad1-cf06-4872-92b7-f9c1caa6a5fd","company_id":"10c1ac82-83d5-423a-b438-4cc7b13d597c","title":"Senior Frontend Engineer, AI Observability \u0026 Evals Platform ","slug":"senior-frontend-engineer-ai-observability-evals-platform-ffb86ec5","description":"ABOUT US\n\n\n\nAt LangChain, our mission is to make intelligent agents ubiquitous. We build the foundation for agent engineering in the real world, helping developers move from prototypes to production-ready AI agents that teams can rely on. We began as widely adopted open-source tools and have grown to also offer a platform for building, evaluating, deploying, and operating agents at scale.\n\nWith $125M raised at Series B from IVP, Sequoia, Benchmark, CapitalG, and Sapphire Ventures, we’re at a stage where we’re continuing to develop new products, growth is accelerating, and all team members have meaningful impact on what we build and how we work together. LangChain is a place where your contributions can shape how this technology shows up in the real world.\n\nToday, our platform includes LangSmith (Observability, Evaluation, Deployment, Fleet, and Sandboxes), our open source frameworks (LangChain, LangGraph, and Deep Agents), and the newly launched LangSmith Engine for autonomous agent improvement. We have 100M+ monthly open source downloads, 6,000+ active LangSmith customers, and 5 of the Fortune 10 use LangSmith in production (+ 35% of the Fortune 500 overall), including teams at Klarna, Clay, Coinbase, Workday, Lyft, Cloudflare, Harvey, Rippling, Vanta, LinkedIn, Monday.com, Nvidia, and Bridgewater.\n\n\n\n\n\n\nABOUT THE ROLE\n\n(In person 5 days/week in San Francisco, CA, New York, Boston )\n\nWe’re looking to bring on an experienced frontend engineer to develop and enhance new features on LangSmith, our enterprise platform product for LLM application observability, testing, and debugging.  \n\n\n\n\nWHAT YOU WILL DO\n\n - Develop new user-facing features using React \u0026 Typescript\n\n - Build reusable components and front-end libraries for future use\n\n - Translate designs and wireframes into high-quality code\n\n - Optimize components for maximum performance across a vast array of web-capable devices and browsers\n\n - Collaborate with fullstack and backend developers and UX/UI designers, to enhance usability\n   \n   \n\n\nWHAT YOU WILL BRING\n\n - Extensive front-end engineering experience with strong proficiency in React, JavaScript and TypeScript\n\n - Hands-on experience with front-end development tools like Babel, Vite, Webpack, NPM, and Yarn\n\n - Familiarity with REST APIs and experience working closely with fullstack and backend engineers\n\n - Knowledge of modern authorization mechanisms, such as JWTs\n\n - A passion for user experience, design and building clean, engaging, and intuitive interfaces\n\n - Strong written and oral communication skills, with the ability to explain technical concepts clearly and concisely to both technical and non-technical stakeholders\n\n - The DNA to thrive in a start-up, fast-paced environment. Views unstructured environments as an opportunity to figure out the most impactful work and help define the future success of the company\n\n - An ownership mind-set. You are self-driven and look for opportunities to make an impact.\n\n - Strong computer science fundamentals, ideally with a degree in computer science or related field\n\n\n\nCompensation \n\nAnnual salary range: $175,000-$240,000 USD\n\nCompensation Philosophy:\n\nWe offer competitive compensation that includes base salary, variable compensation for relevant roles, meaningful equity, benefits, and perks. Actual compensation and offerings will vary based on role, level, and location. Team members in the EU, UK, and APAC receive locally competitive benefits aligned with regional norms and regulations.\n\n\n\n\nBENEFITS\n\nBenefits include medical, dental, and vision coverage, flexible vacation, a 401(k) plan, meals on in-office days in the US and more.","salary_min":175000,"salary_max":240000,"location":"San Francisco, CA","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"senior","tags":["agents","api-design","llm","frontend","evaluation"],"apply_url":"https://jobs.ashbyhq.com/langchain/afb91b9b-46d5-4c9d-aa84-a4f1a3f74263/application","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-13T18:33:10.038Z","expires_at":"2026-09-29T13:32:13.024597Z","created_at":"2026-04-13T09:37:32.570304Z","updated_at":"2026-08-30T13:32:13.16871Z","company_name":"LangChain","company_slug":"langchain","company_logo_url":"https://www.google.com/s2/favicons?domain=langchain.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/eaacbad1-cf06-4872-92b7-f9c1caa6a5fd"},{"id":"258e08a5-3f97-4f5e-bfaf-1bcd3e0ebd55","company_id":"10c1ac82-83d5-423a-b438-4cc7b13d597c","title":"Frontend Engineer, AI Observability \u0026 Evals Platform ","slug":"frontend-engineer-ai-observability-evals-platform-c60f6f55","description":"ABOUT US\n\n\n\nAt LangChain, our mission is to make intelligent agents ubiquitous. We build the foundation for agent engineering in the real world, helping developers move from prototypes to production-ready AI agents that teams can rely on. We began as widely adopted open-source tools and have grown to also offer a platform for building, evaluating, deploying, and operating agents at scale.\n\nWith $125M raised at Series B from IVP, Sequoia, Benchmark, CapitalG, and Sapphire Ventures, we’re at a stage where we’re continuing to develop new products, growth is accelerating, and all team members have meaningful impact on what we build and how we work together. LangChain is a place where your contributions can shape how this technology shows up in the real world.\n\nToday, our platform includes LangSmith (Observability, Evaluation, Deployment, Fleet, and Sandboxes), our open source frameworks (LangChain, LangGraph, and Deep Agents), and the newly launched LangSmith Engine for autonomous agent improvement. We have 100M+ monthly open source downloads, 6,000+ active LangSmith customers, and 5 of the Fortune 10 use LangSmith in production (+ 35% of the Fortune 500 overall), including teams at Klarna, Clay, Coinbase, Workday, Lyft, Cloudflare, Harvey, Rippling, Vanta, LinkedIn, Monday.com, Nvidia, and Bridgewater.\n\n\n\n\n\nAbout The Role: \n\nWe’re looking to add a frontend engineer to develop and enhance new features on LangSmith, working on our enterprise platform product for LLM application observability, testing, and debugging.  \n\n*This role can be based in our San Francisco, New York, or Boston office. Employees within commuting distance work from the office five days per week. Candidates who live outside commuting distance (e.g. \u003e1hr each way), may be eligible for hybrid arrangements depending on location and role requirements.\n\n\nWhat You Will Do:\n\n - Develop new user-facing features using React \u0026 Typescript\n\n - Build reusable components and front-end libraries for future use\n\n - Translate designs and wireframes into high-quality code\n\n - Optimize components for maximum performance across a vast array of web-capable devices and browsers\n\n - Collaborate with fullstack and backend developers and UX/UI designers, to enhance usability\n\n\n\nWhat You Will Bring:\n\n - 2+ years of Front-end engineering experience with strong proficiency in React, JavaScript and TypeScript\n\n - Hands-on experience with front-end development tools like Babel, Vite, Webpack, NPM, and Yarn\n\n - Familiarity with REST APIs and experience working closely with fullstack and backend engineers\n\n - Knowledge of modern authorization mechanisms, such as JWTs\n\n - A passion for user experience, design and building clean, engaging, and intuitive interfaces\n\n - Strong written and oral communication skills, with the ability to explain technical concepts clearly and concisely to both technical and non-technical stakeholders\n\n - The DNA to thrive in a start-up, fast-paced environment. Views unstructured environments as an opportunity to figure out the most impactful work and help define the future success of the company\n\n - An ownership mind-set. You are self-driven and look for opportunities to make an impact.\n\n - Strong computer science fundamentals, ideally with a degree in computer science or related field\n\n\n\nCompensation\n\n - Annual salary range: $160,000-$195,000 USD\n\nCompensation Philosophy:\n\nWe offer competitive compensation that includes base salary, variable compensation for relevant roles, meaningful equity, benefits, and perks. Actual compensation and offerings will vary based on role, level, and location. Team members in the EU, UK, and APAC receive locally competitive benefits aligned with regional norms and regulations.\n\n\n\n\nBENEFITS\n\nBenefits include medical, dental, and vision coverage, flexible vacation, a 401(k) plan, meals on in-office days in the US and more.","salary_min":160000,"salary_max":195000,"location":"San Francisco, CA","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"junior","tags":["api-design","llm","agents","evaluation","frontend"],"apply_url":"https://jobs.ashbyhq.com/langchain/f7de4819-e7aa-4dfb-9acd-8b81ad8caf2c/application","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-12T01:29:10.739Z","expires_at":"2026-09-29T13:32:13.399496Z","created_at":"2026-08-25T18:26:42.681499Z","updated_at":"2026-08-30T13:32:13.539893Z","company_name":"LangChain","company_slug":"langchain","company_logo_url":"https://www.google.com/s2/favicons?domain=langchain.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/258e08a5-3f97-4f5e-bfaf-1bcd3e0ebd55"},{"id":"8901ff8d-733a-4cd4-8657-ca287e37cc7a","company_id":"a0000000-0000-0000-0000-000000000009","title":"Senior Full-Stack Engineer, North Tools \u0026 Retrieval","slug":"senior-full-stack-engineer-north-tools-retrieval-fb589d77","description":"Who are we?\n\nCohere is the leading security-first enterprise AI company.  We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.\n\nWe’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that.\n\nWe obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft.\n\nWe are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us!\n\n\n\nAbout North:\n\nNorth https://cohere.com/north is Cohere's cutting-edge AI workspace platform, designed to revolutionize the way enterprises utilize AI. It offers a secure and customizable environment, allowing companies to deploy AI while maintaining control over sensitive data. North integrates seamlessly with existing workflows, providing a trusted platform that connects AI agents with workplace tools and applications.\n\n\nWhy this role:\n\nThe Data Abstractions \u0026 Tools team is the group inside North responsible for how users bring their data into the platform and put it to work. We own the systems that let users upload, organize, and share their files and knowledge bases, the framework for extending agents with skills and third-party integrations, and the data-sync pipeline that connects enterprise data sources to North's agents.\n\nWe are a small team owning a wide surface area, and we need a senior generalist who can move freely across backend and frontend as the roadmap demands. You will pick up any layer of a feature, from API design and database schema to the frontend that sits on top of it, and own it end-to-end.\n\n\nIn this role, you will:\n\n - Design, build, ship, and maintain features across the full stack for our roadmap surfaces: large-scale file upload and management, shared knowledge bases, the agent extensibility framework (skills and integrations), and the pipeline that syncs external enterprise data into the platform.\n\n - Own features end-to-end, from technical design through implementation, testing, launch, and iteration.\n\n - Write and ship minimal code that runs in low-resource environments with stringent deployment mechanisms. As security and privacy are paramount, you will sometimes need to re-invent the wheel and won't be able to use the most popular libraries or tooling.\n\n - Collaborate closely with Product, UX, and the modeling team to validate requirements and prototype quickly.\n\n - Raise the engineering bar: patterns, testing discipline, code review, and mentoring across the team.\n\n - Use AI actively in your work, while staying accountable for the quality and reliability of what you ship.\n\n\n\nYou may be a good fit if:\n\n - Have shipped (lots of) fullstack code (Python and React) in production\n\n - You excel in fast-paced environments and can execute while priorities and objectives are a moving target\n\n - You have strong coding abilities and are comfortable working across the stack.\n\n - You have built and deployed extremely performant client-side or server-side RAG/agentic applications to millions of users.\n\n - You think like a product owner, not just an engineer. You ask \"should we build this?\" before \"how do we build this?\"\n\n - You have strong opinions on what makes a great user experience and aren't afraid to push back on designs or specs that don't serve the user.\n\n - You've owned features end-to-end: from understanding the user problem, to shipping, to measuring impact.\n\n\n\nWorking Location:\n\nThis role can be based remotely or from one of our office locations listed on the job description - there is no minimum in-office qualification requirement. We care most about hiring exceptional people regardless of locations, though please check the location listed on the posting for guidance around the core time zone or working hours alignment expected for the role.\n\nCompensation:\n\n - For candidates based in California, New York and Washington States, the compensation range is: $215,000 – $260,000\n\n - For candidates based elsewhere in the US, the compensation range is: 180,000 – $220,000\n\n - For candidates based in Canada, the compensation range is: $260,000 – $310,000\n\n\n\n\nFULL-TIME EMPLOYEES AT COHERE ENJOY THESE PERKS:\n\n - A weekly lunch stipend of $75/£75 or equivalent in your local currency for lunch.\n\n - Full health and dental benefits, including a separate budget for mental health.\n\n - RRSP matching, 401K, Pension Scheme.\n\n - 100% Parental Leave top-up for up to 6 months, for either parent.\n\n - Annual enrichment benefits:\n   \n   Arts \u0026 culture, fitness/wellness, quality time, and a workspace improvement credit.\n   \n   Education \u0026 learning ","salary_min":260000,"salary_max":310000,"location":"Toronto, Canada","workplace":"remote","remote_scope":"restricted","job_type":"full-time","experience_level":"senior","tags":["agents","api-design","payments","fullstack","evaluation"],"apply_url":"https://jobs.ashbyhq.com/cohere/6ebac60b-0758-4bb8-8299-44328d4926cb/application","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-11T10:29:11.067Z","expires_at":"2026-09-29T13:31:59.690737Z","created_at":"2026-08-25T18:26:39.48938Z","updated_at":"2026-08-30T13:31:59.830744Z","company_name":"Cohere","company_slug":"cohere","company_logo_url":"https://www.google.com/s2/favicons?domain=cohere.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/8901ff8d-733a-4cd4-8657-ca287e37cc7a"},{"id":"a8bfa422-da52-4371-aecb-c64631f4aaa0","company_id":"d3f1a010-47af-48d2-8b4e-a5953078daac","title":"Senior Product Operations Manager, Evaluation Quality","slug":"senior-product-operations-manager-evaluation-quality-cc54ed93","description":"WHY HARVEY\n\nAt Harvey, we’re transforming how legal and professional services operate. By combining frontier agentic AI, an enterprise-grade platform, and deep domain expertise, we’re reshaping how critical knowledge work gets done for decades to come.\n\nThis is a rare chance to help build a generational company at a true inflection point. We have strong product-market fit and world-class investor support. We’re scaling fast and defining a new category in real time. The work is ambitious, the bar is high, and the opportunity for growth — personal, professional, and financial — is unmatched.\n\nOur team moves fast, takes ownership, and is deeply committed to the mission — operating with intensity, staying close to our customers, and pushing each other for excellence. We live by three values: Decisiveness, Simplicity, and Job's Not Finished. We act quickly on clear judgment over perfect information, we believe simplicity is what scales, and we're never satisfied with where we are. If you want to do the best work of your career alongside people who share that drive, we'd love to build with you.\n\nAt Harvey, the future of professional services is being written today — and we’re just getting started.\n\n\n\n\nROLE OVERVIEW\n\nWe’re looking for a senior operator to own the quality bar behind Harvey’s human evaluations. As we scale globally, the volume of eval work is growing 10x, but volume only matters if the output is trusted. This role makes Human Data’s signal decision-grade: rigorous, calibrated, and reproducible enough that Product, Engineering, and AI Research act on it to ship.\n\n\n\nAs a member of our Evaluation Operations team, you’ll work alongside our Evaluation Operations Manager (who runs throughput and coordination) and partner closely with Applied Legal Researchers, Product, Engineering, and AI Research. You'll set the standard for what \"good\" looks like across eval data and methodology, and own the data analyses to make conclusions that teams will rely on, building the stakeholder trust that lets EPD act on the signal.\n\n\n\n\nWHAT YOU'LL DO\n\n - Own the quality bar for Harvey’s human evaluations: define what “good” looks like for eval methodology and data analysis, and produce decision-grade outputs for EPD\n\n - Author and maintain the evaluation guidelines, instructions, and databases that contract attorneys work from\n\n - Standardize and streamline rubric and evaluation design into repeatable templates and one documented methodology, partnering with Applied Legal Research (ALR), who supplies feature-specific legal depth\n\n - Own contract-attorney quality: onboarding, calibration training, inter-rater reliability, and the feedback loop (including benchmarks and gold references) that keeps judgment consistent across attorneys and over time\n\n - Conducting quantitative and qualitative statistical analyses, diagnosing error states, investigating root causes, and turning raw eval results into a structured, prioritized signal Product and ALR can act on\n\n - Run QA on vendor and contract-attorney deliverables against a defined bar before results inform a launch decision\n\n - Ensure the quality bar holds across jurisdictions and non-English geographies as coverage expands\n\n - Establish one standard, documented way to analyze eval results, and build lightweight operational dashboards to track rater capability and eval-program health\n\n - Support ALR in a review step that certifies an evaluation is sound before it scales to contract attorneys\n\n - Partner with ALR and Analytics to determine where human eval aligns with online signal and where it can provide expanded insights\n\n\n\n\nWHAT YOU HAVE\n\n - 6+ years in product operations, research operations, evaluation/QA operations, or quality program management\n\n - A track record of owning quality inputs (guidelines, instructions, benchmarks, QA procedures) for complex, expert-driven or human-in-the-loop work\n\n - Experience onboarding, training, and calibrating a distributed pool of expert raters, annotators, or reviewers, and running the feedback loop that improves their quality over time\n\n - Enough grounding in measurement concepts (calibration, inter-rater reliability, sampling, rubric design) to independently set up and own the quality of our evaluation loop yourself\n\n - Experience with running quantitative and qualitative data analyses, interpreting and running statistical tests on evaluation data (natively or with AI tool support), and communicating conclusions to various stakeholders\n\n - A record of scaling and streamlining evaluation quality processes under shipping pressure, with a bias toward documentation and reproducibility over one-off analysis\n\n - Ability to work deeply with domain experts (e.g., ALR / lawyers) and translate nuanced judgment into repeatable, documented standards\n\n - Strong cross-functional coordination across Product, Engineering, Research, ALR, and data providers/vendors\n\n - Clear communicator who can build credibility and tru","salary_min":155400,"salary_max":233200,"location":"San Francisco, CA","workplace":"remote","remote_scope":"unknown","job_type":"full-time","experience_level":"senior","tags":["llm","agents","legal","evaluation"],"apply_url":"https://jobs.ashbyhq.com/harvey/f0e4e24a-2a7a-4fc4-8872-9a8a762fecf5/application","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-08-04T16:25:51.563Z","expires_at":"2026-09-29T13:32:50.360186Z","created_at":"2026-08-25T18:26:48.998981Z","updated_at":"2026-08-30T13:32:50.504255Z","company_name":"Harvey AI","company_slug":"harvey-ai","company_logo_url":"https://www.google.com/s2/favicons?domain=harvey.ai\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/a8bfa422-da52-4371-aecb-c64631f4aaa0"},{"id":"862e6737-af2a-43b2-8d28-3f1aca3b1716","company_id":"72014eb6-e84d-48c2-af5c-5424ebec0b3c","title":"Machine Learning Manager, Feed Relevance (Retrieval)","slug":"machine-learning-manager-feed-relevance-retrieval-a91727ea","description":"Reddit is a community of communities. It’s built on shared interests, passion, and trust, and is home to the most open and authentic conversations on the internet. Every day, Reddit users submit, vote, and comment on the topics they care most about. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet’s largest sources of information. For more information, visit www.redditinc.com .\n Reddit is looking for an experienced Engineering Manager to lead our Feed Retrieval team. In this role, you’ll lead a high-impact team of Machine Learning Engineers building the systems that identify, retrieve, and shape the candidate inventory powering Reddit’s personalized feeds. Your team will work at the foundation of Feed Relevance: expanding the set of high-quality content Reddit can recommend, improving personalization and discovery for users across different levels of signal, and building scalable ML systems that directly shape the experiences of over 120M+ daily users. If applying ML / AI in production to improve Reddit Relevance excites you, then you’ve found the right place.\n Responsibilities: \n \n Define Technical Vision \u0026 Strategy: Define the technical vision and long-term roadmap for Feed Retrieval, aligning large-scale recommender-system investments with Reddit’s product, ecosystem, and business objectives.\n Roadmap \u0026 Prioritization: Translate broad Feed Relevance goals into a focused team roadmap, making clear prioritization tradeoffs across model quality, inventory expansion, experimentation velocity, infrastructure cost, and operational reliability.\n Team Leadership \u0026 Development: Coach and support the development of your team, constantly seeking opportunities to grow their skills and impact. \n Technical Execution \u0026 Delivery: Oversee the design, development, and optimization of retrieval systems that source relevant, diverse, fresh, and high-quality candidates for personalized feed experiences.\n Measurement \u0026 Learning: Establish strong measurement, experimentation, and debugging practices so the team can understand retrieval quality, candidate coverage, source incrementality, and downstream impact.\n Platform \u0026 Infrastructure Collaboration: Collaborate with ML platform, infrastructure, ranking, safety, and product teams to build scalable, low-latency retrieval systems that can support the next generation of AI-powered recommendations.\n Operational Excellence: Maintain high standards for system performance, reliability, latency, cost efficiency, and responsible recommendation practices.\n Cross-Functional Partnership: Work with cross-functional partners from across the company to identify key areas of opportunity, set expectations, and communicate your team’s work.\n Recruiting \u0026 Growth: Partner with our incredible recruiting team to attract, interview, and hire diverse and talented machine learning engineers, growing a world-class team.\n \n Qualifications: \n \n Experience Leading ML Teams: 2+ years of experience building and managing high-performing ML or recommender-systems teams.\n Deep ML Expertise: Hands-on experience with large-scale production ML systems, ideally including recommender systems, retrieval models, embedding-based systems, sequence models, transformer-based architectures, or LLM-powered recommendation applications.\n Technical Domain Knowledge: Strong understanding of recommender systems, especially candidate retrieval, embedding/indexing systems, ranking handoffs, feed personalization, exploration, content quality, and measurement strategies. \n Strategic Thinking: Ability to develop and communicate a clear technical strategy across ambiguous problem spaces, balancing user relevance, ecosystem health, system scalability, and business impact.\n Impact-Driven Mindset: Passion for developing scalable, well-designed, and responsible AI solutions that drive business value.\n Exceptional Communication \u0026 Collaboration: Strong interpersonal skills and a collaborative mindset, with the ability to effectively communicate complex technical topics to diverse audiences and build strong relationships with cross-functional partners.\n \n Benefits: \n \n Comprehensive Healthcare Benefits and Income Replacement Programs\n 401k with Employer Match\n Global Benefit programs that fit your lifestyle, from workspace to professional development to caregiving support\n Family Planning Support\n Gender-Affirming Care\n Mental Health \u0026 Coaching Benefits\n Flexible Vacation \u0026 Paid Volunteer Time Off\n Generous Paid Parental Leave \n \n #LI-remote, #LI-JS5\n Pay Transparency: \n This job posting may span more than one career level.\n In addition to base salary, this job is eligible to receive equity in the form of restricted stock units, and depending on the position offered, it may also be eligible to receive a commission. Additionally, Reddit offers a wide range of benefits to U.S.-based employees, including medical, dental, and vision insurance, 401(k) prog","salary_min":253300,"salary_max":354600,"location":"Remote (US)","workplace":"remote","remote_scope":"restricted","job_type":"full-time","experience_level":"junior","tags":["healthcare","llm","fine-tuning","evaluation","machine-learning"],"apply_url":"https://job-boards.greenhouse.io/reddit/jobs/8094985","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-07-29T19:24:35Z","expires_at":"2026-09-29T13:38:57.239784Z","created_at":"2026-07-30T14:09:04.880987Z","updated_at":"2026-08-30T13:38:57.38425Z","company_name":"Reddit","company_slug":"reddit","company_logo_url":"https://www.google.com/s2/favicons?domain=www.reddit.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/862e6737-af2a-43b2-8d28-3f1aca3b1716"},{"id":"729a7b16-2793-4b11-aff6-a0af6753962e","company_id":"28040a6c-6f94-41a4-b15a-f2e4520188ff","title":"AI Evaluation Engineer","slug":"ai-evaluation-engineer-bacea6f4","description":"About Dialpad Dialpad is the AI platform for customer experience, built to resolve customer problems in real time across voice and digital. Our AI agents learn from your best human agents and improve with every interaction, helping organizations understand their customers, deliver better experiences, increase operational efficiencies, and build a lasting competitive advantage. \n Unlike legacy systems built to route and answer, or standalone agentic bot vendors built to deflect, Dialpad was built to resolve. Our AI agents and human agents operate on a single platform with shared context, allowing Agentic AI to resolve issues, advance deals, and eliminate busywork through automation while seamlessly handing conversations to humans when needed, with full context preserved. \n Market-leading brands, including Randstad, Motorola Solutions, Netflix, the San Diego Padres, the Colorado Rockies Baseball Club, and Cal Athletics, trust Dialpad. Dialpad is backed by Andreessen Horowitz, GV, ICONIQ Capital, and T-Mobile. \n Being a Dialer At Dialpad, AI isn’t just a feature; it’s how our teams do their best work every day. We put powerful AI tools in every employee’s hands so they can move faster, think bigger, and achieve more. \n We believe every conversation matters. And we’ve built the platform that turns those conversations into insight and action, for our customers and ourselves. \n We look for people who are intensely curious and hold themselves to a high bar. Our ambition is significant, and achieving it requires a team that operates at the highest level. We seek individuals who embody our core traits: Scrappy, Curious, Optimistic, Persistent, and Empathetic . \n Your role As an AI Evaluation Engineer, you'll be an integral part of our AI Evaluation team, owning evaluation coverage for Dialpad's Agentic AI systems alongside our existing evaluation lead. A key focus will be co-owning LLM-judge metric development and calibration, scenario and benchmark dataset curation, and structured error analysis to support release-readiness decisions for our agentic voice and chat solutions. \n This position reports to the manager of the AI Evaluation team and has the opportunity to be based in our Vancouver office. \n What you’ll do  \n \n You will design and execute validation strategies for agentic, NLP, and speech workflows across staging, beta, and release candidates. \n You will build, run, and improve regression evaluations, A/B comparisons, and red teaming analyses to determine whether product and model changes are ready to move forward. \n You will co-own LLM-judge metric development, calibration, and prompt refinement across evaluation dimensions. \n You will create, configure, and monitor data annotation jobs to keep evaluation and calibration datasets fed on schedule. \n You will develop and maintain QA tooling, notebooks, and pipeline components that make recurring evaluations scalable and reusable across teams. \n You will investigate bugs, triage issues, and decide whether problems should become engineering escalations, test set additions, or follow-up analysis. \n You will collaborate with cross-functional teams, including applied science, engineering, and Product QA. \n \n Skills you’ll bring   \n \n Bachelor's or Master's degree in Computer Science, Software Engineering, Computational Linguistics, or a related field. \n 3+ years of experience in QA, test engineering, model evaluation, or applied ML quality for AI-driven products. \n Experience designing structured test strategies across manual and automated workflows. \n Comfort working with complex AI systems such as speech, NLP, LLM, or agentic products. \n Experience working with evaluation datasets, gold sets, adversarial test sets, or benchmark creation for AI systems. \n Strong analytical skills for investigating failures, comparing outputs, and identifying actionable quality patterns. \n Experience collaborating with cross-functional technical teams and communicating clearly through documentation and reporting. \n For exceptional talent based in Ontario, Canada  the target base salary range for this position is posted below. Our salary ranges are determined by role, level, and location. The range displayed on each job posting reflects the target range for new hire salaries for the position. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training. Your recruiter can share more about the specific salary range for your preferred location during the hiring process. Please note that the compensation details listed in Ontario role postings reflect the base salary only, and do not include bonus, equity, or benefits. \n Ontario Salary Range\n $96,000 — $116,250 CAD \n Why Join Dialpad \n \n Work at the center of the AI transformation in business communications \n Build and ship agentic AI products that are redefining how companies operate \n Join a team where AI amplif","salary_min":96000,"salary_max":116250,"location":"Kitchener, Canada","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"mid","tags":["fine-tuning","alignment","nlp","llm","agents","evaluation"],"apply_url":"https://job-boards.greenhouse.io/dialpad/jobs/8642915002","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-07-23T14:02:57Z","expires_at":"2026-09-29T13:50:48.916615Z","created_at":"2026-07-24T14:20:46.836616Z","updated_at":"2026-08-30T13:50:49.048491Z","company_name":"Dialpad","company_slug":"dialpad","company_logo_url":"https://www.google.com/s2/favicons?domain=dialpad.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/729a7b16-2793-4b11-aff6-a0af6753962e"},{"id":"76dfe6b0-19ed-4122-a27e-0b792fbb7bdd","company_id":"b459414f-fd43-42c4-a6e1-f07225286a75","title":"Evaluations Engineering - Member of Technical Staff","slug":"evaluations-engineering-member-of-technical-staff-0b12929a","description":"ABOUT THE COMPANY\n\nSimile is The Simulation Company. We simulate human behavior to keep people at the center of the decisions that shape the world. With AI, anyone can create a product, a campaign, a policy, or a script — the bottleneck has moved upstream. The hard question is no longer whether you can create something, but what to create, for whom, and how to bring it to life. Those are fundamentally human decisions, and they shouldn't be left to chance or handed off to an algorithm. We're building the infrastructure to understand human behavior at scale and to represent humans in an increasingly agentic world. Our mission is to simulate all eight billion people on earth.\n\n\n\nWe launched five months ago. Since then we've grown revenue 5x, built a new foundation model for human behavior that has run tens of millions of simulations for F100 enterprises, trained a first-of-its-kind confidence model that predicts the accuracy of every simulation, and released the first product that lets organizations verifiably predict the future. The world's leading companies use Simile to make business-critical decisions — from consumer leaders like CVS Health and Wealthfront to professional services organizations like Deloitte and Gallup — strategizing product launches, entering new markets, and forecasting earnings calls.\n\n\n\nWe've raised over $200M at a $2B post-money valuation led by Greenoaks, with Index Ventures, Hanabi, A*, Bain Capital Ventures, and CVS Health Ventures. We've grown from a small home in Palo Alto to a global team of 50+, and we're building a team of the best researchers, engineers, designers, and operators in the world. The future is too important to be left to chance.\n\n\n\n\nABOUT THE ROLE\n\nAs a Member of Technical Staff in Evaluations Engineering, you will build the systems that enable Simile to evaluate whether our simulations of human behavior are accurate, trustworthy, and improving over time.\n\n\n\nYou will work across data and evaluation infrastructure, evaluation execution workflows, backend services, automation, and internal tooling. Your initial focus will include streamlining how evaluations are run across models; strengthening evaluation versioning, data models, and access controls; and automating customer validations, survey operations, and human data workflows.\n\n\n\nEvaluation at Simile presents unusual engineering challenges. Our models predict distributions of human behavior, and the ground truth used to evaluate them can be noisy and heterogeneous. You will partner closely with Evals, Modeling, Product Engineering, and Data Operations to turn complex methods and inputs into systems that are reproducible, scalable, and useful for model development and business decisions.\n\n\n\n\nIN THIS ROLE, YOU WILL:\n\n - Build evaluation execution infrastructure: Develop the services, pipelines, and orchestration needed to run evaluations efficiently across datasets, model versions, populations, and use cases.\n\n - Strengthen evaluation data systems: Design relational schemas, versioning, provenance, permissions, and quality controls that make evaluation results reproducible and trustworthy.\n\n - Automate validation and data collection: Partner with Evals and Data Operations to streamline customer validations, survey deployment, response ingestion, and the integration of new ground truth.\n\n - Build human data workflows: Create labeling and review tools that enable external experts and operators to contribute high-quality judgments to evaluation campaigns.\n\n - Develop evaluation tooling: Build interfaces that help teams manage evals, compare models, investigate results, and identify regressions.\n\n\n\n\nREQUIREMENTS\n\n\nMUST HAVES\n\n - Strong Engineering Fundamentals: Several years of experience building and maintaining production-quality software, with sound judgment in system design, testing, debugging, and maintainability.\n\n - Data and Systems Experience: Experience building backend services, data pipelines, automation workflows, and relational data models.\n\n - End-to-End Execution: Ability to work across data, backend, and interface layers and take ambiguous projects from technical design through deployment and adoption.\n\n - Evaluation Judgment: Strong intuition for what makes evaluation infrastructure reliable, including versioning, provenance, reproducibility, holdout integrity, noisy ground truth, and meaningful model comparisons.\n\n - ML and LLM Fluency: Familiarity with modern model-development and evaluation workflows sufficient to partner effectively with modeling and evaluation researchers.\n\n - Product and User Judgment: Ability to build clear, efficient tools for researchers, engineers, data operators, and other expert users.\n\n - Ownership and Communication: A track record of independently driving important technical work and collaborating effectively across engineering, research, and operations.\n\n\n\n\nNICE TO HAVES\n\nWe do not expect one person to have all of these. We are hiring a team with complementary st","salary_min":200000,"salary_max":400000,"location":"San Francisco, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["llm","generative-ai","data-pipeline","agents","evaluation"],"apply_url":"https://jobs.ashbyhq.com/simile/beaa243c-233f-45f2-9e10-8d54e1eda9d1/application","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-07-17T00:00:46.898Z","expires_at":"2026-09-29T13:41:15.400323Z","created_at":"2026-07-18T14:11:35.756548Z","updated_at":"2026-08-30T13:41:15.535315Z","company_name":"Simile","company_slug":"simile","company_logo_url":"https://www.google.com/s2/favicons?domain=simile.ai\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/76dfe6b0-19ed-4122-a27e-0b792fbb7bdd"},{"id":"9a7b6e3f-9ca5-4d74-8149-e52a00eeffdc","company_id":"3da82454-107f-427f-88e7-01f315ef93fb","title":"Applied Research - Evals \u0026 Data","slug":"applied-research-evals-data-2b9e0702","description":"OWN YOUR INTELLIGENCE\n\n\n\nPrime Intellect is building the open superintelligence stack: the infrastructure frontier AI labs build internally, made available to every ambitious AI team. \n\n\n\nOur platform, Lab, unifies compute, environments, evaluations, secure sandboxes, high-performance training, and deployment into one full-stack system for post-training at frontier scale - from SFT and RL to tool use, agent workflows, and continuously improving production models. We are building open frontier AI: open-source models trained end to end for long-horizon tasks like autonomous research, and the full-stack platform our own research team uses to build them. The next generation of AI companies, enterprises, and research teams do not just need more GPUs. They need the ability to turn their own workflows, tools, data, and feedback loops into superintelligence they own.\n\n\n\nPrime Intellect has raised $150M in total funding from Founders Fund, Radical Ventures, NVIDIA, and exceptional AI, infrastructure, and enterprise operators — including Andrej Karpathy, Dwarkesh Patel, and leaders and founders from Ramp, Perplexity, Harvey, Mercor, Zapier, Datadog, Cognition, OpenAI, Thinking Machines, Together AI, SemiAnalysis, LangChain, Browserbase, Cloudflare, Sierra, Databricks, Airbnb, OpenRouter, Standard Intelligence, Fleet, Core Auto, and more. We are looking for people who want to build at the intersection of frontier research, real infrastructure, and go-to-market for a category that does not fully exist yet.\n\n\nRole Impact\n\nThis is a customer facing role at the intersection of cutting-edge RL/post-training methods, applied data, and agent systems. You’ll have a direct impact on shaping how advanced models are aligned, evaluated, deployed, and used in the real world by:\n\n - Advancing Agent Capabilities: Designing and iterating on next-generation AI agents that tackle real workloads—workflow automation, reasoning-intensive tasks, and decision-making at scale. Working with applied data from real deployments to continuously refine policies, improve reasoning, and enhance reliability and safety.\n\n - Building Robust Infrastructure: Developing the distributed systems, evaluation pipelines, and coordination frameworks that enable these agents to operate reliably, efficiently, and at massive scale. Building data capture, processing, and versioning workflows for feedback, model traces, and reward signals.\n\n - Bridge Between Customers \u0026 Research: Translating customer needs and insights from applied data into clear technical requirements that guide product and research priorities. Collaborating closely with RL and eval teams to ensure real-world signals inform model alignment and reward shaping.\n\n - Prototype in the Field: Rapidly designing and deploying agents, evals, and harnesses alongside customers to validate solutions. Using applied evaluation data to iterate on model performance and discover new capabilities.\n\n\nCustomer-Facing Engineering\n\n - Work side-by-side with customers to deeply understand workflows, data sources, and bottlenecks.\n\n - Prototype agents, data pipelines, and eval harnesses tailored to real use cases, then hand off hardened systems to core teams.\n\n - Translate customer insights and evaluation results into roadmap and research direction.\n\n\nPost-training \u0026 Reinforcement Learning\n\n - Design and implement novel RL and post-training methods (RLHF, RLVR, GRPO, etc.) to align large models with domain-specific tasks.\n\n - Build evaluation harnesses and verifiers to measure reasoning, robustness, and agentic behavior in real-world workflows.\n\n - Integrate applied data collection and analytics into the post-training process to surface regressions, emergent skills, and alignment opportunities.\n\n - Prototype multi-agent and memory-augmented systems to expand capabilities for customer-facing solutions.\n\n\nAgent Development \u0026 Infrastructure\n\n - Rapidly prototype and iterate on AI agents for automation, workflow orchestration, and decision-making.\n\n - Extend and integrate with agent frameworks to support evolving feature requests and performance requirements.\n\n - Architect and maintain distributed training and inference pipelines, ensuring scalability and cost efficiency.\n\n - Develop observability and monitoring (Prometheus, Grafana, tracing) to ensure reliability and performance in production deployments.\n\n\nRequirements\n\n - Strong background in machine learning engineering, with experience in post-training, RL, or large-scale model alignment.\n\n - Experience with applied data workflows and evaluation frameworks for large models or agents (e.g., SWE-Bench, HELM, EvalFlow, internal eval pipelines).\n\n - Deep expertise in distributed training/inference frameworks (e.g., vLLM, sglang, Ray, Accelerate).\n\n - Experience deploying containerized systems at scale (Docker, Kubernetes, Terraform).\n\n - Track record of research contributions (publications, open-source contributions, benchmarks) in ML/RL.\n\n - Passion for advancing the","salary_min":150000,"salary_max":300000,"location":"New York, NY","workplace":"remote","remote_scope":"unknown","job_type":"full-time","experience_level":"senior","tags":["distributed-systems","agents","data-pipeline","llm","reinforcement-learning","evaluation","research"],"apply_url":"https://jobs.ashbyhq.com/PrimeIntellect/bbfe94a6-d1a8-47e9-86af-f117277cdacb/application","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-07-08T18:34:09.743Z","expires_at":"2026-09-29T13:40:21.153213Z","created_at":"2026-04-13T15:01:32.581029Z","updated_at":"2026-08-30T13:40:21.289015Z","company_name":"Prime Intellect","company_slug":"PrimeIntellect","company_logo_url":"https://www.google.com/s2/favicons?domain=primeintellect.ai\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/9a7b6e3f-9ca5-4d74-8149-e52a00eeffdc"},{"id":"4aa05e90-1b2a-4e5a-b2e0-4035bd5fc1fe","company_id":"b459414f-fd43-42c4-a6e1-f07225286a75","title":"Evaluations - Member of Technical Staff","slug":"evaluations-member-of-technical-staff-8a4ec4b6","description":"ABOUT THE COMPANY\n\nSimile is The Simulation Company. We simulate human behavior to keep people at the center of the decisions that shape the world. With AI, anyone can create a product, a campaign, a policy, or a script — the bottleneck has moved upstream. The hard question is no longer whether you can create something, but what to create, for whom, and how to bring it to life. Those are fundamentally human decisions, and they shouldn't be left to chance or handed off to an algorithm. We're building the infrastructure to understand human behavior at scale and to represent humans in an increasingly agentic world. Our mission is to simulate all eight billion people on earth.\n\n\n\nWe launched five months ago. Since then we've grown revenue 5x, built a new foundation model for human behavior that has run tens of millions of simulations for F100 enterprises, trained a first-of-its-kind confidence model that predicts the accuracy of every simulation, and released the first product that lets organizations verifiably predict the future. The world's leading companies use Simile to make business-critical decisions — from consumer leaders like CVS Health and Wealthfront to professional services organizations like Deloitte and Gallup — strategizing product launches, entering new markets, and forecasting earnings calls.\n\n\n\nWe've raised over $200M at a $2B post-money valuation led by Greenoaks, with Index Ventures, Hanabi, A*, Bain Capital Ventures, and CVS Health Ventures. We've grown from a small home in Palo Alto to a global team of 50+, and we're building a team of the best researchers, engineers, designers, and operators in the world. The future is too important to be left to chance.\n\n\n\n\n\n\n\nABOUT THE ROLE\n\nAs a Member of Technical Staff, Model Evaluations at Simile, you will build the measurement systems that determine whether our simulations of human behavior are accurate, trustworthy, and useful enough to guide real-world decisions. You will help shape what Simile measures, the quality bars we defend, and how evaluation evidence guides model, product, and customer decisions.\n\n\n\nEvaluation at Simile brings together model evals, statistics, behavioral science, research methodology, product quality, and human judgment. Our models simulate people, populations, markets, and groups, which means our evals must reason about distributions, noisy human ground truth, uncertainty, qualitative outputs, behavioral data, and customer decision-making. You will work with unusually rich data about human behavior, including surveys, long-form interviews, customer studies, qualitative research, and behavioral signals such as transactions, product interactions, and other real-world traces.\n\n\n\nWe are hiring across several forms of expertise. Some candidates may be deep in LLM evaluation, model training, and research engineering. Others may bring exceptional strength in statistics, behavioral science, survey methodology, human data, product evaluation, or experimentation. Across backgrounds, we are looking for people who can reason clearly, build quickly, use agentic coding tools fluently, and take hands-on ownership of ambiguous evaluation problems.\n\n\n\nThe core question for this role is simple: How do we know when a simulation of human behavior is good enough to trust?\n\n\n\n\nIN THIS ROLE, YOU WILL:\n\n - Build the measurement layer for behavioral simulation: Design evals, metrics, rubrics, datasets, dashboards, and workflows that measure whether Simile’s models are accurately predicting human behavior across customer use cases, populations, question types, and decision contexts.\n\n - Partner with modeling to improve models: Evaluate new model versions, diagnose regressions, identify priority areas for model-improvement cycles, and maintain stable eval suites that represent capabilities customers actually care about.\n\n - Contribute to product and applied evals: Build evals for qualitative responses, retrieval, survey generation, AI-generated research reports, customer-facing outputs, and other product surfaces where model quality directly shapes customer trust. Turn subjective quality concerns into concrete rubrics, labeled data, automated graders, release criteria, and model-improvement signals.\n\n - Make ground truth and uncertainty legible: Develop rigorous ways to compare simulated responses against human data, customer studies, Simile-collected ground truth, and behavioral datasets. Help the company reason about sampling error, uncertainty, calibration, margin of error, representativeness, and what “ground truth” means when human behavior is inherently noisy.\n\n - Automate evaluation workflows: Use modern agentic coding tools to rapidly build internal tools, inspect model outputs, create labeling workflows, validate evals, and turn fuzzy evaluation questions into working systems. We value people who can compress long, ambiguous projects into fast, useful prototypes without losing sight of rigor or reliability.\n\n - Help define th","salary_min":200000,"salary_max":400000,"location":"San Francisco, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["llm","agents","search","generative-ai","research","evaluation"],"apply_url":"https://jobs.ashbyhq.com/simile/33d75074-c23b-4a1f-bfdb-129bcc5be662/application","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-06-30T22:51:51.838Z","expires_at":"2026-09-29T13:41:15.307514Z","created_at":"2026-07-01T14:10:44.230358Z","updated_at":"2026-08-30T13:41:15.440338Z","company_name":"Simile","company_slug":"simile","company_logo_url":"https://www.google.com/s2/favicons?domain=simile.ai\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/4aa05e90-1b2a-4e5a-b2e0-4035bd5fc1fe"},{"id":"255321d4-5ae5-4873-9105-ae267dcc102c","company_id":"b6db41bc-ba14-4906-b2f7-a3ce9289a346","title":"User Researcher, AI Evaluations","slug":"user-researcher-ai-evaluations-f04d2a5f","description":"WHO WE ARE\n\nNotion is the collaborative AI workspace where teams and agents think together https://www.youtube.com/watch?v=vkpYpWfEK5s. We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion.\n\n\n\nNotinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work.\n\n\n\n\nABOUT THE ROLE:\n\nWe’re seeking an experienced UX Researcher to define and scale how we evaluate Notion’s AI-powered experiences—focusing on what “good” looks like not only for model output quality, but for the end-to-end product experience where people discover, set goals, delegate work, review results, and build trust over time with AI.\n\n\n\nThis role sits at the intersection of research craft and evaluation operations: you’ll run studies that uncover user mental models, expectations, and failure/recovery behaviors, then translate those insights into reusable rubrics, workflows, and measurement approaches that product, design, engineering, and data science can apply consistently.\n\n\n\nThis role can be based in either San Francisco or New York City. We work from our offices on Mondays, Tuesdays and Thursdays (our Anchor Days) because we do our best thinking and building together in person. We’re looking for someone who’s excited to work alongside the team during those days. \n\n\n\n\nWHAT YOU'LL ACHIEVE:\n\n - Define what “good” looks like (frameworks \u0026 rubrics): Establish clear, reusable evaluation criteria that reflect real user expectations—helpfulness, trust, tone, control, and transparency. You’ll translate qualitative insight into scoring guidance that can be applied consistently across teams and over time.\n\n - Run recurring evals (longitudinal \u0026 feature-specific): Run recurring longitudinal and feature-specific surveys and studies to measure experience quality over time against defined rubrics. Lead qualitative studies, side-by-side comparisons, and human-in-the-loop evaluation efforts to deepen understanding of where experiences break down and how they can improve. You’ll help teams spot regressions, benchmark improvements, and understand when expectations shift.\n\n - Anchor evaluation in real workflows (context \u003e isolated feedback): Ensure evals reflect jobs-to-be-done, user intent, and the full interaction journey (goal setting, delegation, review, iteration), not just decontextualized thumbs up/down. You’ll help teams understand who is evaluating, what they’re trying to do, and why outputs succeed or fail.\n\n - Identify failure modes \u0026 recovery behavior (guardrails): Uncover breakdowns, regressions, and edge cases across the system—from model behavior to UI and integrations—and study how people notice issues, correct them, and continue their work. You’ll turn these insights into actionable guidance for guardrails, fixes, and prioritization.\n\n - Operationalize evaluation with partners (process \u0026 tooling): Collaborate closely with Product, Design, Engineering, and Data Science to align on target use cases and build scalable evaluation loops (human-in-the-loop review, longitudinal studies, and calibration of automated/LLM-judge approaches against human judgment).\n\n\n\n\nSKILLS YOU'LL NEED TO BRING: \n\n - Ability to operationalize insight into measurement: You’re comfortable turning “soft” user expectations (trust, tone, usefulness, clarity) into concrete rubrics, scoring guidelines, and observable metrics.\n\n - AI fluency and systems thinking: You’re curious and hands-on with AI products, and can reason about how model behavior, uncertainty, and system constraints shape user experience. You also have experience evaluating AI-enabled products (LLMs, agents, generative UI/workflow automation) and working with Data Science/ML partners on measurement strategy and evaluation tooling.\n\n - Clear communication and impact orientation: You can align diverse partners around shared definitions of quality and create artifacts that enable teams to act consistently. You tailor storytelling to different audiences, connect research to business outcomes, and drive follow-through so insights translate into product change.\n\n - Strong UX research craft (quant + qual): You can choose the right methods for the question— interviews, benchmarking, surveys, experiments—and synthesize into actionable guidance. You also can prioritize ruthlessly, work through ambiguity, and balance scrappy iteration with deep","salary_min":196000,"salary_max":230000,"location":"San Francisco, CA","workplace":"remote","remote_scope":"unknown","job_type":"full-time","experience_level":"senior","tags":["agents","llm","research","evaluation"],"apply_url":"https://jobs.ashbyhq.com/notion/0e9114bd-4603-4bdf-a86f-a7a4f390fae8/application","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-06-18T23:15:14.006Z","expires_at":"2026-09-29T13:33:43.47371Z","created_at":"2026-06-28T14:03:08.935234Z","updated_at":"2026-08-30T13:33:43.617687Z","company_name":"Notion","company_slug":"notion","company_logo_url":"https://www.google.com/s2/favicons?domain=notion.so\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/255321d4-5ae5-4873-9105-ae267dcc102c"},{"id":"edd81527-9709-4b01-8b95-c78ef07b4bc1","company_id":"6ce2d21e-b00f-4343-9bd0-5ac62ff81431","title":"Senior Machine Learning Engineer, Simulation Evaluation","slug":"senior-machine-learning-engineer-simulation-evaluation-2b2339ea","description":"Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver—The World's Most Experienced Driver™—to improve access to mobility while saving thousands of lives now lost to traffic crashes. The Waymo Driver powers Waymo’s fully autonomous ride-hail service and can also be applied to a range of vehicle platforms and product use cases. The Waymo Driver has provided over ten million rider-only trips, enabled by its experience autonomously driving over 100 million miles on public roads and tens of billions in simulation across 15+ U.S. states.\n The Challenge \n Waymo’s simulator is one of the most complex virtual environments ever built. It blends deterministic logic, physical dynamics, and state-of-the-art Generative AI to create a training ground for the Waymo Driver. The Simulator Evaluation team faces the ultimate data challenge: How do you mathematically prove that a virtual world is \"real\" ?\n We are seeking visionary machine learning engineers and researchers to architect the scalable deep learning systems, novel data workflows, and eval tools that power our research roadmap. In this role, you will pioneer the machine learning and generative vision paradigms required to define and measure the realism of our multimodal world models. Your work will define the state of the art for autonomous simulation, directly steering our research trajectory and the capabilities of the Waymo Driver.\n You will: \n \n Lead the design, development and deployment of cutting-edge evaluation approaches to assess realism of state-of-the-art multimodel world models and generative systems for simulation use cases at Waymo.\n Architect and implement robust and scalable machine learning pipelines for tuning, evaluating, and deploying large-scale discriminator models for the purposes of simulator realism evaluation. \n Evaluate open-source and production-ready video generation techniques that measure realism (e.g. temporal stability, multi-modal consistency, geometric discrepancy, condition following, etc.)\n Apply vision language models to evaluate semantic understanding and controllability across our world simulation products.\n Collaborate with research teams across Waymo and Alphabet to integrate advancements in 4D world modeling and generative AI into production systems.\n Mentor engineers on the team and provide technical guidance on architecture and execution.\n \n You have: \n \n Bachelor's, Master's, or PhD in computer science, machine learning, robotics, or a related field.\n Five or more years of experience in machine learning engineering or applied deep learning, supported by a portfolio of shipped products or peer-reviewed publications.\n Proficient programming skills in Python and hands-on experience with modern machine learning frameworks such as Jax, Flax, or PyTorch.\n Experience designing and implementing evaluation frameworks for complex systems or machine learning models.\n \n We prefer: \n \n Track record of training large-scale generative models (diffusion models, flow matching, vision language models, etc.)\n A PhD and demonstrated success delivering machine learning products focused on 3D generative models, world models, or video generation.\n Experience simulating sensor data, including camera, lidar, and radar, or modeling semantic scenes.\n Experience developing autonomous systems, robotics software, or autonomous vehicle simulations.\n Experience training and optimizing large-scale models on GPU or TPU clusters for efficient production serving.\n Professional experience writing C++ for high-performance production systems.\n The expected base salary range for this full-time position across US locations is listed below. Actual starting pay will be based on job-related factors, including exact work location, experience, relevant training and education, and skill level. Your recruiter can share more about the specific salary range for the role location or, if the role can be performed remote, the specific salary range for your preferred location, during the hiring process.  \n Waymo employees are also eligible to participate in Waymo’s discretionary annual bonus program, equity incentive plan, and generous Company benefits program, subject to eligibility requirements.  \n Salary Range\n $213,000 — $263,000 USD","salary_min":213000,"salary_max":263000,"location":"Mountain View, CA","workplace":"onsite","remote_scope":"not_remote","job_type":"full-time","experience_level":"senior","tags":["autonomous-vehicles","generative-ai","deep-learning","robotics","pytorch","diffusion-models","machine-learning","evaluation"],"apply_url":"https://careers.withwaymo.com/jobs?gh_jid=8001797","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-06-13T17:49:43Z","expires_at":"2026-09-29T13:35:06.515878Z","created_at":"2026-06-28T14:04:26.926315Z","updated_at":"2026-08-30T13:35:06.659952Z","company_name":"Waymo","company_slug":"waymo","company_logo_url":"https://www.google.com/s2/favicons?domain=waymo.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/edd81527-9709-4b01-8b95-c78ef07b4bc1"},{"id":"4b255973-955d-4d22-ba64-dc9d0c08e001","company_id":"053355fc-0162-4bb9-b414-cbf7679ee9c8","title":"Director, Research - Evaluation \u0026 Training","slug":"director-research-evaluation-training-067f35f5","description":"About Snorkel \n At Snorkel, we believe meaningful AI doesn’t start with the model, it starts with the data.\n We’re on a mission to help enterprises transform expert knowledge into specialized AI at scale. The AI landscape has gone through incredible changes since 2015, when Snorkel started as a research project in the Stanford AI Lab, to the generative AI breakthroughs of today. But one thing has remained constant: the data you use to build AI is the key to achieving differentiation, high performance, and production-ready systems. We work with some of the world’s largest organizations to empower scientists, engineers, financial experts, product creators, journalists, and more to build custom AI with their data faster than ever before. Excited to help us redefine how AI is built? Apply to be the newest Snorkeler!\n ABOUT THE ROLE  \n We're looking for a manager to lead a team of researchers to focus on data evaluation, error analysis and data valuation methods to predict model performance. This team is responsible for showcasing the value and quality of Snorkel’s data for model training and evaluation, understanding where today's frontier models fall short, and turning that understanding into a point of view on what benchmarks and datasets these models will benefit from. \n You and your team will be responsible for Snorkel’s data design flywheel by analyzing model failures, finding capability and skill gaps in current models, suggesting the next benchmarks to invest in and then proving the value of this data for our customers. \n  \n MAIN RESPONSIBILITIES  \n \n Own a multi-quarter roadmap centered on novel evaluation, error analysis, and data valuation techniques\n Synthesize and share trends from model-failure analysis and benchmarking into recommendations on the datasets the community should focus on and the ones Snorkel should invest in — making this team a primary input to the company's data strategy.\n Focus on data valuation techniques that quantify how Snorkel data meaningfully improves model performance\n Lead and grow a team of researchers, setting a high bar for quality, rigor and speed of execution\n Act as the primary bridge between the team's findings and Product, GTM, and our customers\n \n  \n PREFERRED QUALIFICATIONS  \n \n 7+ years in applied AI, ML, or research roles, with 4+ years managing technical teams.\n A leader who has repeatedly turned research and analysis into business outcomes, and who instinctively connects technical findings to market and customer needs.\n Strong business and market judgment in the AI/ML space — you understand the competitive and frontier-lab landscape and can prioritize accordingly.\n Technically conversant and credible: enough depth in LLM evaluation, benchmarking, and model behavior analysis to set direction, judge experimental quality, and pressure-test results — without needing to be the deepest technical expert in the room.\n A nose for trends: able to look across many evaluation results and failure cases and extract the signal that should drive what gets built next.\n Excellent communication and storytelling skills, with the ability to make technical results legible and persuasive to non-research audiences.\n Familiarity with data valuation or data attribution research is a strong plus.\n Bonus: experience working with frontier labs, public benchmarks, or commercial AI data/eval products.\n Actual compensation will be determined based on factors including skills, qualifications, experience, and geographic location.\n Salary range(s) for this role\n $275,000 — $425,000 USD \n Be Your Best at Snorkel \n Joining Snorkel AI means becoming part of a company that has market proven solutions, robust funding, and is scaling rapidly—offering a unique combination of stability and the excitement of high growth. As a member of our team, you’ll have meaningful opportunities to shape priorities and initiatives, influence key strategic decisions, and directly impact our ongoing success. Whether you’re looking to deepen your technical expertise, explore leadership opportunities, or learn new skills across multiple functions, you’re fully supported in building your career in an environment designed for growth, learning, and shared success.\n Snorkel AI is proud to be an Equal Employment Opportunity employer and is committed to building a team that represents a variety of backgrounds, perspectives, and skills. Snorkel AI embraces diversity and provides equal employment opportunities to all employees and applicants for employment. Snorkel AI prohibits discrimination and harassment of any type on the basis of race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state, or local law. All employment is decided on the basis of qualifications, performance, merit, and business need. \n We will ensure that individuals with dis","salary_min":275000,"salary_max":425000,"location":"New York, NY","workplace":"remote","remote_scope":"unknown","job_type":"full-time","experience_level":"lead","tags":["generative-ai","llm","research","evaluation"],"apply_url":"https://job-boards.greenhouse.io/snorkelai/jobs/6020877004","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-06-11T22:43:31Z","expires_at":"2026-09-29T13:34:00.099084Z","created_at":"2026-06-28T14:03:22.023885Z","updated_at":"2026-08-30T13:34:00.244061Z","company_name":"Snorkel AI","company_slug":"snorkel-ai","company_logo_url":"https://www.google.com/s2/favicons?domain=snorkel.ai\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/4b255973-955d-4d22-ba64-dc9d0c08e001"},{"id":"d8e4dfdd-920a-4d58-9605-44f850da2a35","company_id":"e8c9f3a5-9310-43f5-9341-321fe6d93a92","title":"Director of Platform Management for Simulation, Evaluation \u0026 Validation","slug":"director-of-platform-management-for-simulation-evaluation-validation-74af71dd","description":"About us    \n Founded in 2017, Wayve is the leading developer of Embodied AI technology.  Our advanced AI software and foundation models enable vehicles to perceive, understand, and navigate any complex environment, enhancing the usability and safety of automated driving systems.\n Our vision is to create autonomy that propels the world forward.  Our intelligent, mapless, and hardware-agnostic AI products are designed for automakers, accelerating the transition from assisted to automated driving.  In our fast-paced environment big problems ignite us—we embrace uncertainty, leaning into complex challenges to unlock groundbreaking solutions. We aim high and stay humble in our pursuit of excellence, constantly learning and evolving as we pave the way for a smarter, safer future.\n At Wayve, your contributions matter.  We value diversity, embrace new perspectives, and foster an inclusive work environment; we back each other to deliver impact.  \n Make Wayve the experience that defines your career!  \n The Role  \n As Director of Platform Management for Simulation, Evaluation \u0026 Validation , you will lead the overall product vision and director for the platforms that Wayve uses to develop, evaluate, and validate the AI Driver. This includes open and closed loop simulation platforms, the evaluation test platform, synthetic data production, critical event pipelines, validation, and the workflows that connect them.\n You will directly lead and indirectly influence sub team platform leads across Simulation, Evaluation, Data, and Validation to convert a complex set of cross-cutting initiatives into a single, opinionated product strategy. The job is part developer-tools PM (your users are engineers and scientists who will tell you immediately when something is slow, wrong, or in the way), part platform PM (your roadmap has to compose across simulators, datasets, evaluation, and triage), and part safety-systems PM (the outputs gate releases that put cars on real roads).\n You will help shape and set vision for how our tools and services should function and work in an AV2.0 context - a Wayve-pioneered approach to both the driving stack and offline development.\n Key Responsibilities \n \n Set product vision and strategy for the platforms. Define a multi-year, multi-quarter product strategy that takes Wayve from a collection of capable tools to a unified, opinionated platform for evaluating and validating the AI Driver. Anchor the strategy in measurable customer outcomes — developer velocity, signal quality, cost per evaluation, time-to-insight, validation credibility.\n Own the product roadmap end-to-end. Translate company-level priorities into a coherent roadmap across simulation, evaluation, and validation initiatives. Drive trade-offs between different internal and external customer groups. Ensure resourcing decisions are clearly communicated to other director level stakeholders.\n Lead and grow the team. Manage and mentor the product managers embedded across. Hire to fill gaps, level up the craft, and build a highly capable lean team.\n Build world class developer experiences. Run regular user research with Autonomy, Science, Validation, Release, and Product teams. Maintain a clear picture of what each customer segment needs from the platform, where the friction is, and which gaps are quietly bleeding velocity. \n Define and defend quality bars for an internal platform. Set platform-wide standards for reliability, latency, cost per unit (per simulation, per evaluation, per enriched hour), self-service, observability, and API stability. Treat platform regressions the way a consumer team treats a churn spike.\n Establish regular operational cadences. Lead monthly business reviews and close the loop on requests and progress reporting with the leadership team. Lead quarterly planning, KPI reviews, and the trade-off conversations.\n Represent the platform externally. Help Wayve's leadership tell a credible story about how we develop and validate the AI Driver — to OEM partners, regulators, and at relevant industry forums.\n \n About You  \n In order to set you up for success at Wayve, we’re looking for the following skills and experience.  \n Essential  \n \n 8+ years of product management experience , with at least 3–5 years leading PM teams (PMs and/or senior PMs reporting to you).\n Track record building internal platforms, developer tools, or ML/AI evaluation infrastructure — products whose users are engineers and scientists, where the bar is set by power users with strong opinions and the ability to route around you if the tool isn't good enough.\n Data driven . You always seek to answer questions with data and lean into defining the right customer facing KPIs indicative of success\n Systems thinking across hardware, AI, and product. You understand how decisions in one part of the stack (data, simulator fidelity, metric design) propagate into outcomes (release confidence, on-road performance, regulator trust) els","salary_min":332000,"salary_max":415000,"location":"Sunnyvale, CA","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["generative-ai","autonomous-vehicles","data-pipeline","robotics","evaluation"],"apply_url":"https://wayve.firststage.co/jobs?gh_jid=8585026002","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-06-11T19:43:11Z","expires_at":"2026-09-29T13:43:22.717035Z","created_at":"2026-06-28T14:12:39.528441Z","updated_at":"2026-08-30T13:43:22.84833Z","company_name":"Wayve","company_slug":"wayve","company_logo_url":"https://www.google.com/s2/favicons?domain=wayve.ai\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/d8e4dfdd-920a-4d58-9605-44f850da2a35"},{"id":"d0c707f8-46ab-4546-a1fe-71cf6c2db09a","company_id":"6ce2d21e-b00f-4343-9bd0-5ac62ff81431","title":"Senior Software Engineer, ML/Eval Data Platforms \u0026 Infrastructure","slug":"senior-software-engineer-mleval-data-platforms-infrastructure-02234f34","description":"Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver—The World's Most Experienced Driver™—to improve access to mobility while saving thousands of lives now lost to traffic crashes. The Waymo Driver powers Waymo’s fully autonomous ride-hail service and can also be applied to a range of vehicle platforms and product use cases. The Waymo Driver has provided over ten million rider-only trips, enabled by its experience autonomously driving over 100 million miles on public roads and tens of billions in simulation across 15+ U.S. states.\n The Planner Evaluation team works on one of the key challenges in autonomous driving: measuring and improving the quality of the software that drives the car. We are looking for experienced data-minded software engineers and data scientists to help us improve how we characterize and evaluate changes to the Onboard software stack (Planner, Perception, etc). If you are passionate about autonomous vehicles and how to use rich, complex data to drive decision making, this is the role for you!\n This role follows a hybrid work schedule and reports to an Engineering Manager. \n  \n You will:\n \n Develop and productionize data pipelines to generate high-quality ML and evaluation datasets, streamlining curation, sampling, and slicing for training and testing the Waymo Driver software.\n Design and implement robust tools and infrastructure for data mining, exploration, and analysis to extract insights from large-scale datasets and drive data-driven decisions.\n Architect, build, and maintain large-scale data platforms to process Waymo driving logs and simulation data, ensuring dataset generation is fresh, accurate, and complete.\n Own and execute complex projects, successfully translating ambiguous requirements into high-impact deliverables.\n Collaborate cross-functionally with Data Science and Quantitative Analytics teams to identify their data needs and engineer infrastructure solutions.\n \n You have:\n \n Education: Master’s degree or PhD in Computer Science, Engineering, or a related technical field.\n Experience: 3+ years of professional software engineering experience.\n Distributed Systems: Hands-on experience with systems that ingest, store, transform, and output data at scale.\n Programming Languages: Proficiency in C++ or Python within a production environment.\n Communication: Strong technical communication and collaboration skills.\n \n  \n We prefer:\n \n Experience with A/B experiment infrastructure.\n Exposure to ad-hoc data analysis utilizing SQL.\n Prior experience working within the autonomous vehicle (AV) industry.\n \n  \n #Hybrid\n The expected base salary range for this full-time position across US locations is listed below. Actual starting pay will be based on job-related factors, including exact work location, experience, relevant training and education, and skill level. Your recruiter can share more about the specific salary range for the role location or, if the role can be performed remote, the specific salary range for your preferred location, during the hiring process.  \n Waymo employees are also eligible to participate in Waymo’s discretionary annual bonus program, equity incentive plan, and generous Company benefits program, subject to eligibility requirements.  \n Salary Range\n $213,000 — $263,000 USD","salary_min":213000,"salary_max":263000,"location":"Mountain View, CA","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"senior","tags":["autonomous-vehicles","fine-tuning","data-pipeline","distributed-systems","evaluation","infrastructure"],"apply_url":"https://careers.withwaymo.com/jobs?gh_jid=7991303","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-06-09T17:21:44Z","expires_at":"2026-09-29T13:35:07.871524Z","created_at":"2026-06-28T14:04:28.292816Z","updated_at":"2026-08-30T13:35:08.007144Z","company_name":"Waymo","company_slug":"waymo","company_logo_url":"https://www.google.com/s2/favicons?domain=waymo.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/d0c707f8-46ab-4546-a1fe-71cf6c2db09a"},{"id":"2087e015-ee4e-49c3-b24d-576611c371ec","company_id":"a0000000-0000-0000-0000-000000000001","title":"Staff+ Software Engineer, Safeguards Evals ","slug":"software-engineer-safeguards-evals-c7ee0bf5","description":"About Anthropic \n Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.\n About the role \n How do we know whether a model is safe - and how do we know whether the systems we built to catch misuse actually catch it?\n Anthropic answers both questions with evaluations. We measure model behavior across misuse, prompt injection, and user well-being to inform training and deployment decisions. We also use AI to investigate potential misuse of Claude, analyzing real-world traffic to surface bad actors and emerging threats that drive enforcement actions. Neither is worth much unless the evaluations behind them are representative, robust, and trustworthy.\n This role builds the methods and infrastructure that make them so. Sitting at the intersection of applied ML research and engineering, you'll design experiments to improve how we evaluate both model behavior and the agentic systems that govern it, build datasets that represent real abuse rather than clean approximations of it, and ship those methods into the pipelines that gate model training, agent changes, and launch decisions.\n Responsibilities \n Evaluation research and methodology. Design and run experiments to improve evaluation quality — developing methods to generate representative test data, simulate realistic user behavior, and validate grading accuracy. Research how different factors (multi-turn conversations, tools, long context, user diversity) impact model safety behavior. Analyze evaluation coverage to identify measurement gaps, and evolve evals so they remain unsaturated and high-signal as model and agent capabilities advance.\n Agentic investigation evals. Build and own the evaluation harness for an agentic investigation system — defining metrics, test cases, and grading approaches for a complex, long-horizon agent. Measure agent performance end-to-end (detection precision and recall, investigation quality, robustness) and drive hill-climbing on the hardest harm areas. Construct RL environments to improve Claude's safety investigation capabilities.\n Datasets grounded in real harm. Construct high-quality eval datasets representing real-world misuse across harm areas such as cyber attacks, bio weapons, and influence operations, drawing from both real traffic patterns and synthetic generation. Collaborate with Policy and Enforcement to translate observed harm patterns into measurable evaluations.\n Productionization and tooling. Ship successful research into evaluation, regression, and release pipelines that run during model training, on every agent change, prompt update, and underlying model upgrade, and beyond launch. Build tooling that enables policy experts to author, run, and iterate on evaluations without engineering support. Surface findings to research and training teams to drive upstream model improvements.\n Minimum qualifications \n \n 8+ years of industry software engineering or ML engineering experience\n Experience building and maintaining data pipelines\n Experience working with LLMs and a working understanding of their capabilities and failure modes — especially agentic systems with tool use and multi-step reasoning\n Strong data analysis skills — you can draw reliable insights from large datasets\n Ability to move fluidly between research prototyping and production-quality code\n Ability to translate ambiguous problems into concrete, testable experiments\n Care deeply about AI safety and want your work to have real impact\n \n Preferred qualifications \n \n Expertise in building or contributing to LLM or agent evaluation frameworks, benchmarks, or automated grading systems\n Extensive experience in trust and safety, content moderation, or abuse detection systems\n Experience in red teaming, adversarial testing, or jailbreak research on AI systems\n Experience with synthetic data generation or data augmentation\n Experience with distributed systems or large-scale data processing\n Experience with prompt engineering or building LLM-powered applications\n The annual compensation range for this role is listed below. \n For sales roles, the range provided is the role’s On Target Earnings (\"OTE\") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role.\n Annual Salary:\n $320,000 — $485,000 USD \n Logistics \n Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience\n Required field of study:  A field relevant to the role as demonstrated through coursework, training, or professional experience\n Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position\n Location-based hybrid policy: Currently, we expe","salary_min":320000,"salary_max":485000,"location":"San Francisco, CA","workplace":"hybrid","remote_scope":"not_remote","job_type":"full-time","experience_level":"lead","tags":["alignment","llm","distributed-systems","agents","data-pipeline","rust","evaluation"],"apply_url":"https://job-boards.greenhouse.io/anthropic/jobs/5251671008","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-06-09T15:02:53Z","expires_at":"2026-09-29T13:30:42.494473Z","created_at":"2026-06-28T14:00:32.865428Z","updated_at":"2026-08-30T13:30:42.637675Z","company_name":"Anthropic","company_slug":"anthropic","company_logo_url":"https://www.google.com/s2/favicons?domain=anthropic.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/2087e015-ee4e-49c3-b24d-576611c371ec"},{"id":"d3c94339-dd13-4c3f-ad05-a18b15c19e75","company_id":"f36ec848-cb19-4b95-a680-6733e58086c0","title":"Machine Learning Engineer II - Autonomous Driving Performance Evaluation","slug":"machine-learning-engineer-ii-autonomous-driving-performance-evaluation-8e6d5a1d","description":"May Mobility is transforming cities through autonomous technology to create a safer, greener, more accessible world. Based in Ann Arbor, Michigan, May develops and deploys autonomous vehicles (AVs) powered by our innovative Multi-Policy Decision Making (MPDM) technology that literally reimagines the way AVs think. Our vehicles do more than just drive themselves - they provide value to communities, bridge public transit gaps and move people where they need to go safely, easily and with a lot more fun. We’re building the world’s best autonomy system to reimagine transit by minimizing congestion, expanding access and encouraging better land use in order to foster more green, vibrant and livable spaces. Since our founding in 2017, we’ve given more than 500,000 autonomous rides to real people around the globe. And we’re just getting started. We’re hiring people who share our passion for building the future, today, solving real-world problems and seeing the impact of their work. Join us. \n Job Summary \n May Mobility is entering an exciting phase of growth as we expand our first-of-its-kind autonomous shuttle and mobility services across the nation. Launched in 2017 with a strong team of experienced roboticists and software engineers with decades of experience fielding robotic systems in the wild, May Mobility is looking to expand its team of robotics engineers with a background in robotics or autonomous vehicles.\n We are seeking ML-Oriented Software Engineers with experience in robotics applications. As part of our Autonomous Driving ML team, you will use ML Engineering concepts to measure, analyze and systematically improve the performance of May's Autonomous Driving stack through data, metrics, evaluation and test/hillclimbing suites.\n Essential Responsibilities \n \n Design, implement and own ML metrics and evaluation pipelines spanning offline model evaluation, simulation and on-road performance. \n Build and maintain test, regression and hillclimbing suites that gate model and stack releases, including automated triage of regressions to root cause. \n Drive model improvement through loss analysis, error mining, and data balancing/curation strategies for training and evaluation sets.\n \n Skills and Abilities \n Success in this role typically requires the following competencies: \n \n Designing quantitative metrics and statistical analyses that translate model behavior into actionable, decision-grade signals (significance, slicing, long-tail analysis).\n Building evaluation and analytics frameworks in production, including dataset slicing, result aggregation and dashboarding at scale. \n Applying data-centric ML methods such as hard-example mining, resampling/reweighting and curriculum or balance adjustments to lift model performance.\n \n Qualifications and Experience \n Candidates most successful in this role typically hold the following qualifications or comparable knowledge or experience: \n Required \n \n Bachelor's or Master's degree in Robotics, Computer Science, Statistics, or a related field with strong mathematical and engineering foundations. \n A minimum of 2 years building evaluation, metrics, or data analysis systems for ML in production. \n Proficiency in Python (NumPy/Pandas or equivalent dataframe tooling) with experience in Linux environments. \n Familiarity with basic concepts in Machine Learning (losses, train/eval splits, common failure modes) and basic Perception and Planning concepts in Autonomous Driving.\n \n Desirable \n \n Proficiency in Go or C++. \n Familiarity with experiment tracking and evaluation tooling such as MLflow, Weights \u0026 Biases, or in-house equivalents. \n Familiarity with statistical methods for A/B comparison, regression detection and noisy-metric analysis. \n Familiarity with data mining and curation at scale (embedding-based retrieval, active learning, auto-labeling). \n Familiarity with visualization and dashboarding tools (Plotly, Grafana, Streamlit or similar).\n \n Physical Requirements \n \n Standard office working conditions which includes but is not limited to:\n \n Prolonged sitting\n Prolonged standing\n Prolonged computer use\n \n \n Travel required? -  Low 5-10%\n \n \n \n \n \n \n \n Benefits and Perks \n \n Comprehensive healthcare suite including medical, dental, vision, life, and disability plans. Domestic partners who have been residing together at least one year are also eligible to participate. \n Health Savings and Flexible Spending Healthcare and Dependent Care Accounts available.\n Rich retirement benefits, including an immediately vested employer safe harbor match.\n Generous paid parental leave as well as a phased return to work. \n Flexible vacation policy in addition to paid company holidays.\n Total Wellness Program providing numerous resources for overall wellbeing   \n \n Don’t meet every single requirement? Studies have shown that women and/or people of color are less likely to apply to a job unless they meet every qualification. At May Mobility, we’re committe","salary_min":172000,"salary_max":210000,"location":"Remote (US)","workplace":"remote","remote_scope":"restricted","job_type":"full-time","experience_level":"junior","tags":["healthcare","autonomous-vehicles","robotics","evaluation","machine-learning"],"apply_url":"https://job-boards.greenhouse.io/maymobility/jobs/8187068002","is_featured":false,"is_sticky":false,"status":"active","published_at":"2026-06-08T20:01:04Z","expires_at":"2026-09-29T13:47:54.127214Z","created_at":"2026-06-28T14:16:46.812396Z","updated_at":"2026-08-30T13:47:54.257672Z","company_name":"May Mobility","company_slug":"may-mobility","company_logo_url":"https://www.google.com/s2/favicons?domain=maymobility.com\u0026sz=128","quality_score":90,"url":"https://aidevboard.com/job/d3c94339-dd13-4c3f-ad05-a18b15c19e75"}],"page":1,"per_page":20,"total":99,"total_is_exact":true,"total_pages":5}
