Research Engineer

Turing · Brazil
full-time senior Posted 5 months ago

Before you apply

Build my evidence-backed draft — free Apply on company site →

Paste your relevant resume section or 2–4 true bullets. See supported requirements and honest gaps. No account and no application sent.

Get weekly job alerts like this →

About this role

About Turing Turing’s mission is to accelerate superintelligence to drive real economic progress. Headquartered in San Francisco, Turing works with frontier AI labs to generate high-quality datasets, reinforcement learning environments, and frontier research benchmarks that improve model capabilities in software engineering, enterprise knowledge work, and advanced STEM reasoning. In software engineering, Turing is the largest and longest-running data provider in the category. Turing also works with Fortune 500 enterprises across financial services, life sciences, healthcare, retail, automotive, and CPG to build and deploy end-to-end agentic AI systems inside mission-critical workflows. By operating on both sides, Turing closes the loop between frontier research and enterprise deployment, turning real-world deployment signals into better data, evaluations, and more capable models. Learn more at www.turing.com .    * This is a remote role and can be performed anywhere in Brazil.* The Role We are looking for a Research Engineer to help deliver frontier-quality datasets, RL environments, and evaluations that improve state-of-the-art models for leading AI labs and enterprise clients. This is a hands-on, research-facing technical leadership role. You will work directly with customer researchers/engineers to translate their model and post-training goals into concrete data and environment specifications, and drive the production of data that meets extremely high standards for correctness, realism, diversity, difficulty, and measurable model lift. This role is designed for candidates with roughly 4 to 5 years of experience building and improving deep learning systems, especially where strong results depend on data quality, data curation, denoising, synthetic data generation, and rigorous evaluation. You’ll operate in one or more of the following capability areas: Coding and software engineering agents (repositories, unit tests, debugging, tool use, code reviews, long-horizon workflows) RL environments and verifier-based training (tasks, rewards/verifiers, trajectories, evaluation harnesses) Multimodal data and reasoning (text + images + documents + tables/charts; optional audio/video) STEM reasoning (math, physics, chemistry, bio, engineering – solution verification and error analysis) Modern embodied AI / VLM-driven agents (vision-language(-action) models, embodied task suites, tool/sensor/action abstractions, long-horizon interaction data) What You’ll Do 1) Own data and environment quality from an AI researcher perspective Translate ambiguous research goals into clear data requirements: target skills, failure modes, difficulty calibration, coverage, and success metrics. Define what “good” looks like by creating detailed rubrics, counterexamples, and boundary cases (what to include vs. exclude). Perform deep, detail-oriented audits of produced data: spot subtle errors, reward hacking opportunities, leakage, ambiguity, inconsistent assumptions, and distribution shifts. Drive iterative improvements using evidence: error taxonomies, slice-based quality metrics, and model-behavior-informed refinements. 2) Design and build datasets and RL environments for your capability area(s) Contribute to or lead the design of: Task suites (single-step and long-horizon workflows) Ground-truth signals (verifiers, unit tests, structured checks, reward functions, automatic validators) Environment interfaces (APIs, tool schemas, state abstractions, database schemas, simulator-like dynamics) Depending on your mapped capability area(s), you may focus on: Coding / SWE agents: data reflecting real development work (codebase navigation, bug localization, patching, tests, code reviews, CI-like constraints, refactors, security fixes). Multimodality: tasks that test true multimodal reasoning (chart reading, document QA, UI understanding, diagram-based STEM reasoning, OCR-aware tasks). STEM: tasks with verifiable solutions (symbolic checks, reference solvers, numerical validation, step consistency, unit sanity). Modern embodied AI / VLM-driven agents: interaction data and environments for vision-language(-action) models (long-horizon tasks, instruction following grounded in visual context, robust action selection, safety/constraint adherence, adversarial state coverage). 3) Build robust validation, denoising, and synthetic data systems Implement automated validation and filtering to achieve frontier-grade signal-to-noise: Deduplication, decontamination, leakage checks Consistency checks (format, schema, invariants) Difficulty and diversity controls (coverage, novelty, long-tail) Develop synthetic data generation and augmentation pipelines where appropriate: Programmatic task generators Controlled perturbations to create hard negatives Scenario templating with diversity constraints Simulator-/tool-driven rollouts for trajectory data Create documentation and

Similar Jobs

Related searches:

On-site Jobs Senior Jobs On-site Senior Jobs Senior Generative AISenior Data EngineeringSenior Machine LearningSenior AI ResearchSenior AI Agents & RAGSenior Robotics & AutonomySenior Healthcare AI agentshealthcarefine-tuningsearchreinforcement-learningdeep-learningresearch

Get jobs like this delivered weekly

Free AI jobs newsletter. No spam.