Senior Platform AI Engineer
full-time
senior
Posted 3 months ago
Before you apply
Build my evidence-backed draft — free Apply on company site →Paste your relevant resume section or 2–4 true bullets. See supported requirements and honest gaps. No account and no application sent.
About this role
Drata is building the trust layer between great companies - automating compliance, managing risk, and helping organizations prove trust continuously as they scale. We're Dratanauts: a global crew of 600+ professionals united by a culture that rewards integrity, ownership, and raising the bar, no matter where in the world we're working from.
Why Join the Drata Team?
At Drata, you're not maintaining legacy compliance software - you're building the agentic AI platform defining what trust looks like for the next generation of companies. Here's what makes the work itself worth showing up for:
- Problems without a playbook: You'll work at the edge of AI and security, building agentic governance, continuous compliance, and real-time trust verification to solve problems that don't have an established answer yet. You're writing it as you go.
- Real ownership, not just process: Our values center on owning outcomes and raising the bar, not checking boxes. You're expected to have opinions and back them.
- A seat at the table: Your perspective is unique and valued. Open debate and diverse viewpoints are built into how decisions actually get made here, at every level.
- Growth at rocketship speed: Drata is scaling fast, which means scope grows fast too. High performers get more ownership, visibility, and experience.
- A crew, not just coworkers: Dratanauts consistently describe a "come as you are" culture with sharp, curious people—the kind of team that makes hard problems genuinely fun to solve. See what they say here https://drata.com/about/careers/life and follow us on LinkedIn https://www.linkedin.com/company/drata/posts/?feedView=all for company news, employee stories, and career updates.
Job Summary:
Drata's AI Platform team builds the production infrastructure that powers AI features across our compliance platform — from MCP servers that make Drata's data available to AI agents, to LLM workflow orchestration that automates SOC 2, TPRM, and policy analysis. You'll own the systems that sit between our AI models and our customers: tool definitions that agents actually understand, deployment pipelines that handle model upgrades without breaking output quality, and orchestration layers that manage multi-step agent workflows with persistent state.
This is not a traditional infrastructure role. You'll debug prompt templates alongside Terraform modules. You'll design API schemas optimized for LLM token budgets, not just HTTP throughput. When a model upgrade changes behavior across 15 workflows, you'll assess quality impact — not just confirm the containers are healthy.
You'll work closely with our agent developers, product engineers, and an embedded SRE partner, sitting at the intersection of AI development and production reliability.
Our north star is simple: minimize the time it takes to launch a new agent in production. You're someone who asks "are we solving the right problem?" before writing the first line of code, who builds systems that make five other engineers faster, not just yourself, and who's equally proud of what they chose not to build.
What you'll do:
MCP SERVER DEVELOPMENT & AI-OPTIMIZED API DESIGN
- Design and build MCP (Model Context Protocol) servers that expose Drata's platform to AI agents. This means making architectural decisions about tool granularity, naming conventions for agent disambiguation, response compression for LLM context windows, and workspace isolation for multi-tenant access. You'll own the protocol layer that determines whether agents can reliably find and use the right tools — writing semantic parameter descriptions, contextual hints, and tool schemas that optimize for model comprehension, not just developer ergonomics.
AGENT ORCHESTRATION & WORKFLOW INFRASTRUCTURE
- Build and operate the infrastructure for deploying multi-step agent workflows — state management across complex reasoning chains, tool routing and execution runtimes, and long-running agentic processes that persist over time. Own the orchestration layer that coordinates agent planning, tool calls, and human-in-the-loop patterns. Design systems that handle agent failure modes gracefully: retries on ambiguous tool outputs, fallback strategies when models produce unexpected results, and observability into multi-step execution traces.
LLM OPERATIONS & MODEL LIFECYCLE MANAGEMENT
- Own the operational side of our LLM workflows: model upgrades across production pipelines (assessing behavior changes, not just version bumps), prompt versioning and A/B testing, AI workflow deployment with custom container compatibility, and output quality monitoring.
- Manage token capacity planning — understanding model costs, context limits, batching strategies, and rate governance across workflows. When an AI workflow fails, you'll investigate whether it's a prompt template issue, a model behavior change, or an infrastructure problem. Making that distinction requires understanding both systems.
PR
Similar Jobs
Related searches:
Remote Jobs
Senior Jobs
Remote Senior Jobs
Senior AI Agents & RAGSenior Backend & SystemsSenior NLP & Language AISenior Data EngineeringSenior Data ScienceSenior Machine LearningSenior AI Infrastructure
AI Jobs in San Francisco
AI Agents & RAG in San FranciscoBackend & Systems in San FranciscoNLP & Language AI in San FranciscoData Engineering in San FranciscoData Science in San FranciscoMachine Learning in San FranciscoAI Infrastructure in San Francisco
ragagentsmlopssearchdistributed-systemscloudembeddingsapi-design
Get jobs like this delivered weekly
Free AI jobs newsletter. No spam.