Staff AI Engineer - Agent Architecture & Behavior

Artisan · San Francisco, CA · $250k - $325k
full-time lead Posted 14 hours ago

Before you apply

Build my evidence-backed draft — free Apply on company site →

Paste your relevant resume section or 2–4 true bullets. See supported requirements and honest gaps. No account and no application sent.

Get weekly job alerts like this →

About this role

Artisan · Full-time · In person in San Francisco Base salary: $250,000–$325,000 USD annually. Equity: 0.15%–0.30%. US visa sponsorship available. BUILD SOMETHING NEW AT THE FRONTIER OF APPLIED AI At Artisan, we're working on a new, ambitious project that will push the boundaries of what agentic AI can do. We're keeping the product details private ahead of launch, but we can tell you this: the technical problems are substantial, the scope for invention is real, and this hire will shape the core technology. We're looking for a hands-on technical lead to design and build the underlying AI architecture. You'll work across agent behavior, complex multi-agent systems, tool use, context, and evaluation, taking promising ideas through to dependable production software. This is an individual-contributor role with broad technical ownership. You'll make consequential architecture decisions, write the hardest parts of the system, and work closely with our existing engineers and leadership. You should enjoy both exploring an uncertain problem and doing the detailed engineering required to make a solution work. WHAT YOU'LL OWN - Agent architecture and behavior. Design and implement agent execution loops, planning strategies, tool interfaces, and verification. Turn ambiguous technical requirements into clear system boundaries and working code. - Multi-agent systems. Build delegation, coordination, context sharing, and result synthesis. Handle concurrent work, conflicting updates, cancellation, and stale results. Establish when a multi-agent approach improves on a simpler baseline. - Reliable execution. Make complex, stateful workflows resilient to interruptions and partial failures. Build checkpoints, recovery strategies, and appropriate human intervention into the architecture. - Context, memory, and reusable methods. Improve retrieval, context construction, persistent state, and skill representation. Investigate how systems can use feedback and prior experience to perform better without introducing regressions. - Evaluation and experimentation. Build realistic evaluations, analyze task trajectories, and turn observed failures into measurable improvements. Compare approaches using quality, reliability, latency, and cost. - Model and tooling decisions. Evaluate models and emerging techniques, prototype promising approaches, and make informed build-versus-buy decisions. Choose tools because they solve the problem, and be willing to replace them when the evidence changes. - Technical leadership. Set engineering standards, review important design decisions, and help the team implement a coherent AI system. Stay close to the product and accountable for what ships. You'll partner with product and infrastructure engineers on production services, integrations, secure execution, and observability. You'll own the AI architecture and its effectiveness, with implementation shared across the team. WHAT WE'RE LOOKING FOR - You have personally built and shipped a substantial agentic system. Production use or rigorous, reproducible open-source work matters more than the name of a framework or employer. - You have deep practical experience with LLM tool use, planning, context engineering, and evaluations. You have implemented multi-agent coordination or substantial parallel agent/tool execution and can explain its failure modes. - You have hands-on experience with browser or computer automation in an agentic system, including observing state, verifying effects, and recovering when an interface or execution path fails. - You are an excellent software engineer in Python, TypeScript, or a comparable language. You are comfortable with asynchronous services, state machines, persistence, concurrency, retries, and cancellation. - You know which decisions belong to a model and which guarantees must be enforced in code. You can reason carefully about permissions, untrusted inputs, uncertain external outcomes, and human approvals. - You can design meaningful experiments, debug real system behavior, and explain what improved, why it improved, and where the evidence is still weak. - You can take technical ownership of an unclear problem, work effectively with other engineers, and ship with urgency and care. USEFUL ADDITIONAL EXPERIENCE Depth in agent memory and retrieval, skill acquisition, reinforcement learning or post-training, trajectory datasets, sandboxed execution, distributed systems, inference optimization, or multimodal and voice models would be valuable. We expect strong foundations and particular depth in a few areas, rather than prior specialization in every one. There is no required degree, publication record, previous employer, or agent framework. We're hiring for demonstrated engineering ability, judgment, and ownership. HOW WE WORK This role is based in our San Francisco office. Expect a small team, short feedback cycles, direct communication, and high standards. We value people wh

Similar Jobs

Related searches:

On-site Jobs Lead Jobs On-site Lead Jobs Lead AI InfrastructureLead Machine LearningLead NLP & Language AILead AI Agents & RAGLead Robotics & AutonomyLead Backend & Systems AI Jobs in San Francisco AI Infrastructure in San FranciscoMachine Learning in San FranciscoNLP & Language AI in San FranciscoAI Agents & RAG in San FranciscoRobotics & Autonomy in San FranciscoBackend & Systems in San Francisco llmreinforcement-learningagentsdistributed-systems

Get jobs like this delivered weekly

Free AI jobs newsletter. No spam.