Research Infrastructure - Member of Technical Staff
full-time
lead
Posted 1 day ago
Apply Now
Stand out: build a proof-of-work pitch →
Free GitHub-based preview. Direct apply stays one click away.
Get weekly job alerts like this →Hiring for this role?
AI Market Demand Pack · $29 one-time
Compare this role's skills with the full AI hiring market. Get ranked demand, salary bands, leading companies, public source URLs, and a decision brief.
About this role
ABOUT THE COMPANY
Simile is The Simulation Company. We simulate human behavior to keep people at the center of the decisions that shape the world. With AI, anyone can create a product, a campaign, a policy, or a script — the bottleneck has moved upstream. The hard question is no longer whether you can create something, but what to create, for whom, and how to bring it to life. Those are fundamentally human decisions, and they shouldn't be left to chance or handed off to an algorithm. We're building the infrastructure to understand human behavior at scale and to represent humans in an increasingly agentic world. Our mission is to simulate all eight billion people on earth.
We launched five months ago. Since then we've grown revenue 5x, built a new foundation model for human behavior that has run tens of millions of simulations for F100 enterprises, trained a first-of-its-kind confidence model that predicts the accuracy of every simulation, and released the first product that lets organizations verifiably predict the future. The world's leading companies use Simile to make business-critical decisions — from consumer leaders like CVS Health and Wealthfront to professional services organizations like Deloitte and Gallup — strategizing product launches, entering new markets, and forecasting earnings calls.
We've raised over $200M at a $2B post-money valuation led by Greenoaks, with Index Ventures, Hanabi, A*, Bain Capital Ventures, and CVS Health Ventures. We've grown from a small home in Palo Alto to a global team of 50+, and we're building a team of the best researchers, engineers, designers, and operators in the world. The future is too important to be left to chance.
ABOUT THE TEAM
Research Infrastructure builds the systems that every step of the model lifecycle runs on: data ingestion and schema design, distributed training, evaluation, serving, and monitoring. We are the reason a researcher's hypothesis can become a production simulation in days rather than quarters.
Two things make this problem unusual. First, our research-to-product pipeline is unusually tight - the experimental methods we validate on Monday are integrated into systems customers use to make high-stakes decisions. Second, simulating a society means running inference over populations of agents, not single requests. A single customer study can mean millions of model calls with interdependent state. Cost per simulation and latency per agent are not back-office metrics for us; they determine what research is even possible to run.
ABOUT THE ROLE
As a Member of Technical Staff in Research Infrastructure, you will build the platform our researchers train, evaluate, and deploy on - and own it through the last mile, where a trained checkpoint becomes a production service serving millions of interdependent agent calls at a cost per simulation we can afford.
This is a role for someone who is energized by both halves of that. You will spend some weeks designing the data schemas and training pipelines a research team depends on, others profiling a serving path to find where the FLOPs and GPU memory are going, and others still bringing up cluster nodes or deleting the third redundant copy of a code path. The common thread is leverage: every improvement you make compounds across every researcher and every simulation we run.
We are looking for engineers who find it gratifying to see their work pushed to its absolute limits, and who own problems end-to-end - including the last mile of deployment that most people would rather hand off.
IN THIS ROLE, YOU WILL
- Build the ML platform our researchers live in. Design and operate the services, libraries, and tooling that cover the full lifecycle - data exploration, feature generation, experiment tracking, training orchestration, evaluation, and deployment. Success is defined by your ability to increase experiment velocity, streamlining the researcher’s path from ideation to a fully validated, production-ready model.
- Make training and data pipelines fast. Own throughput end to end: model FLOPs utilization across our training configs, tokenization cost when the data mix changes, and ingestion paths that take hours today where they should take minutes. Profile where the time and GPU memory actually go, then fix it, including the observability that makes the next bottleneck obvious before it bites.
- Make serving fast and cheap enough to run a society. Own the inference path our simulations run on: batching and scheduling, KV cache reuse across agents sharing context, quantization, and the request patterns unique to population-scale runs where one study is millions of interdependent calls. Cost per simulation and latency per agent decide what research we can afford to run at all, so treat them as research constraints, not ops metrics.
- Scaling simulation Data. Lead the redesign of our data architecture to handle the complexity and sheer volume of our simulation mode
Similar Jobs
Related searches:
On-site Jobs
Lead Jobs
On-site Lead Jobs
Lead AI Agents & RAGLead AI InfrastructureLead Data EngineeringLead Generative AILead Machine LearningLead Backend & Systems
AI Jobs in San Francisco
AI Agents & RAG in San FranciscoAI Infrastructure in San FranciscoData Engineering in San FranciscoGenerative AI in San FranciscoMachine Learning in San FranciscoBackend & Systems in San Francisco
mlopspytorchgpugenerative-aidistributed-systemsagentssearchdata-pipeline
Get jobs like this delivered weekly
Free AI jobs newsletter. No spam.