Member of Data Staff

Simile · New York, NY · $200k - $300k
full-time lead Posted 1 week ago

Before you apply

Build my evidence-backed draft — free Apply on company site →

Paste your relevant resume section or 2–4 true bullets. See supported requirements and honest gaps. No account and no application sent.

Get weekly job alerts like this →

About this role

ABOUT THE COMPANY Simile is The Simulation Company. We simulate human behavior to keep people at the center of the decisions that shape the world. With AI, anyone can create a product, a campaign, a policy, or a script — the bottleneck has moved upstream. The hard question is no longer whether you can create something, but what to create, for whom, and how to bring it to life. Those are fundamentally human decisions, and they shouldn't be left to chance or handed off to an algorithm. We're building the infrastructure to understand human behavior at scale and to represent humans in an increasingly agentic world. Our mission is to simulate all eight billion people on earth. We launched five months ago. Since then we've grown revenue 5x, built a new foundation model for human behavior that has run tens of millions of simulations for F100 enterprises, trained a first-of-its-kind confidence model that predicts the accuracy of every simulation, and released the first product that lets organizations verifiably predict the future. The world's leading companies use Simile to make business-critical decisions — from consumer leaders like CVS Health and Wealthfront to professional services organizations like Deloitte and Gallup — strategizing product launches, entering new markets, and forecasting earnings calls. We've raised over $200M at a $2B post-money valuation led by Greenoaks, with Index Ventures, Hanabi, A*, Bain Capital Ventures, and CVS Health Ventures. We've grown from a small home in Palo Alto to a global team of 50+, and we're building a team of the best researchers, engineers, designers, and operators in the world. The future is too important to be left to chance. ABOUT THE TEAM Every agent in our simulation is grounded in data from a real person. That makes the supply chain that drives data acquisition and first-party collection the raw material of our product. This is what drives the difference between a model that predicts human behavior and one that approximates it. Data sits upstream of research, engineering, and every customer deployment. We decide which populations we can credibly simulate, which datasets are worth buying, and how faithfully our agents reflect the people they are modeled on. We work in a small, high-ownership team with direct access to the researchers and customers who consume what we build. ABOUT THE ROLE As a Member of Data Staff, you will own the full picture of how data enters and flows through Simile - both the third-party datasets we license and the first-party data we collect. On the sourcing side, you will map the frontier of the data landscape and secure the datasets that make our simulations predictive across new domains and geographies. On the collection side, you will run the supply chain that turns data from real people into grounded agents. This includes designing data collection instruments, interacting with vendors and partners, and the quality and representativeness standards that determine whether a simulation can be trusted. Your core responsibilities will include: - Expanding our coverage of the world: Deciding which populations Simile should be able to simulate next, then going and getting the data that makes it possible. Much of what you want will not be for sale, which means finding who holds it and showing them our vision for the future. - Running Simile’s data machine: Expanding and running the operations behind our own human data collection - running the supply chain behind Simile’s data engine, which includes panel and field vendor management, incentive structures, throughput, and cost per completed participant. - Finding the richest datasets to improve our simulation of the world: Structuring agreements around how we actually use data - training, fine-tuning, and derivative agent behavior that persists long after a contract term ends. Most data agreements are not written with foundation models in mind, and getting these terms right is the difference between an asset we own and one we license. - Building our always-on feedback loop: Turning what research and forward deployed teams need into a concrete sourcing and supply chain roadmap - and, just as importantly, tracking which data measurably improved the model so the next round of spend is better informed than the last. - Defending data fidelity: Owning the question of whether our agents actually resemble the people they are modeled on. You will set the bar for sample composition and response quality, catch fraud and low-effort participants before they reach a model, and hold the line when a dataset is convenient but not credible. - Trust and compliance: Working with legal so that consent, privacy, and usage rights hold up to the scrutiny of enterprise and government partners. Our access to sensitive populations depends on getting this right the first time. REQUIREMENTS MUST HAVES - You are excited about enhancing Simile’s data supply chain:

Similar Jobs

Related searches:

On-site Jobs Lead Jobs On-site Lead Jobs Lead Generative AILead Machine LearningLead AI Agents & RAG AI Jobs in New York Generative AI in New YorkMachine Learning in New YorkAI Agents & RAG in New York generative-aiagentsfine-tuning

Get jobs like this delivered weekly

Free AI jobs newsletter. No spam.