Research Engineer - Evals

Firecrawl · San Francisco, CA · $210k - $224k
full-time senior Posted 2 months ago

Before you apply

Build my evidence-backed draft — free Apply on company site →

Paste your relevant resume section or 2–4 true bullets. See supported requirements and honest gaps. No account and no application sent.

Get weekly job alerts like this →

About this role

RESEARCH ENGINEER - EVALS You'll build the evaluation systems that tell us whether Firecrawl actually works. That sounds simple. It isn't. Our core promise, convert any URL into clean, structured, LLM-ready data reliably, is hard to measure rigorously across millions of different websites, formats, and edge cases. As the systems we're measuring get more complex, the question "did that work?" gets harder, not easier. This isn't an eval role where you inherit a framework and run benchmarks. You'll design the metrics, build the pipelines, generate the datasets, and own the feedback loop from output quality back to model and product decisions. If you care about what "good" actually means and have the engineering depth to measure it, this is the role. Salary Range: $250,000–$290,000 USD/year (SF) / $210,000–$224,000 CAD/year (Toronto) Equity Range: Competitive equity. Details shared during the process. Location: San Francisco, CA (SF HQ) or Toronto, ON (Toronto Hub). On-site, five days a week. Job Type: Full-Time Experience: 4+ years in ML, research engineering, or data-heavy backend, with real evaluation work Work Authorization: Must be authorized to work in the United States or Canada. We're not able to sponsor US visas right now. For Canada, we'll consider sponsorship on a case-by-case basis through our Toronto Hub. ABOUT FIRECRAWL Firecrawl is the easiest way to turn the web into data AI agents can use. One API call converts any URL into clean, LLM-ready markdown or structured data. It's the boring-hard problem everyone building with LLMs eventually hits, solved. We hit 8 figures in ARR in year one and more than doubled it in year two. We have 180k+ GitHub stars, putting us in the top 50 repositories of all time, and developers, agents, and category-defining AI companies build on us every day. Growth like this is rare, and we're just getting started. We're a small team punching far above our weight, working out of SF HQ and our new Toronto Hub. Everyone here owns a real piece of the product and company, end to end, and runs it themselves. No hiding behind process or headcount. This is a place for people who want to work at the frontier: an AI company building the infrastructure other AI companies run on, not one bolting AI onto an existing product. We move fast, go deep, and are building the tools superintelligence will rely on to gather data from the web. WHAT YOU'LL DO - Design the metrics that define what "good output" actually means across millions of sites, formats, and edge cases - Build the pipelines and harnesses that measure quality rigorously and at scale - Generate and curate the datasets that make evaluation trustworthy - Own the feedback loop from output quality back to model and product decisions - Turn "did that work?" into an answer the whole team can act on WHAT WE'RE LOOKING FOR - You have the engineering depth to build real evaluation systems, not just run existing ones - You care deeply about what "good" means and how to measure it rigorously - You're comfortable owning ambiguous problems where the metric itself has to be invented - You move fast and close the loop. You'd rather ship, measure, and iterate than perfect on paper WHAT WE'RE NOT LOOKING FOR - Someone who only wants to run benchmarks someone else designed - A pure researcher who won't build the systems, or a pure engineer who won't think about methodology - Someone who needs a fully-specced ticket to start A NOTE ON PACE We operate at an absurd level of urgency because the window for what we're building won't stay open forever. If that excites you, keep reading. If it doesn't, no hard feelings, but this role probably isn't for you. BENEFITS & PERKS AVAILABLE TO ALL EMPLOYEES - Salary that makes sense: $250,000–$290,000 USD/year (SF) / $210,000–$224,000 CAD/year (Toronto), based on impact, not tenure - Own a piece: Gain competitive equity in what you're helping build - Generous PTO: 15 days mandatory, anything after 24 days, just ask (holidays excluded). Take the time you need to recharge - Parental leave: 12 weeks fully paid, for all parents - Wellness stipend: $100 USD/month for the gym, therapy, massages, or whatever keeps you human - Learning & Development: Expense up to $1,000 USD/year toward anything that helps you grow professionally - Team offsites: A change of scenery, minus the trust falls - Sabbatical: 3 paid months off after 4 years, do something fun and new AVAILABLE TO US-BASED FULL-TIME EMPLOYEES - Full coverage, no red tape: Medical, dental, and vision (100% for employees, 50% for partner and kids). No weird loopholes, just care that works - Life & Disability insurance: Employer-paid basic life and AD&D, short-term disability, and long-term disability. Coverage for life's curveballs - Virtual care and a health guide: Teladoc for the couch doctor visit, plus Rightway to answer coverage questions and fight billin

Similar Jobs

Related searches:

On-site Jobs Senior Jobs On-site Senior Jobs Senior Fintech & Payments AISenior Machine LearningSenior NLP & Language AISenior AI ResearchSenior AI Agents & RAGSenior Data Engineering AI Jobs in San Francisco Fintech & Payments AI in San FranciscoMachine Learning in San FranciscoNLP & Language AI in San FranciscoAI Research in San FranciscoAI Agents & RAG in San FranciscoData Engineering in San Francisco data-pipelineagentspaymentsllmsearchresearchevaluation

Get jobs like this delivered weekly

Free AI jobs newsletter. No spam.