Machine Learning Engineer
full-time
senior
Posted 4 days ago
Before you apply
Build my evidence-backed draft — free Apply on company site →Paste your relevant resume section or 2–4 true bullets. See supported requirements and honest gaps. No account and no application sent.
About this role
MACHINE LEARNING ENGINEER
You'll build the ML behind Firecrawl — the models and the systems that serve them. That starts with search: training and shipping the ranking and relevance models for one of our fastest-growing products, then extending that work across extraction quality and LLM-driven features. You'll also own how we measure: A/B testing launches and building the experimentation frameworks the whole team ships against. If you ship models into production — whether your title says ML engineer or data scientist — this is for you.
Salary Range: $210,000–$240,000/year
Equity Range: Competitive equity — details shared during the process.
Location: San Francisco, CA (Hybrid, on-site required)
Job Type: Full-Time
Experience: 3+ years building ML or data-heavy systems in production
Visa: Must be legally authorized to work in the United States. We're not able to sponsor visas right now, though that may change down the line.
ABOUT FIRECRAWL
Firecrawl is the easiest way to turn the web into data AI agents can use. One API call converts any URL into clean, LLM-ready markdown or structured data - the boring-hard problem everyone building with LLMs eventually hits, solved.
We hit 8 figures in ARR in year one and more than doubled it in year two. We have 170k+ GitHub stars, and developers, agents, and category-defining AI companies build on us every day. Growth like this is rare, and we're just getting started.
We're a small team punching far above our weight. Everyone here owns a real piece of the product and company, end to end, and runs it themselves - no hiding behind process or headcount.
This is a place for people who want to work at the frontier: an AI company building the infrastructure other AI companies run on, not one bolting AI onto an existing product. We move fast, go deep, and are building the tools superintelligence will rely on to gather data from the web.
WHAT YOU'LL DO
- Improve ranking and relevance for Firecrawl Search — from feature engineering to model training to production
- Build and tune models for learning-to-rank, query understanding, and LLM-driven retrieval
- Extend ML across Firecrawl's products — extraction quality, content classification, and evaluation of LLM-driven features
- Mine query logs and behavioral data at scale to find where our products win and where they fail
- Build the data pipelines that turn web-scale crawl and query data into training data and features
- Work hands-on with platform, search engineers and cloud DevOps to get models running fast and cheap in production
- Design and formulate our testing strategy — the A/B testing frameworks and offline evaluation the team ships against
- Partner on product launches across Firecrawl: define success metrics, run the experiments, and make the ship/no-ship call on evidence
- Report on how releases perform post-launch and turn the findings into the next iteration
WHAT WE'RE LOOKING FOR
- You've shipped ML models into production systems and owned them after launch — deploying, monitoring, and retraining them, not handing them off
- You have real ranking or relevance-modeling experience — learning-to-rank, recommendations, or search quality
- You're comfortable in large, data-heavy systems: query logs, pipelines, and datasets that don't fit in memory
- You write production-quality code (Python at minimum) and can work inside a real backend codebase
- You're rigorous about measurement — you've designed and analyzed A/B tests and know when a lift is real
- You can communicate results clearly to the team — what shipped, what moved, and what to do next
NICE TO HAVE
- MLOps experience — MLflow, experiment tracking, model registries, or feature stores; Kubernetes is a plus
- Experience building or standardizing an experimentation framework at a previous company
- Experience with embedding models, vector retrieval, or LLM-based relevance evaluation
- Experience evaluating LLM outputs at scale — quality scoring, structured-extraction accuracy, or agent behavior
- Spark or similar large-scale data processing experience
WHAT WE'RE NOT LOOKING FOR
- A pure statistician or analyst who needs an engineering team to productionize their work
- Someone who wants to specialize narrowly and hand off everything else
- Someone who optimizes for process over shipping
A NOTE ON PACE
We operate at an absurd level of urgency because the window for what we're building won't stay open forever. If that excites you, keep reading. If it doesn't, no hard feelings — but this role probably isn't for you.
BENEFITS & PERKS
AVAILABLE TO ALL EMPLOYEES
- Salary that makes sense — $210,000–$240,000/year, based on impact, not tenure
- Own a piece — Gain competitive equity in what you're helping build
- Generous PTO — 15 days mandatory, anything after 24 days, just ask (holidays excluded); take the time you need to rec
Similar Jobs
Related searches:
Hybrid Jobs
Senior Jobs
Hybrid Senior Jobs
Senior NLP & Language AISenior Data EngineeringSenior Data ScienceSenior Machine LearningSenior AI InfrastructureSenior AI Agents & RAG
AI Jobs in San Francisco
NLP & Language AI in San FranciscoData Engineering in San FranciscoData Science in San FranciscoMachine Learning in San FranciscoAI Infrastructure in San FranciscoAI Agents & RAG in San Francisco
searchagentsmlopsdata-pipelinellmembeddingsmachine-learning
Get jobs like this delivered weekly
Free AI jobs newsletter. No spam.