Research Engineer Benchmarks
San Francisco, CA - USA
Job Summary
Every week someone asks which data provider is actually best: for company records for finding the right person for fresh job listings. Right now the answers come from the vendors themselves. Youll build the independent version. Youll run rigorous automated public benchmarks of the providers in Alexandria and beyond and publish them weekly. The same results will feed straight back into Alexandria so it learns which provider to call for which job.
This isnt our internal evals role. Youre measuring the market in public where every number will be challenged by the vendors it ranks. Youll own it end to end: the datasets the ground truth the scoring the harness the weekly release and the loop into Alexandrias provider selection. No one hands you a methodology. You write it defend it and ship it every week.
Salary Range: $250000$290000 USD/year (SF) / $210000$224000 CAD/year (Toronto)
Equity Range: Competitive equity. Details shared during the process.
Location: San Francisco CA (SF HQ) or Toronto ON (Toronto Hub). Hybrid onsite 3 days a week.
Equity Range: Competitive equity. Details shared during the process.
Location: San Francisco CA (SF HQ). On-site five days a week.
Job Type: Full-Time
Experience: 4 years in ML research engineering or data engineering with evaluation or benchmark work youve shipped
Work Authorization: Must be authorized to work in the United States or Canada. Were not able to sponsor US visas right now. For Canada well consider sponsorship on a case-by-case basis through our Toronto Hub.
Firecrawl is the easiest way to turn the web into data AI agents can use. One API call converts any URL into clean LLM-ready markdown or structured data. Its the boring-hard problem everyone building with LLMs eventually hits solved.
In September 2026 we raised a $75M Series B led by Smash Capital and were spending it building the largest repository of knowledge in the world. We hit 8 figures in ARR in year one and more than doubled it in year two. We have 187k GitHub stars putting us in the top 40 repositories of all time and developers agents and category-defining AI companies build on us every day. Growth like this is rare and were just getting started.
Were a small team punching far above our weight. Everyone here owns a real piece of the product and company end to end and runs it themselves. No hiding behind process or headcount.
This is a place for people who want to work at the frontier: an AI company building the infrastructure other AI companies run on not one bolting AI onto an existing product. We move fast go deep and are building the tools superintelligence will rely on to gather data from the web. That library is called Alexandria and it starts now.
Design and run head-to-head benchmarks of data providers against verified ground truth. For example:
Apollo vs. FullEnrich vs. DataLegion on company details.
FullEnrich vs. DataLegion on finding the right person and their current employer and title.
Built In vs. ZipRecruiter on relevant fresh non-duplicate job listings.
Build and maintain the test datasets and ground truth and keep them from going stale or leaking.
Own the automated harness and the weekly benchmark release. Every run has to be reproducible versioned and defensible.
Measure what buyers actually care about: accuracy coverage freshness speed and cost.
Work with marketing to ship public leaderboard pages that are useful and hold up under scrutiny.
Close the loop into Alexandria so benchmark results change which provider gets called for what.
Keep expanding the categories we benchmark inside Alexandria and beyond it.
Youve shipped evals or benchmarks and you can explain exactly why your results were trustworthy.
Youve done this somewhere that matters: a frontier lab a data company like Scale Surge micro1 or Mercor a third-party benchmark org or a public benchmark project. OSS contributors very welcome.
Youre strong in Python and API integrations and comfortable with messy vendor APIs rate limits and inconsistent schemas.
You know how to build test sets and scoring methods: sampling labeling inter-rater agreement and when to trust an LLM judge and when not to.
You use AI heavily and keep upgrading your own workflow. You ship without waiting for instructions.
You write findings clearly for engineers marketers and the vendors on the other side of the leaderboard.
Someone who wants to run benchmarks someone else designed.
Someone who picks the metric that makes the story look good. Our numbers have to survive the vendors who lose.
A pure researcher who wont build the harness or a pure engineer who wont think hard about methodology.
Someone who needs a fully specced ticket or a quarter to ship the first result.
We operate at an absurd level of urgency because the window for what were building wont stay open forever. If that excites you keep reading. If it doesnt no hard feelings but this role probably isnt for you.
Salary that makes sense: $250000$290000 USD/year / $210000$224000 CAD/year (Toronto) based on impact not tenure
Own a piece: Gain competitive equity in what youre helping build
Generous PTO: 15 days mandatory anything after 24 days just ask (holidays excluded). Take the time you need to recharge
Parental leave: 12 weeks fully paid for all parents
Wellness stipend: $100 USD/month for the gym therapy massages or whatever keeps you human
Learning & Development: Expense up to $1000 USD/year toward anything that helps you grow professionally
Team offsites: A change of scenery minus the trust falls
Sabbatical: 3 paid months off after 4 years do something fun and new
Full coverage no red tape: Medical dental and vision (100% for employees 50% for partner and kids). No weird loopholes just care that works
Life & Disability insurance: Employer-paid basic life and AD&D short-term disability and long-term disability. Coverage for lifes curveballs
Virtual care and a health guide: Teladoc for the couch doctor visit plus Rightway to answer coverage questions and fight billing errors for you
Mental health: Talkspace therapy and psychiatry on your schedule
Fertility and family building: Carrot covering you and your partner
EAP: Free confidential counseling legal and financial consults and online will prep through Guardian
401(k) plan: Retirement might be a ways off but future-you will thank you
Pre-tax benefits: HSA FSA and commuter benefits to help your wallet out a bit
Supplemental options: Extra life and AD&D accident critical illness hospital indemnity plus pet legal and identity protection through MetLife
Full coverage no red tape: Extended health dental and vision through Manulife (Diamond the top tier) 100% employer-paid for you your partner and your kids
Life & Disability insurance: Employer-paid life AD&D short-term disability and long-term disability. Coverage for lifes curveballs
Virtual care: Dialogue Premium so you can see a doctor or nurse from your couch any hour
Mental health: Talkspace Elite therapy and psychiatry on your schedule
Fertility and family building: Carrot covering you and your partner
Retirement: Group RRSP through Wealthsimple so future-you can thank you
SF HQ perks: Snacks drinks team lunches intense ping pong and peak startup energy
E-Bike transportation: A loaner electric bike to get you around the city on us
Toronto Hub perks: Snacks drinks team lunches glass-walled views down University Avenue and a home base steps from Union Station
Transit covered: A PRESTO card loaded for GO Transit subway and streetcar plus station parking if you drive to the train. Winter-proof on us
Application Review: Send us your work. We want the benchmark leaderboard eval or dataset you built and how you knew it was right. We care about what youve shipped not where you went to school.
Intro Chat (25 min): A quick conversation to get to know each other. Well cover what youve been working on what drew you to Firecrawl and what you want next. Time for your questions too.
Technical Chat (45 min): A real problem from our world: design a benchmark that decides whether FullEnrich or DataLegion finds the right persons current title. Well cover where the ground truth comes from and how youd defend the result to the vendor who loses. Come ready to think out loud.
Workflow Chat (30 min): Show us how you actually work: your AI tools your setup and a recent thing you shipped faster than you would have a year ago.
Founder Chat (25 min): Culture pace ownership and how you like to work. Time for your questions too.
Paid Work Trial (40 hours): Ship a small benchmark end to end on a real provider category paid at a contractor rate. Its the truest signal for both sides. Remote-friendly and well flex around your current commitments.
Decision: We move fast after the trial.
If you want to be the person the whole market checks before picking a data provider you should join us.
Apply now.
Required Experience:
IC
About Company
The web crawling, scraping, and search API for AI. Built for scale. Firecrawl delivers the entire internet to AI agents and builders. Clean, structured, and ready to reason with.