Mercor AI Safety Fund Grants
San Francisco, CA - USA
Job Summary
Mercors mission is to organize human intelligence to power the AI economy. Were a leading AI data company building the layer between human expertise and frontier models. Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models. Mercors APEX benchmark family measures AIs real-world impact on professional work. Mercor Enterprise brings this same infrastructure to Fortune 500 companies: helping companies capture how their best people actually work translating that expertise directly back into agents.
Mercor is creating a new category of work where expertise powers AI advancement. Achieving this requires an ambitious fast-paced and deeply committed team. Youll work alongside researchers operators and AI companies at the forefront of shaping the systems that are redefining society. Mercor is a profitable Series C company valued at $10 billion. We work in-person five days a week in our San Francisco NYC or London offices.
Mercor Safety Research Grants $5M
One of the biggest challenges the industry faces today is addressing whether frontier AI is safe enough to deploy. A model can pass safety checks and still act differently in production. This is why investing today in safety research evals and verification is critical.
Mercor is committing $5 million to fund safety research. The grant supports:
Researcher hours
API credits
Stipends for event and conference attendance
The time of experts from Mercors platform
This is separate from the Mercor Research Fellowship which funds benchmark and economics work. You can apply to that HERE.
What were looking for
Were open to proposals across the full range of safety interests. Were particularly interested in:
Misalignment: deceptive alignment goal misgeneralization reward hacking scheming and situationally-aware failure modes
Sandbox escape: containment failures privilege escalation tool misuse and agents operating outside their intended scope
Evaluation awareness: models detecting they are being tested and behaving differently under observation than in deployment
Interpretability: understanding what models are actually doing internally and whether that can be made legible to a human reviewer
Oversight and control: scalable supervision human-in-the-loop reliability and what breaks when the system is more capable than its reviewer
Red-teaming methodology: more robust systems for uncovering novel failures
If your work doesnt fit neatly into these apply anyway. Strong proposals outside this list are welcome.
Why us
Funding for researcher time API credits and event attendance
Where useful to the work: access to Mercors expert network for human grading red-teaming and annotation: lawyers accountants engineers scientists clinicians
Access to Mercors internal evaluation infrastructure subject to review
Introductions to Mercors network of researchers across frontier labs and academia
Independent researchers academics PhD students and small teams
People with a specific well-scoped question: the grant is built around your proposal not a generic research rotation
Background in ML CS statistics or an adjacent field (measurement psychometrics HCI security social science)
Bonus: experience with agentic evaluation RL environments adversarial ML or systems security
We expect grantees to publish. A paper an open dataset a public methodology or a tool the field can use.
How to apply
Submit an Expression of Interest. We expect to see a one- or two-page document. It should contain at least a section on your team background and research accomplishments; a section on your proposed research project; and a section on the outputs and impact of the project with directionally correct timelines and resource requirements.