Mercor AI Safety Fund Grants
mercor · San Francisco, CA, Mexico · Research
Also hiring in United States, United Kingdom
About the role
About Mercor
Mercor's mission is to organize human intelligence to power the AI economy. We're a leading AI data company, building the layer between human expertise and frontier models. Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models. Mercor's APEX benchmark family measures AI's real-world impact on professional work. Mercor Enterprise brings this same infrastructure to Fortune 500 companies: helping companies capture how their best people actually work, translating that expertise directly back into agents.
Mercor is creating a new category of work where expertise powers AI advancement. Achieving this requires an ambitious, fast-paced and deeply committed team. You’ll work alongside researchers, operators, and AI companies at the forefront of shaping the systems that are redefining society. Mercor is a profitable Series C company valued at $10 billion. We work in-person five days a week in our San Francisco, NYC, or London offices.
Mercor Safety Research Grants $5M
One of the biggest challenges the industry faces today is addressing whether frontier AI is safe enough to deploy. A model can pass safety checks and still act differently in production. This is why investing today in safety research, evals and verification is critical.
Mercor is committing $5 million to fund safety research. The grant supports:
Researcher hours
API credits
Stipends for event and conference attendance
The time of experts from Mercor's platform
This is separate from the Mercor Research Fellowship, which funds benchmark and economics work. You can apply to that HERE.
What we're looking for
We're open to proposals across the full range of safety interests. We're particularly interested in:
Misalignment: deceptive alignment, goal misgeneralization, reward hacking, scheming, and situationally-aware failure modes
Sandbox escape: containment failures, privilege escalation, tool misuse, and agents operating outside their intended scope
Evaluation awareness: models detecting they are being tested and behaving differently under observation than in deployment
Interpretability: understanding what models are actually doing internally, and whether that can be made legible to a human reviewer
Oversight and control: scalable supervision, human-in-the-loop reliability, and what breaks when the system is more capable than its reviewer
Red-teaming methodology: more robust systems for uncovering novel failures
If your work doesn't fit neatly into these, apply anyway. Strong proposals outside this list are welcome.
Why us
Funding for researcher time, API credits, and event attendance
Where useful to the work: access to Mercor's expert network for human grading, red-teaming, and annotation: lawyers, accountants, engineers, scientists, clinicians
Access to Mercor's internal evaluation infrastructure, subject to review
Introductions to Mercor's network of researchers across frontier labs and academia
Who should apply
Independent researchers, academics, PhD students, and small teams
People with a specific, well-scoped question: the grant is built around your proposal, not a generic research rotation
Background in ML, CS, statistics, or an adjacent field (measurement, psychometrics, HCI, security, social science)
Bonus: experience with agentic evaluation, RL environments, adversarial ML, or systems security
We expect grantees to publish. A paper, an open dataset, a public methodology, or a tool the field can use.
How to apply
Submit an Expression of Interest. We expect to see a one- or two-page document. It should contain at least a section on your team, background, and research accomplishments; a section on your proposed research project; and a section on the outputs and impact of the project, with directionally correct timelines and resource requirements.