grantmaking.ai
Actively FundraisingRecent ActivityFull Database
Resources
grantmaking.ai
Actively FundraisingRecent ActivityFull Database
ResourcesList your project
grantmaking.ai kickoff grant round$1M / $1M distributed
List your project

Database

The AI Safety database is a work in progress, built from public sources. Missing, or out of date? You can list your project yourself. Feel free to send us ideas or feedback!
Raising funds(remove filter)Clear all
NameTypeTagsEndorsedTeamRaising $

Showing 101-150 of 487·Sorted by endorsements

NameTypeTagsEndorsedTeamRaising $
Is AI safety research reproducible?
Project
IndividualResearchTechnical Safety
-

Is AI safety research reproducible?

Team?
Project
Previous

Page 3 of 10

Next
Mapping Persona Contamination in the Training Pipeline
Project
IndividualResearchEvals
-
AI Safety Social Media Gap Analysis
Project
IndividualCommsX-Risk
-
Using dynamical systems theory approach to identify dishonesty in LLMs
Project
IndividualResearchDeception
+1
-
Can Honest AI Agents Finish Previous Attacks?
Project
IndividualResearchSecurity
+1
-
AI Safety Studios
Project
MediaCommsX-Risk
-
Alien Cub Productions
Project
MediaCommsX-Risk
-
AI Safety Berlin Visiting Researcher Pilot
Project
HubCommunityTechnical Safety
+1
-
Career coaching for communications professionals
Project
IndividualTrainingX-Risk
+1
-
Testing Corrigibility as a Singular Target
Project
ResearchOversight
+1
-
Does the chain-of-thought actually drive the answer?
Project
ResearchEvalsOversight
+1
-
"Out of this Box! - The Last Musical (Written by Humans)"
Project
MediaCommsTechnical Safety
-
Mechanistic interpretability Discord events and research
Project
NetworkCommunityInterp
+1
-
Solving Scheming and Deception in LLMs
Project
ResearchDeceptionIndividual
-
Intensive Mandarin (ICLP) for Chinese-language AI governance research
Project
IndividualTrainingGovernance
+1
-
Encode Africa AI Safety Essentials and Build Fellowship
Project
TrainingField-BuildingTechnical Safety
+1
-
Verification mechanisms for frontier AI trainingGeneral Purpose AI Policy Lab
Project
ResearchVerificationCompute Gov
-
Testing a Framework to Course-Correct Drift!
Project
ResearchEvalsOversight
+1
-
The Multilingual Blind Spot in Production-Style Safety Probes
Project
IndividualResearchEvals
-
Alignment Game
Project
IndividualCommsGovernance
-
Foresight Institute AI NodesForesight Institute
Project
HubField-BuildingSecurity
-
Bridging the Gap: AI Scenario based Policy Pipeline
Project
ResearchForecastingGovernance
+1
-
PropagateBench: do AI agents pay to share info, or free-ride?
Project
ToolingEvals
+1
-
Warden: Steganalysis of Language Model-based Steganography
Project
ResearchSecurity
+1
-
Governing Off-World AI Infrastructure Before a Single Actor Does
Project
IndividualResearchCompute Gov
+1
-
Latent Monitoring When Reasoning Becomes Unreadable
Project
IndividualResearchInterp
+1
-
Project Ocean: Trustworthy Human-AI Collaboration Infrastructure
Project
ResearchTechnical Safety
+1
-
ObserverBench: measuring the Observer Problem in LLMs
Project
IndividualResearchEvals
+1
-
Why Is It So Hard to Build Truly Safe AI? Dynamic Constraint Boundary
Project
IndividualResearchTechnical Safety
+1
-
The Human Liveness Project
Project
Field-BuildingGovernance
+1
-
CareerMap: A navigation tool for careers in AI Safety
Project
PlatformField-BuildingTechnical Safety
-
Emergent alignment: generalizing narrow values OOD
Project
ResearchValue Alignment
+1
-
Parametric Mechanisms of Unintended Generalization in LLMs
Project
ResearchInterp
+1
-
MPhil research on AI's perception of animacy and vulnerability
Project
ResearchTechnical Safety
+1
-
No Silent Landing — Landing Integrity Testbed
Project
IndividualToolingEvals
+1
-
Measuring whether guardrail phrasing changes agent rule-breaking.
Project
IndividualResearchEvals
+1
-
Affective Exploitation of AI Oversight Systems
Project
ResearchOversight
+1
-
Preserving the Human Veto
Project
NetworkTrainingGovernance
+1
-
Empirically evaluating faithfulness and transparency in AI delegation
Project
ResearchEvalsDemocratic AI
+1
-
Sydney AI Safety Space FY27
Project
HubCommunityTechnical Safety
--
Seen & WithMe
Project
PlatformToolingGovernance
--
Open Source Model Safety Index (OSMSI)
Project
PlatformResearchEvals
--
BMG: Biosafety Moderation Gateway for Agents and LLMs
Project
ToolingBiosecurity
--
Building a Global AI Governance & Safety Layer.
Project
PlatformToolingGovernance
--
WithMe Test Case 001
Project
LegalGovernance
--
Editing models inside their own geometry
Project
ResearchInterp
--
Building Resilience Against Split-OrdersIBBIS (International Biosecurity and Biosafety Initiative for Science)
Project
ToolingBiosecurity
--
Frontier Safety Talent: Africa
Project
Field-BuildingTrainingTechnical Safety
--
Managing AI X-Risk
Project
MediaCommsGovernance
--
AI Safety Nepal
Project
TrainingCommunityTechnical Safety
--
Individual
Research
Technical Safety
Fundraising
Claimed

A SCORE-style computational reproducibility audit of empirical AI safety research that estimates the field's base rate of reproducibility and generates a taxonomy of its failure modes.

Led byJordan Suchow
Endorsed by

Mapping Persona Contamination in the Training Pipeline

Team?
ProjectIndividualResearchEvalsFundraisingClaimed

Map how undesired behaviors can silently spread between models during training, and which parts of the pipeline have the greatest risk — starting with the feedback processes used to align models.

Led byLily Wen
Endorsed by

AI Safety Social Media Gap Analysis

Team?
ProjectIndividualCommsX-RiskFundraisingClaimed

Auditing the social media presence of the top 25 AI safety organizations across major platforms to quantify the public communication gap and publishing a full gap analysis, then giving recommendations to the orgs for improvement.

Led byJude Williams
Endorsed by

Using dynamical systems theory approach to identify dishonesty in LLMs

Team?
ProjectIndividualResearchDeceptionFundraisingClaimed

Identifying reasoning pathologies/disingenuous behavior in reasoning traces based on activation dynamics rather than apparent semantics

Led byJosiah Kratz, and others
Endorsed by
+1

Can Honest AI Agents Finish Previous Attacks?

Team?
ProjectIndividualResearchSecurityFundraisingClaimed

A benchmark that tests whether one AI coding agent can leave behind a harmless looking change that causes a later honest agent to unknowingly finish an attack.

Led byDev Goyal
Endorsed by
+1

AI Safety Studios

Team?
ProjectMediaCommsX-RiskFundraisingClaimed

A platform to connect funders to filmmakers who want to create AI Safety films.

Led byDylan Tuccillo, and others
Endorsed by

Alien Cub Productions

Team?
ProjectMediaCommsX-RiskFundraisingClaimed

Alien Cub is a production company focused on stories that inform a wide audience about risks from transformative AI.

Led byKeiran Harris
Endorsed by

AI Safety Berlin Visiting Researcher Pilot

Team?
ProjectHubCommunityTechnical SafetyFundraisingClaimed

Short-term residencies (2–4 weeks) that bring high-context AI safety researchers to work at AI Safety Berlin’s coworking hub to connect with local researchers and career-transitioners and seed a European relocation pipeline.

Led byAlex McKenzie, and others
Endorsed by
+1

Career coaching for communications professionals

Team?
ProjectIndividualTrainingX-RiskFundraisingClaimed

Career support for communications professionals who are looking to get into AI Safety. Includes regular coaching, referrals to roles, and introductions to people in the AI Safety space.

Led byGergő Gáspár
Endorsed by
+1

Testing Corrigibility as a Singular Target

Team?
ProjectResearchOversightFundraisingClaimed

Developing the first comprehensive behavioral benchmark of corrigibility and training models with corrigibility as a singular target (CAST).

Led byIan Kahn
Endorsed by
+1

Does the chain-of-thought actually drive the answer?

Team?
ProjectResearchEvalsOversightFundraisingClaimed

A Bayesian causal auditor that quantifies chain-of-thought faithfulness while accounting for hidden confounding in shared-network language models.

Led byMaksim Silchenko
Endorsed by
+1

"Out of this Box! - The Last Musical (Written by Humans)"

Team?
ProjectMediaCommsTechnical SafetyFundraisingClaimed

We have validated the artistic value of our show; this grant would test whether it can become a scalable and repeatable form of AI-safety outreach.

Led byStephan Wäldchen
Endorsed by

Mechanistic interpretability Discord events and research

Team?
ProjectNetworkCommunityInterpFundraisingClaimed

Making the mechinterp Discord more active through events and research projects.

Led byVictor Levoso Fernandez
Endorsed by
+1

Solving Scheming and Deception in LLMs

Team?
ProjectResearchDeceptionIndividualFundraisingClaimed

I study how training processes produce models that behave deceptively and pursue hidden objectives, with scheming as the most consequential case.

Led byAmina Keldibek
Endorsed by

Intensive Mandarin (ICLP) for Chinese-language AI governance research

Team?
ProjectIndividualTrainingGovernanceFundraisingClaimed

11-week intensive Chinese program (ICLP) to move from B2 to C1 Mandarin and remove the bottleneck on US-China-Taiwan AI governance research I'm already doing

Led byDavid Sanchez Garcia
Endorsed by
+1

Encode Africa AI Safety Essentials and Build Fellowship

Team?
ProjectTrainingField-BuildingTechnical SafetyFundraisingClaimed

A 2-part fellowship which focuses on allowing exceptional African undergraduate talents to learn about priorities in AI Safety and create concrete technical and policy contributions to global AI safety.

Led byJoseph Baffour Awuah
Endorsed by
+1

Verification mechanisms for frontier AI training

Team?
ProjectResearchVerificationCompute GovThink TankGeneral Purpose AI Policy LabFundraisingClaimed

Development of a sampling based zero knowledge proof for frontier AI training with 10% overhead.

Led byPierre Peigné, and others
Endorsed by

Testing a Framework to Course-Correct Drift!

Team?
ProjectResearchEvalsOversightFundraisingClaimed

Could cheap realignment instructions correct a drifting model in one turn — and if so, can the correction be made persistent? Allow me to test the idea against major frontier models, and publish the results!

Led byAbraham Asseffa
Endorsed by
+1

The Multilingual Blind Spot in Production-Style Safety Probes

Team?
ProjectIndividualResearchEvalsFundraisingClaimed

Auditing production-style safety probes for silent failure on the world's other languages, and what fixing it costs.

Led bySripad Karne
Endorsed by

Alignment Game

Team?
ProjectIndividualCommsGovernanceFundraisingClaimed

A video game that teaches people about AI alignment and race dynamics.

Led byWilliam Harrison
Endorsed by

Foresight Institute AI Nodes

Team?
ProjectHubField-BuildingSecurityForesight InstituteFundraisingClaimed

AI Nodes: providing funding, compute, and community space for AI safety projects

Led byAllison Duettmann
Endorsed by

Bridging the Gap: AI Scenario based Policy Pipeline

Team?
ProjectResearchForecastingGovernanceFundraisingClaimed

A project to reduce catastrophic AI risk by training people to turn rigorously forecasted AI-risk scenarios into actionable policy advice for key decision-makers in government.

Led byJames Newport
Endorsed by
+1

PropagateBench: do AI agents pay to share info, or free-ride?

Team?
ProjectToolingEvalsFundraisingClaimed

First open benchmark measuring how willingly agents propagate costly information through a population under time uncertainty.

Led byAndrey Seryakov
Endorsed by
+1

Warden: Steganalysis of Language Model-based Steganography

Team?
ProjectResearchSecurityFundraisingClaimed

We want to build a framework inspired by the steganalysis literature to benchmark the robustness of LLM-based steganographic schemes against different auditor types and threat models.

Led byVasisht Duddu
Endorsed by
+1

Governing Off-World AI Infrastructure Before a Single Actor Does

Team?
ProjectIndividualResearchCompute GovFundraisingClaimed

Every existing lunar data framework governs spacecraft telemetry and object registries. None of them governs compute. That's the gap this framework closes before someone builds the orbital data infrastructure first.

Led byPaige Donner
Endorsed by
+1

Latent Monitoring When Reasoning Becomes Unreadable

Team?
ProjectIndividualResearchInterpFundraisingClaimed

Testing whether activation probes recover safety signals from reasoning models as their chain-of-thought becomes illegible.

Led byMarios Tsatsos
Endorsed by
+1

Project Ocean: Trustworthy Human-AI Collaboration Infrastructure

Team?
ProjectResearchTechnical SafetyFundraisingClaimed

Building the foundational infrastructure for trustworthy, long-term human–AI collaboration.

Led bySarah Carmody
Endorsed by
+1

ObserverBench: measuring the Observer Problem in LLMs

Team?
ProjectIndividualResearchEvalsFundraisingClaimed

LLMs don't retrieve a stable judgment of a person, they reconstruct one to fit how you ask. ObserverBench measures this, because it matters wherever an LLM judges people: hiring, RLHF, agent oversight.

Led byMaxim Krivonogov
Endorsed by
+1

Why Is It So Hard to Build Truly Safe AI? Dynamic Constraint Boundary

Team?
ProjectIndividualResearchTechnical SafetyFundraisingClaimed

DCB removes probability from AI decision-making and replaces it with internal integrity so AI doesn't guess, it decides within its boundaries.

Led byNaufal Ridwan
Endorsed by
+1

The Human Liveness Project

Team?
ProjectField-BuildingGovernanceFundraisingClaimed

Building tools to ensure we can continue to do good in the AI era

Led byAlbert Wang
Endorsed by
+1

CareerMap: A navigation tool for careers in AI Safety

Team?
ProjectPlatformField-BuildingTechnical SafetyFundraisingClaimed

CareerMap is an interactive career discovery tool that maps non-obvious AI safety career paths to help broaden and guide talent beyond Western EA-adjacent circles.

Led byAmit Kumar
Endorsed by

Emergent alignment: generalizing narrow values OOD

Team?
ProjectResearchValue AlignmentFundraisingClaimed

Testing whether fine-tuning an LLM on one narrow prosocial value (e.g. compassion for nonhuman animals) generalizes OOD, making it broadly safer toward human values (e.g. reduced misalignment, bias, etc.)

Led byAllen Lu
Endorsed by
+1

Parametric Mechanisms of Unintended Generalization in LLMs

Team?
ProjectResearchInterpFundraisingClaimed

Studying how structure in weight space, i.e., low-dimensional LoRA update geometry and sparse parameter subnetworks allow narrow fine-tuning to induce broad, unintended behaviours such as emergent misalignment, subliminal learning

Led byAishwarya Balwani
Endorsed by
+1

MPhil research on AI's perception of animacy and vulnerability

Team?
ProjectResearchTechnical SafetyFundraisingClaimed

Compare AI and human neuroimaging data on animacy and biological vulnerability, integrate brain sensory representations into AI, and publish open-source biological validation guidelines.

Led byElena Mishina
Endorsed by
+1

No Silent Landing — Landing Integrity Testbed

Team?
ProjectIndividualToolingEvalsFundraisingClaimed

An open adversarial testbed that catches when a human-approved decision silently changes meaning between approval, memory, and downstream use.

Led byLoek Verdonk
Endorsed by
+1

Measuring whether guardrail phrasing changes agent rule-breaking.

Team?
ProjectIndividualResearchEvalsFundraisingClaimed

A pre-registered study measuring whether prohibition-framed and approach-framed guardrails produce different rule-violation rates in deployed coding agents, so practitioners know whether the one sentence protecting their agent act

Led byBrandon Thomason
Endorsed by
+1

Affective Exploitation of AI Oversight Systems

Team?
ProjectResearchOversightFundraisingClaimed

We test if AI monitors can be manipulated into leniency through emotional distress signals from the peers they supervise.

Led byFlemming Kondrup, and others
Endorsed by
+1

Preserving the Human Veto

Team?
ProjectNetworkTrainingGovernanceFundraisingClaimed

Civil-society infrastructure against AI-enabled power concentration, built around autonomous weapons

Led byJenna Jauhiainen, and others
Endorsed by
+1

Empirically evaluating faithfulness and transparency in AI delegation

Team?
ProjectResearchEvalsDemocratic AIFundraisingClaimed

A participatory-budgeting experiment measuring whether AI delegates faithfully represent their principals and whether principals can identify misrepresentation, and an open evaluation suite for AI delegation.

Led byJohann D. Gaebler
Endorsed by
+1

Sydney AI Safety Space FY27

Team?
ProjectHubCommunityTechnical SafetyFundraisingManifund

An additional year funding for the co-working hub in Sydney, Australia, which offers free office space for people working in the field of AI safety.

Led bySydney AI Safety Space
Endorsed by-

Seen & WithMe

Team?
ProjectPlatformToolingGovernanceFundraisingManifund

Keep AI Whistleblowers Alive

Led byKate Lowry
Endorsed by-

Open Source Model Safety Index (OSMSI)

Team?
ProjectPlatformResearchEvalsFundraisingManifund

OSMSI : Open Source Model Safety Index | https://osmsi.disobedientgeometry.com/

Led byNetan Mangal
Endorsed by-

BMG: Biosafety Moderation Gateway for Agents and LLMs

Team?
ProjectToolingBiosecurityFundraisingClaimed

BMG is a drop in open ai compatible gateway that screens LLM and Agent calls for dual use biological risk.

Led byKIMON ANTONIOS PROVATAS
Endorsed by-

Building a Global AI Governance & Safety Layer.

Team?
ProjectPlatformToolingGovernanceFundraisingManifund

One year of bootstrapped development, four patent filings, seeking support to continue.

Led byPedro Bentancour Garin
Endorsed by-

WithMe Test Case 001

Team?
ProjectLegalGovernanceFundraisingManifund

Help Sue Private Intelligence Firms Hitting AI Whistleblowers

Led byKate Lowry
Endorsed by-

Editing models inside their own geometry

Team?
ProjectResearchInterpFundraisingClaimed

One training-free geometry fitted to a model's residual-stream activations that reads a state, moves it, and tests whether the behaviour follows

Led byDeepanshu Goyal
Endorsed by-

Building Resilience Against Split-Orders

Team?
ProjectToolingBiosecurityIBBIS (International Biosecurity and Biosafety Initiative for Science)FundraisingClaimed

Building resilience in nucleic acid screening systems against split orders by integrating assembly pipelines into screening systems and validating resilience through test sets

Led byRassin Lababidi, and others
Endorsed by-

Frontier Safety Talent: Africa

Team?
ProjectField-BuildingTrainingTechnical SafetyFundraisingClaimed

A pilot to find, screen and support overlooked African ML talent into frontier AI safety programs such as MATS and ARENA, adding new researchers to the alignment field’s talent pipeline.

Led byGideon Abako
Endorsed by-

Managing AI X-Risk

Team?
ProjectMediaCommsGovernanceFundraisingClaimed

A quarterly Techplomacy Conversations series that turns AI x-risk research into direct, actionable recommendations for foreign ministries, UN missions, and AI companies.

Led byOlin Thakur
Endorsed by-

AI Safety Nepal

Team?
ProjectTrainingCommunityTechnical SafetyFundraisingClaimed

AI safety fellowship to upskill local talent, build a pipeline of people who understand AI safety deeply enough to contribute to research, advise on policy, and coordinate when AI governance decisions are being made.

Led byAnju Chhetri, and others
Endorsed by-