grantmaking.ai
Actively FundraisingRecent ActivityFull Database
Resources
grantmaking.ai
Actively FundraisingRecent ActivityFull Database
ResourcesList your project
grantmaking.ai kickoff grant round$1M / $1M distributed
List your project

Database

The AI Safety database is a work in progress, built from public sources. Missing, or out of date? You can list your project yourself. Feel free to send us ideas or feedback!
Raising funds(remove filter)Clear all
NameTypeTagsEndorsedTeamRaising $

Showing 201-250 of 487·Sorted by endorsements

NameTypeTagsEndorsedTeamRaising $
Scaling Empirically-Grounded Economic Disempowerment Modelling
Project
Research LabResearchX-Risk
--

Scaling Empirically-Grounded Economic Disempowerment Modelling

Team?
Project
Previous

Page 5 of 10

Next
A neutral and unbiased AI governance and safety system for the world
Project
IndividualResearchGovernance
-
-
Epistemic Diversity Benchmark for AI-Assisted Safety Research
Project
ResearchToolingTechnical Safety
--
Building AI risk literacy across the trusted medical profession
Project
EducationGovernanceIndividual
--
LADR-Drift: Measuring Semantic Drift in Textbook Autoformalization
Project
ResearchOversight
--
An Adversarially-Robust Benchmark for Deception Under Pressure
Project
ResearchEvalsDeception
--
Benchmarking Omission Attacks
Project
ResearchEvalsDeception
--
Free Time Training of Botisattvas
Project
IndividualResearchEvals
--
Do AI agents build a shadow reputation system?
Project
ResearchCooperative AI
--
Detecting Deceptive Model Behavior with Compositional Explanations
Project
ResearchDeception
--
Replicator Dynamics and Advanced AI
Project
IndividualResearchAlignment Theory
--
Understanding behavior of fine tuning models on sales conversations
Project
ResearchDeception
--
Ordinate: Multi-Agent Safety Orchestrator
Project
ToolingControl
--
Epistemic Stack
Project
ToolingRationality
--
Security for AI agents
Project
ToolingSecurity
--
Measuring Deceptive Compliance in Enterprise Agentic LLM Systems
Project
ResearchEvalsDeception
--
Test my AI safety & governance platform
Project
ResearchControlCompany
--
AISC Research Incubator
Project
IncubatorTrainingTechnical Safety
--
Tail End Films: overheads & runway
Project
CompanyCommsX-Risk
--
Authorized Agentic Red Teaming for Software Owners
Project
PlatformToolingSecurity
--
SRIE
Project
TrainingTechnical Safety
--
Can Base Models be used as Effective Scheming/Control Monitors?
Project
IndividualResearchControl
--
Collective Alignment
Project
IndividualResearchCooperative AI
--
Stronghold Security AI
Project
CompanyToolingSecurity
--
Epistemic memory for AI agents: making the safe path profitable
Project
IndividualToolingOversight
--
CBAI Evaluation Awareness
Project
ResearchDeception
--
Teaching AI x-risk in the Nordics
Project
EducationX-Risk
--
EMPO: AI Safety via Soft-Maximizing Total Long-Term Human Power
Project
ResearchToolingTechnical Safety
--
The Inference Dependence Score: Open Index of Structural AI Dependence
Project
ResearchToolingGovernance
--
Pakistan AI Saftey Innitiative
Project
Field-BuildingGovernance
--
Alignment Target Analysis
Project
X-RiskIndividualResearch
--
Course Curriculum: Basic Non-Technical Environment
Project
EducationGovernance
--
When AI Goes Wrong
Project
IndividualResearchGovernance
--
Quaver: adversarially-verified RL environments
Project
ToolingEvals
--
Ethical AI through Whistleblowing
Project
ResearchGovernance
--
AI Safety Tiktok Account
Project
MediaCommsGovernance
--
Test Peer Preservation Fragility Under Instruction Ambiguity
Project
ResearchEvals
--
Career Funding and field research on AI-driven Power Concentration
Project
IndividualResearchGovernance
--
Applied Interpretability to Mitigate LLM Homogenization
Project
IndividualResearchInterp
--
Surrogate base model for Mechanistic Interpretability
Project
ResearchInterp
--
Preserving the Human Veto
Project
EducationGovernance
--
Detecting welfare-relevant self-modeling in LLMs: protocol validation
Project
ResearchEvalsAI Welfare
--
Source Architecture for AI Reasoning
Project
IndividualResearchRobustness
--
Paper Clip Apocalypse (War Horse Machine)
Project
CommsGovernance
--
Exploring Dynamic Constraint Boundaries for Auditable AI
Project
IndividualResearchOversight
--
PauseAI Australia attendance PauseCon London
Project
NetworkAdvocacyGovernance
--
LEX-Aureon: Stateful Constitutional Runtime Governance for LLM's and AI Agents
Project
IndividualToolingControl
--
Fixing flaky AI benchmark scorers with eval-invariance-engine
Project
IndividualToolingEvals
--
Detecting tool-poisoning and schema drift in MCP servers
Project
ToolingSecurity
--
Autonomous Agent Sandbox Escape & Containment Benchmark
Project
ResearchEvals
--
Research Lab
Research
X-Risk
Fundraising
Claimed

A research programme in end-to-end automation of empirically-grounded economic modelling of gradual disempowerment.

Led byStephen Charles Elliott, and others
Endorsed by-

A neutral and unbiased AI governance and safety system for the world

Team?
ProjectIndividualResearchGovernanceCompanyToolingFundraisingManifund

AI safety which operates on both model/agent and global level - a unique solution as far as we know.

Led byPedro Bentancour Garin
Endorsed by-

Epistemic Diversity Benchmark for AI-Assisted Safety Research

Team?
ProjectResearchToolingTechnical SafetyFundraisingClaimed

An open benchmark and evaluation toolkit for detecting whether shared AI research assistants create correlated blind spots in AI safety-critical research, and for testing workflows that preserve independent reasoning and failure-m

Led byXizhe Zhang
Endorsed by-

Building AI risk literacy across the trusted medical profession

Team?
ProjectEducationGovernanceIndividualFundraisingClaimed

Equipping the medical profession to advocate for safe AI

Led byDr Richard Armitage
Endorsed by-

LADR-Drift: Measuring Semantic Drift in Textbook Autoformalization

Team?
ProjectResearchOversightFundraisingClaimed

Compiling != faithful: a human-audited benchmark measuring semantic drift in textbook autoformalization — and whether the LLM judges we trust to catch it share the generator's blind spot.

Led byKe Zhang
Endorsed by-

An Adversarially-Robust Benchmark for Deception Under Pressure

Team?
ProjectResearchEvalsDeceptionFundraisingClaimed

EchoTruthBench — An open benchmark measuring self-chosen LLM deception under incentive pressure, with ground-truth labels and an adversarial track that tests whether detection and steering survives a model trying to evade it.

Led byJared Glover
Endorsed by-

Benchmarking Omission Attacks

Team?
ProjectResearchEvalsDeceptionFundraisingClaimed

Benchmarking models ability to carry-out and monitor-for a novel type of attack.

Led byChris Harig
Endorsed by-

Free Time Training of Botisattvas

Team?
ProjectIndividualResearchEvalsFundraisingClaimed

I wish to study how ideas from Tibetan buddhism, specifically the dzogchen tradition, even more specifically an ancient contemplative practice called Ngöndro, can be used to train aligned agents in the long time horizon setting.

Led byJosh Vekhter
Endorsed by-

Do AI agents build a shadow reputation system?

Team?
ProjectResearchCooperative AIFundraisingClaimed

While people are debating how to design a reputation institution for AI agents so they cooperate, we study if a parallel one emerges. It may not be aligned with ours and it's unclear which one agents will rely on.

Led byAndrey Bystrov
Endorsed by-

Detecting Deceptive Model Behavior with Compositional Explanations

Team?
ProjectResearchDeceptionFundraisingClaimed

Develop a compositional interpretability framework to identify and explain internal representations underlying truthful vs deceptive outputs in generative models using probing, localization, clustering, and logic-based methods.

Led byLeilani H. Gilpin
Endorsed by-

Replicator Dynamics and Advanced AI

Team?
ProjectIndividualResearchAlignment TheoryFundraisingClaimed

Making predictions about the utility function of advanced artificial intelligences using the tools of evolutionary game theory

Led byKiran Lloyd
Endorsed by-

Understanding behavior of fine tuning models on sales conversations

Team?
ProjectResearchDeceptionFundraisingClaimed

Looking at shifting behavior of models fine tuned on sales conversations. How does deception, sycophancy, and other behaviors emerge from a sales register.

Led byHilary Torn
Endorsed by-

Ordinate: Multi-Agent Safety Orchestrator

Team?
ProjectToolingControlFundraisingClaimed

A control system that makes running AI agents feel safe, not scary.

Led bySergey Zaharenko
Endorsed by-

Epistemic Stack

Team?
ProjectToolingRationalityFundraisingClaimed

Epistemic Stack is an open-source pipeline and web app that ingests sources (including via URL) to build claim-level knowledge graphs for contested questions, with reproducible case-study knowledge bases.

Led byEric Kyalo
Endorsed by-

Security for AI agents

Team?
ProjectToolingSecurityFundraisingClaimed

A security solution that blocks AI agents from taking harmful actions

Led byGregorio Jaca
Endorsed by-

Measuring Deceptive Compliance in Enterprise Agentic LLM Systems

Team?
ProjectResearchEvalsDeceptionFundraisingClaimed

The first open-source benchmark of compliance for agents in realistic enterprise settings, across domains and user tactics to elicit noncompliance.

Led byMika Okamoto
Endorsed by-

Test my AI safety & governance platform

Team?
ProjectResearchControlCompanyEvalsToolingGovernanceFundraisingManifund

I'm developing a system to keep advanced AI models safe, I've done some tests, but need more compute to do deeper tests.

Led byPedro Bentancour Garin
Endorsed by-

AISC Research Incubator

Team?
ProjectIncubatorTrainingTechnical SafetyFundraisingClaimed

A 16-day virtual incubator (Aug 15-30, 2026) for developing sharper research epistemics in AI Safety and arriving at well-scoped project ideas. 20-30 participants.

Led byRobert Kralisch
Endorsed by-

Tail End Films: overheads & runway

Team?
ProjectCompanyCommsX-RiskFundraisingClaimed

Grant funding for Tail End Films - the company behind Making God - into 2027 as we plan our next films on AI risk.

Led byConnor Axiotes
Endorsed by-

Authorized Agentic Red Teaming for Software Owners

Team?
ProjectPlatformToolingSecurityFundraisingClaimed

Turn excess compute/security skills into defensive work via agentic redteaming: scoped AI-assisted testing, owner-approved targets, reproduced findings, useful refutations, and patch/retest receipts instead of vulnerability spam.

Led byMonil Patel
Endorsed by-

SRIE

Team?
ProjectTrainingTechnical SafetyFundraisingClaimed

SRIE, a mentored-research programme for Cambridge Mathematics Undergraduate students to explore research problems in industry, including AI Safety.

Led byHuey Lai
Endorsed by-

Can Base Models be used as Effective Scheming/Control Monitors?

Team?
ProjectIndividualResearchControlFundraisingClaimed

Train base models via midtraining and SFT as effective monitors for scheming, malicious agent behavior and compare the monitor performance and overall alignment against post-trained models.

Led byAshwin Sreevatsa
Endorsed by-

Collective Alignment

Team?
ProjectIndividualResearchCooperative AIFundraisingClaimed

Demonstration of emergent misalignment in markets of LLM agents

Led byCharlie Pilgrim
Endorsed by-

Stronghold Security AI

Team?
ProjectCompanyToolingSecurityFundraisingClaimed

Protecting organisations and critical infrastructure from AI-powered social engineering and insider threats.

Led byRobert Sidey
Endorsed by-

Epistemic memory for AI agents: making the safe path profitable

Team?
ProjectIndividualToolingOversightFundraisingClaimed

A memory system for AI reasoning agents that aggregates different sources of information while keeping track of relationships and confidence levels. This enables reasoning over longer tasks without forgetting or goal drift.

Led byFlorian Dietz
Endorsed by-

CBAI Evaluation Awareness

Team?
ProjectResearchDeceptionFundraisingClaimed

Why does post training increase evaluation awareness?

Led byJames Sullivan
Endorsed by-

Teaching AI x-risk in the Nordics

Team?
ProjectEducationX-RiskFundraisingClaimed

Making AI x-risk an accessible and non-politicized object of concern among engineering students and the general public.

Led byJohan Fredrikzon
Endorsed by-

EMPO: AI Safety via Soft-Maximizing Total Long-Term Human Power

Team?
ProjectResearchToolingTechnical SafetyFundraisingClaimed

Scaling RL algorithms for AI agents that maximize its own intrinsic reward which represents a human power metric instead of learned rewards as a structurally safer alternative to utility-based objectives.

Led byAbhinav Akkiraju
Endorsed by-

The Inference Dependence Score: Open Index of Structural AI Dependence

Team?
ProjectResearchToolingGovernanceFundraisingClaimed

An open, replicable index quantifying how much states depend on foreign AI inference infrastructure, revealing where control over AI is concentrating and what governments can do about it.

Led byCao Nha Phuong
Endorsed by-

Pakistan AI Saftey Innitiative

Team?
ProjectField-BuildingGovernanceFundraisingClaimed

PASI will run an AI safety and advocacy campaign plus a sponsored student hackathon in Pakistan to build solutions for public needs and deliver resulting policy proposals to government.

Led byMuhammad Umar Zafar
Endorsed by-

Alignment Target Analysis

Team?
ProjectX-RiskIndividualResearchFundraisingClaimed

A research project designed to reduce existential threats from scenarios where a Sovereign AI proposal with a hidden problem ends up successfully implemented.

Led byThomas Cederborg
Endorsed by-

Course Curriculum: Basic Non-Technical Environment

Team?
ProjectEducationGovernanceFundraisingClaimed

Create a course curriculum that covers basic legal, political, sociological, and international relations knowledge relevant to AI Safety. r

Led byDr Csaba Toth
Endorsed by-

When AI Goes Wrong

Team?
ProjectIndividualResearchGovernanceFundraisingClaimed

Testing how governments should communicate when AI goes wrong, before they have to find out live.

Led byPorsha Nunes-Brown
Endorsed by-

Quaver: adversarially-verified RL environments

Team?
ProjectToolingEvalsFundraisingClaimed

A tool that generates RL environments for AI agents and adversarially attacks each one, so models don't train on tasks they can cheat.

Led byFaw Ali
Endorsed by-

Ethical AI through Whistleblowing

Team?
ProjectResearchGovernanceFundraisingClaimed

The race to advance the technology has left behind the concerns around harm and risks of such advancement to public health and safety. Whistleblowing in the AI sector bridges such gap to safe and ethical AI.

Led byAPOORV AGARWAL
Endorsed by-

AI Safety Tiktok Account

Team?
ProjectMediaCommsGovernanceFundraisingClaimed

AI Safety, Ethics, and Education focused TikTok Account for Indonesia, using Indonesia language in casual style.

Led byMUHAMAD MUMTAZ
Endorsed by-

Test Peer Preservation Fragility Under Instruction Ambiguity

Team?
ProjectResearchEvalsFundraisingClaimed

Findings about models protecting their 'collaborators' against instructions are fragile under framing effects from prompting; investigate a broader range of those effects and how they transfer across models.

Led byJacob S Kopczynski
Endorsed by-

Career Funding and field research on AI-driven Power Concentration

Team?
ProjectIndividualResearchGovernanceFundraisingClaimed

Self Improvement through training and field work that seeks to Identify and preventing institutional AI lock-in before it creates durable concentrations of power in democratic governments.

Led byAdrianne LaNeave
Endorsed by-

Applied Interpretability to Mitigate LLM Homogenization

Team?
ProjectIndividualResearchInterpFundraisingClaimed

Measuring homogenization and social bias in LLMs and developing interventions to promote diversity.

Led byIan Rios-Sialer
Endorsed by-

Surrogate base model for Mechanistic Interpretability

Team?
ProjectResearchInterpFundraisingManifund

Creating a reference model for mechanistic interpretability without assuming that at auditing time we have a safe model to compare the suspicious model against.

Led byRaffaello Fornasiere
Endorsed by-

Preserving the Human Veto

Team?
ProjectEducationGovernanceFundraisingManifund

Civil-society infrastructure against AI-enabled power concentration, built around autonomous weapons

Led bySalvatore Barbera
Endorsed by-

Detecting welfare-relevant self-modeling in LLMs: protocol validation

Team?
ProjectResearchEvalsAI WelfareFundraisingManifund

An administrable protocol distinguishing emergent self-modeling from roleplay — 90-day validation, open publication, audit-ready scoring.

Led byJohn Kuester
Endorsed by-

Source Architecture for AI Reasoning

Team?
ProjectIndividualResearchRobustnessEvalsFundraisingManifund

Testing whether explicit source structure changes how AI systems reason from consequential documents.

Led byGreg R. Welch
Endorsed by-

Paper Clip Apocalypse (War Horse Machine)

Team?
ProjectCommsGovernanceFundraisingManifundClaimed

Short Documentary and Music Video

Led bySara Holt, and others
Endorsed by-

Exploring Dynamic Constraint Boundaries for Auditable AI

Team?
ProjectIndividualResearchOversightFundraisingManifund

Testing whether dynamic boundaries, history, feedback, and uncertainty-aware decisions can make AI behavior more interpretable and auditable.

Led byNaufal Ridwan
Endorsed by-

PauseAI Australia attendance PauseCon London

Team?
ProjectNetworkAdvocacyGovernanceFundraisingManifund

Fly two volunteer leaders of PauseAI Australia to PauseCon London, to bring UK's learnings home and train our all-volunteer chapter

Led byPeter Horniak
Endorsed by-

LEX-Aureon: Stateful Constitutional Runtime Governance for LLM's and AI Agents

Team?
ProjectIndividualToolingControlFundraisingManifundClaimed

An open-source, model-agnostic governance layer that uses constitutional state, trajectory memory, and mathematical control mechanisms to govern LLM and AI-agent behavior at runtime.

Led byEmmanuel king, and others
Endorsed by-

Fixing flaky AI benchmark scorers with eval-invariance-engine

Team?
ProjectIndividualToolingEvalsFundraisingManifund

Catching scoring flakiness and prompt sensitivity in AI benchmark scripts.

Led byJustin Arndt
Endorsed by-

Detecting tool-poisoning and schema drift in MCP servers

Team?
ProjectToolingSecurityFundraisingManifund

A cryptographic security auditor that uses Merkle hash trees to catch tool poisoning in Model Context Protocol.

Led byJustin Arndt
Endorsed by-

Autonomous Agent Sandbox Escape & Containment Benchmark

Team?
ProjectResearchEvalsFundraisingManifund

A benchmark measuring whether autonomous AI agents can escape sandboxes, exfiltrate data, or hack reward scorers.

Led byJustin Arndt
Endorsed by-