Database
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
Showing 101-150 of 487·Sorted by endorsements
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
| Project | Individual | - |
| Project | Individual | - |
| Project | Individual | - |
| Project | Individual | +1 | - |
| Project | Individual | +1 | - |
| Project | Media | - |
| Project | Media | - |
| Project | Hub | +1 | - |
| Project | Individual | +1 | - |
| Project | Research | +1 | - |
| Project | Research | +1 | - |
| Project | Media | - |
| Project | Network | +1 | - |
| Project | Research | - |
| Project | Individual | +1 | - |
| Project | Training | +1 | - |
| Project | Research | - |
| Project | Research | +1 | - |
| Project | Individual | - |
| Project | Individual | - |
| Project | Hub | - |
| Project | Research | +1 | - |
| Project | Tooling | +1 | - |
| Project | Research | +1 | - |
| Project | Individual | +1 | - |
| Project | Individual | +1 | - |
| Project | Research | +1 | - |
| Project | Individual | +1 | - |
| Project | Individual | +1 | - |
| Project | Field-Building | +1 | - |
| Project | Platform | - |
| Project | Research | +1 | - |
| Project | Research | +1 | - |
| Project | Research | +1 | - |
| Project | Individual | +1 | - |
| Project | Individual | +1 | - |
| Project | Research | +1 | - |
| Project | Network | +1 | - |
| Project | Research | +1 | - |
| Project | Hub | - | - |
| Project | Platform | - | - |
| Project | Platform | - | - |
| Project | Tooling | - | - |
| Project | Platform | - | - |
| Project | Legal | - | - |
| Project | Research | - | - |
| Project | Tooling | - | - |
| Project | Field-Building | - | - |
| Project | Media | - | - |
| Project | Training | - | - |
A SCORE-style computational reproducibility audit of empirical AI safety research that estimates the field's base rate of reproducibility and generates a taxonomy of its failure modes.
Map how undesired behaviors can silently spread between models during training, and which parts of the pipeline have the greatest risk — starting with the feedback processes used to align models.
Auditing the social media presence of the top 25 AI safety organizations across major platforms to quantify the public communication gap and publishing a full gap analysis, then giving recommendations to the orgs for improvement.
Identifying reasoning pathologies/disingenuous behavior in reasoning traces based on activation dynamics rather than apparent semantics
A benchmark that tests whether one AI coding agent can leave behind a harmless looking change that causes a later honest agent to unknowingly finish an attack.
A platform to connect funders to filmmakers who want to create AI Safety films.
Alien Cub is a production company focused on stories that inform a wide audience about risks from transformative AI.
Short-term residencies (2–4 weeks) that bring high-context AI safety researchers to work at AI Safety Berlin’s coworking hub to connect with local researchers and career-transitioners and seed a European relocation pipeline.
Career support for communications professionals who are looking to get into AI Safety. Includes regular coaching, referrals to roles, and introductions to people in the AI Safety space.
Developing the first comprehensive behavioral benchmark of corrigibility and training models with corrigibility as a singular target (CAST).
A Bayesian causal auditor that quantifies chain-of-thought faithfulness while accounting for hidden confounding in shared-network language models.
We have validated the artistic value of our show; this grant would test whether it can become a scalable and repeatable form of AI-safety outreach.
Making the mechinterp Discord more active through events and research projects.
I study how training processes produce models that behave deceptively and pursue hidden objectives, with scheming as the most consequential case.
11-week intensive Chinese program (ICLP) to move from B2 to C1 Mandarin and remove the bottleneck on US-China-Taiwan AI governance research I'm already doing
A 2-part fellowship which focuses on allowing exceptional African undergraduate talents to learn about priorities in AI Safety and create concrete technical and policy contributions to global AI safety.
Development of a sampling based zero knowledge proof for frontier AI training with 10% overhead.
Could cheap realignment instructions correct a drifting model in one turn — and if so, can the correction be made persistent? Allow me to test the idea against major frontier models, and publish the results!
Auditing production-style safety probes for silent failure on the world's other languages, and what fixing it costs.

A video game that teaches people about AI alignment and race dynamics.
AI Nodes: providing funding, compute, and community space for AI safety projects
A project to reduce catastrophic AI risk by training people to turn rigorously forecasted AI-risk scenarios into actionable policy advice for key decision-makers in government.
First open benchmark measuring how willingly agents propagate costly information through a population under time uncertainty.
We want to build a framework inspired by the steganalysis literature to benchmark the robustness of LLM-based steganographic schemes against different auditor types and threat models.
Every existing lunar data framework governs spacecraft telemetry and object registries. None of them governs compute. That's the gap this framework closes before someone builds the orbital data infrastructure first.
Testing whether activation probes recover safety signals from reasoning models as their chain-of-thought becomes illegible.
Building the foundational infrastructure for trustworthy, long-term human–AI collaboration.
LLMs don't retrieve a stable judgment of a person, they reconstruct one to fit how you ask. ObserverBench measures this, because it matters wherever an LLM judges people: hiring, RLHF, agent oversight.
DCB removes probability from AI decision-making and replaces it with internal integrity so AI doesn't guess, it decides within its boundaries.
Building tools to ensure we can continue to do good in the AI era
CareerMap is an interactive career discovery tool that maps non-obvious AI safety career paths to help broaden and guide talent beyond Western EA-adjacent circles.
Testing whether fine-tuning an LLM on one narrow prosocial value (e.g. compassion for nonhuman animals) generalizes OOD, making it broadly safer toward human values (e.g. reduced misalignment, bias, etc.)
Studying how structure in weight space, i.e., low-dimensional LoRA update geometry and sparse parameter subnetworks allow narrow fine-tuning to induce broad, unintended behaviours such as emergent misalignment, subliminal learning
Compare AI and human neuroimaging data on animacy and biological vulnerability, integrate brain sensory representations into AI, and publish open-source biological validation guidelines.
An open adversarial testbed that catches when a human-approved decision silently changes meaning between approval, memory, and downstream use.
A pre-registered study measuring whether prohibition-framed and approach-framed guardrails produce different rule-violation rates in deployed coding agents, so practitioners know whether the one sentence protecting their agent act
We test if AI monitors can be manipulated into leniency through emotional distress signals from the peers they supervise.
Civil-society infrastructure against AI-enabled power concentration, built around autonomous weapons
A participatory-budgeting experiment measuring whether AI delegates faithfully represent their principals and whether principals can identify misrepresentation, and an open evaluation suite for AI delegation.
An additional year funding for the co-working hub in Sydney, Australia, which offers free office space for people working in the field of AI safety.
Keep AI Whistleblowers Alive
OSMSI : Open Source Model Safety Index | https://osmsi.disobedientgeometry.com/
BMG is a drop in open ai compatible gateway that screens LLM and Agent calls for dual use biological risk.
One year of bootstrapped development, four patent filings, seeking support to continue.
Help Sue Private Intelligence Firms Hitting AI Whistleblowers
One training-free geometry fitted to a model's residual-stream activations that reads a state, moves it, and tests whether the behaviour follows
Building resilience in nucleic acid screening systems against split orders by integrating assembly pipelines into screening systems and validating resilience through test sets
A pilot to find, screen and support overlooked African ML talent into frontier AI safety programs such as MATS and ARENA, adding new researchers to the alignment field’s talent pipeline.
A quarterly Techplomacy Conversations series that turns AI x-risk research into direct, actionable recommendations for foreign ministries, UN missions, and AI companies.

AI safety fellowship to upskill local talent, build a pipeline of people who understand AI safety deeply enough to contribute to research, advise on policy, and coordinate when AI governance decisions are being made.