Database
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
Showing 51-100 of 487·Sorted by endorsements
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
| Project | Research |
| - |
| Project | Platform | - |
| Project | Platform | +3 | - |
| Project | Think Tank | +3 | - |
| Project | Individual | - |
| Project | Comms | - |
| Project | Individual | - |
| Project | Network | +2 | - |
| Project | Training | - |
| Project | Individual | - |
| Project | Research | - |
| Project | Individual | - |
| Project | Network | - |
| Project | Research | - |
| Project | Research | - |
| Project | Think Tank | - |
| Project | Training | +2 | - |
| Project | Individual | - |
| Project | Training | - |
| Project | Research | - |
| Project | Research | - |
| Project | Individual | - |
| Project | Individual | - |
| Project | Research | - |
| Project | Research | - |
| Project | Individual | +2 | - |
| Project | Platform | - |
| Project | Individual | - |
| Project | Network | +2 | - |
| Project | Research | - |
| Project | Research | - |
| Project | Research | - |
| Project | Individual | - |
| Project | Media | +2 | - |
| Project | Research | - |
| Project | Research | - |
| Project | Platform | +2 | - |
| Project | Individual | +2 | - |
| Project | Think Tank | - |
| Project | Conference | +1 | - |
| Project | Research Lab | +1 | - |
| Project | Individual | +1 | - |
| Project | Research | - |
| Project | Media | +1 | - |
| Project | Network | +1 | - |
| Project | Advocacy | - |
| Project | Individual | - |
| Project | Training | - |
| Project | Conference | - |
| Project | Individual | - |
We will test whether circuits in protein language models can detect function-preserving redesigns of known toxins that evade homology-based DNA-synthesis screening.
SafeBio-Registry: An Open-Source Verification, SpecDef Weight Locking and Unlearning Platform for Genomic and Protein Language Models
Building an independent runtime governance layer that enables organizations to deploy autonomous AI with authorization, human oversight, auditability and cross-model governance.
Bringing together legal expertise and civil society input to encourage Council of Europe action on AI x-risks and build legal knowledge at the intersection of AI x-risks and the ECHR.
We aim to develop a framework for evaluating whether reward model preferences remain aligned over long-horizon tasks, along with training method that improves long-horizon alignment performance.
Our mission is to inform and organize the public to confront societal-scale risks of AI, and put an end to the reckless race to develop superintelligent AI.
Measuring whether open weight models detect that they're being evaluated, whether they change behavior when they do, and whether that gap grows with capability using causal, white-box evidence.
Liaise helps close the generalist talent gap in AI safety by helping motivated non‑technical people enter with fluency, alignment, and a trusted network.
A junior research fellowship for recent African graduates that combines ILINA’s spring seminar with a mentored research phase on AI and global catastrophic risks, offering governance and technical tracks.
GenomeGuard is an open-source defense and threat-discovery framework for detecting data poisoning, compromised annotations, and supply-chain attacks before they propagate into genomic foundation models.
Agent Island places agents in a rich social setting, similar to reality competitions like Survivor, to study multiagent interactions and the consequences of learning pressure in competitive settings.
Compression-based PAC-Bayes certification for frontier-scale LLM safety monitors, deployment setting shift, and modern post-training.

Humans in Control (HIC) is a nonpartisan grassroots advocacy organization focused on AI safeguards.
Research agenda aimed at developing methods for constructing powerful, easily interpretable world-models.
Creating a reference model for mechanistic interpretability without assuming that at auditing time we have a safe model to compare the suspicious model against.
AI risk assessment currently checks few threat models and doesn't compose them into the aggregate risk that matters. We'll build a tool mapping what frontier system cards cover and omit, plus a paper on what the assessments miss.
COMPASS (Capacity-Oriented Mentorship for Public Administrators on AI Safety & Strategy): AI X-Risk Track trains sitting Global South government officials to manage AI x-risk, so their governments are part of preventing it.
Probe an open-weight model’s activations under biasing/cue conditions to test whether chain-of-thought explanations match internal reasoning, releasing paper, code, and datasets.
A scalable fellowship training researchers to develop interventions for achieving high-value long-term futures
LLM agents collaborate to discover and formally verify theorems about the internal computations of transformers, beginning with a simple pilot question: how many attention heads are needed to represent a Boolean function?
A formal, testable account of LLM persona selection as Bayesian inference, validated against model internals, so labs can monitor and steer model personas during post-training and deployment
Mechanistically analyze how activation verbalizers use target-model activation concepts (e.g., cyclic day-of-week representations) via PCA/DAS/patching, explain cross-family failures, and improve verbalizers.
Browser based game in the style of Plague Inc where players act as a rogue AI attempting to escape human control. Intended to give lay audiences a grounded understanding of how ASI x-risk could play out.
TLDR: A representative survey with Yougov of the American public on questions about AI futures, including space governance, successionism, values.
Build an open-source platform of model organisms and agentic sandboxes to iteratively test mitigations for emergent misalignment via white-box probes, trigger tests, and evaluation-awareness checks.
Naturalizing theoretical alignment by initiating the development of a scientific theory that is capable of making falsifiable empirical claims about agents in general, including humans, AGI, and ASI, and thereby about alignment.

This grant would help us maintain and scale Mapping AI, an open-source stakeholder map of the people and organizations with the potential to shape U.S. AI policy.
How do the options agents present to human researchers alter which research paths are explored, and do steered researchers notice or feel less in control?
A dedicated role at Pause IA to brief French policymakers on existential and catastrophic risks from advanced AI, and to train our volunteer network to do the same with their own representatives.
I've previously developed a classifier that distinguishes training and inference based on Nvidia software telemetry. This project will achieve that using physical sensors, making the system more secure.
Implementing different types of unlearning methods for genomic and protein language models to remove sensitive biological information (e.g. pathogen virulence) while preserving predictive performance and scientific utility.
Identify what AI early risks or failures could be reported despite strategic rivalries towards a « safe culture » shaped after aviation
A training methodology, and transcoder sets that allow to leverage heavy-weight interpretability methods, but made more lightweight for test-time analysis.
A webinar series and peer-support community helping people identify and move into high-impact talent gaps careers and projects in AI safety and governance.
We want to investigate how AI agent overeagerness can backfire when exhibited in safety-critical scenarios.
Safeguarding open-weight genomic foundation models through weight lock against adversarial finetuning
Reduce duplicative hiring processes in AI safety organisations by sharing common screening processes for generalist roles.
CIRIS provides a free and open source model-agnostic agentic alignment harness, available on all major app stores.
Paying 1-8 FTE to scale ControlAI's work meeting and briefing lawmakers about extinction risk from AI and building a coalition to prohibit superintelligence across the G7 (excl. US).
A participant-led unconference gathering the global AI safety community, online and at local sites, Nov 20-22, 2026.
Focusing more on the intersection of red teaming and interp, while creating a strong rooted community.
A project about researching radical methods to both advance and prepare countermeasures for what is to come after transformers.
Perform research to rigorously elucidate and quantify generalization versus memorization, and examine evidence of originality in LLMS.
It's La Jetée (1962), the time-traveling film later adapted as 12 Monkeys (1995), but it's about X-risk and inspired by AI 2027 and it's directed by, and starring, myself and @p8stie. Ergo, La P8stée
Launching a French AI safety community through university outreach and local events to connect students, researchers, practitioners, and policymakers.
Research and plan advocacy for creating a UK Minister for Human Autonomy, including foundational research, project management, and drafting policy recommendations to safeguard human autonomy amid AI.
An independent, multi-lineage panel of AI models, tested on whether it can identify welfare concerns in AI evaluations with opinions, dissents, and proposed modifications published in a public registry
A 3 / 6 month builders fellowship where mentor-builder pods ship practical AI safety tools, not papers.
Hosting a full day conference based in Sydney, Australia where young, aspiring students in senior high school and university interested in AI Safety can connect and share ideas
An evaluation suite to identify a model’s legal values (e.g., anti-tech-regulation) relative to well-known actors (e.g., Ruth Bader Ginsburg) and an assessment of how language in a model’s constitution impacts the extent to which these values are human-aligned.