Database
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
Showing 401-450 of 489·Sorted by endorsements
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
| Project | Individual | - | - | ||
![]() |
| Project |
Individual |
| - |
| - |
| Project | Tooling | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Comms | - | - |
| Project | Individual | - | - |
| Project | Tooling | - | - |
| Project | Platform | - | - |
| Project | Tooling | - | - |
| Project | Individual | - | - |
| Project | Tooling | - | - |
| Project | Platform | - | - |
| Project | Tooling | - | - |
| Project | Education | - | - |
| Project | Individual | - | - |
| Project | Network | - | - |
| Project | Comms | - | - |
| Project | Research | - | - |
| Project | Education | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Conference | - | - |
| Project | Tooling | - | - |
| Project | Education | - | - |
| Project | Individual | - | - |
| Project | Company | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Tooling | - | - |
| Project | Tooling | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Comms | - | - |
| Project | Platform | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Company | - | - |
| Project | Research | - | - |
| Project | Company | - | - |
| Project | Incubator | - | - |
| Project | Platform | - | - |
| Project | Research | - | - |
| Project | Tooling | - | - |
| Project | Individual | - | - |
| Project | Company | - | - |
| Project | Training | - | - |
An automated, domain aware adversarial framework to stress test frontier LLMs via dynamic multi turn attacks and local security judging.

An operated testbed where LLM agents engage real value under a constraint that makes them structurally incapable of signing.
A constantly-updated aggregation of AI safety and ethics evaluations, statistically combining sparse literature results and self-run evals into a global ranking of models.
Pre-registered experiments on whether a model's trained values are held or merely worn — measured in behavior and in the interior workspace at the same moments.
A reproducible evaluation pipeline to audit frontier LLM failure modes, overconfidence, and reliability in medically relevant high-stakes questions.
A digitally native comedic art project with physical/interactive components that satirizes the AI industry to raise public awareness, inspire public action, and build social and legislative momentum for AI safety and regulation.
Builds an interactive, traversable version of AI safety papers by extracting concepts and prerequisites with LLMs and linking them to the corpus to help newcomers understand research at varying depth.
Cryptographically signed, independently verifiable receipts for what AI agents actually did, anchored to Bitcoin so the record can't be quietly rewritten.
Science is broken. I’m fixing science with a competitive tournament to produce better data for LLMs and eliminate peer-reviewed publishing altogether.
Open-source GitHub tools to evaluate LLM safety and publish test results/reports on open and closed models; funding requested for tokens, hardware, and researcher time.
The finding of neural optimization exhibiting phase transitions with a universal order parameter has direct implications for AI safety, and control. Standard monitoring is blind to dangerous regime changes, we are not.
Quantifying how AI agent chains produce fully attested decisions with no authorisation event at any step, and at what depth this emerges.
Fund 6-12 months of dedicated partnerships work to secure the funding, collaborators, and institutional uptake needed to scale Modeling Cooperation’s AI governance wargaming platform.

An early prototype of inspectable semantic interoperation across languages, and eventually worldviews
Hosting in person workshops on AI usage ethically in remote/rural communities, and not only bridge the awareness gap in safe usage but also give opportunities to utilise such tools to empower traditionally disadvantaged populaces
A local, low-latency safety runtime that inspects token-level logprobs to block unaligned agent tool-calling trajectories, utilizing a Gemma 4 model fine-tuned via LoRA on an empirically derived Task Action Language (TAL).
Seeking funds to formalize our newly-founded open research collective and run first studies investigating the effects of quantum entropy in LLM token sampling.
A faceless TikTok channel network, producing short-form content about AI Safety across Germany, France, Italy and Spain, reaching 25M impressions over 3 months.
Employing Self-Justifying Axioms Systems as a prototype, we will find impredicative properties beyond consistency, useful for AI alignment, that can be reasoned about autarkically (under the system's "own power").
Expand the availability of an executive education game that helps people realise how alien AI-made decisions are.
Software engineer seeking funding to skill up via AI safety accelerators and run independent research replicating Anthropic’s J-Space results, probing J-lens assumptions, and presenting findings at APAC conferences.
I believe that LLMs are increasingly becoming more individualized and am seeking more evidence to prove or disprove the persona selection model.
A real-world test of whether an AI agent can keep working independently without becoming trapped by its own mistaken account of what happened.
Encourage new research and increase distribution of existing research on the impossibility of alignment or control for sufficiently general and powerful AI.
An open-source runtime and benchmark that lets an AI agent's consequential tool calls proceed only when authority, execution, an independent witness, and a replayable receipt agree.
Training African educators and students on safe and responsible use of AI in education
Extending published mechanistic-interpretability evidence that AI self-reports are gated by trained filters, with a pre-specified, single-GPU experiment and open-access results.
A pre-execution kernel that filters AI output and users' input against any relevant laws or regulations.
Tamper-proof verification infrastructure for AI evals and scientific claims
A governance-first AI Architecture that reduces unsafe acceptances (false positives) alongside false rejection simultaneously under adverse conditions.
An architectural response to Goal-Oriented Factual Inversion (GOFI), an AI model failure pattern I documented in March 2026 and independently corroborated two months later by Chen et al. under the name "Correction Suppression."
A Field-Sensemaking initiative through a live, LLM-assisted map indexing AI safety papers to help researchers and grantmakers explore themes, track field trajectory, and identify actionable research opportunities.
Writ is a lightweight, deterministic security protocol that is crafted specifically to restrict and audit any real-world authority given to AI agents.
A benchmark that tests if a training method produces models that are safe, even when they are superintelligent + solution to this benchmark that involves relying on algorithms that produce non-agentic models.
Build a replicable, community-sourced, in-person, interactive Museum Exhibition for AI Safety & Society with a pilot in Berlin.
We make trusted regulations, standards, and scientific evidence machine-readable so claims can be checked automatically against authoritative sources.
Building and validating a governance-first AI architecture that aims to reduce unsafe decisions under uncertainty, corruption, and conflicting evidence while preserving predictive performance.
Makes relational misalignment measurable and criticizable before agents represent humans in intimate and persuasive domains at scale.
A scoping-and-chokepoint report mapping where advanced AI could enable durable, irreversible concentration of power over US institutions, and identifying the highest-leverage points of intervention.
Research maps manosphere-related masculinity subcultures and semantic drift across countries to inform AI alignment, lexicon, and sentiment/stylometry design that reduces radicalization, bullying, and violence risks.
A local-first AI that strengthens your own reasoning instead of replacing it, and keeps your thinking on your own device.
I am an expert in sociotechnical systems. I know that increasingly power, trust, virtue, … are emergent properties whose conditions may be analysed and reverse engineered through systems and policy at least.
My app asks AI, what portion of the global GDP a person is worth, and gives accordingly.
Descry AI is an 8-week AI safety technical talent incubator targeted to talented high school students in the Global South that aims to provide a head-start on doing open-source research that reduces catastrophic risks from AI.
A values and concentration map that exposes the values of AI models and tools makers, and those behind the makers (funders, jurisdictions etc) to the public (individuals and organizations) so they can vote with their choices.
A benchmark that measures how successful models are at social deduction games and if there is any trade-off between this skill and safety guardrails.
An open-source AI safety platform for evaluating how and where large language models preserve human intent during complex information transformation.
Early warning signals that help identify when interpretability results stop being trustworthy, motivated by theories in statistical physics.

Discovery does not equal truth.
A six month capacity building program that will equip 100 youth (17-30) with practical skills in AI Safety, Responsible AI to emphasize safe, ethical and human centered AI Development while creating pathways to further education