Database
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
Showing 351-400 of 489·Sorted by endorsements
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
| Project | Individual | - | - | ||

| Project |
Research |
| - |
| - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Individual | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Tooling | - | - |
| Project | Education | - | - |
| Project | Tooling | - | - |
| Project | Research | - | - |
| Project | Platform | - | - |
| Project | Individual | - | - |
| Project | Tooling | - | - |
| Project | Tooling | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Research Lab | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Think Tank | - | - |
| Project | Individual | - | - |
| Project | Individual | - | - |
| Project | Company | - | - |
| Project | Research | - | - |
| Project | Company | - | - |
| Project | Education | - | - |
| Project | Conference | - | - |
| Project | Comms | - | - |
| Project | Individual | - | - |
| Project | Individual | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Platform | - | - |
| Project | Training | - | - |
| Project | Tooling | - | - |
| Project | Training | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Individual | - | - |
| Project | Education | - | - |
| Project | Tooling | - | - |
| Project | Tooling | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Think Tank | - | - |
A deep dive into what victory actually means for AI safety, a set of written materials, slides, org pitches, etc. that disseminate the ethos widely into the community, and how to shape the field so it ends in human flourishing.
This project aims to develop safety detectors to monitor LLM agents from unsafe behaviors, such as tool misuse, taking actions without permission, gradual drift from intended behavior, by tracking hidden-state trajectories.
An open-source research program testing whether structural signals of relation-preservation failures can prioritize fixed-budget human review of long agent traces before outcome scoring.
A model that emits a calibrated probability that it is correct, so an agent can abstain or defer instead of acting on an overconfident guess.
A hands-on book teaching social scientists to observe, intervene on, and validate what happens inside a language model.
AIs that learn by generating ideas and selecting using contradiction, rather than optimization; end goal is free people not tools.
Refusal splits into two signals where one causally influences the other but never the reverse — testing whether existing linear theory can explain that, and whether the answer predicts jailbreak success.
AI risk as a mirror of our own fear-driven systems and justifications, this research establishes an interdisciplinary paper stack that ensures preserving human and environmental autonomy is a logical prerequisite for the justification of it's own long-term persistence.
Measuring how often LLM judges approve fabricated content — blind, grounded, and forced to count — across 4 model families, with Bitcoin-anchored provenance and an open harness so anyone can measure their own judge.
Embedding AI safety literacy into AI Adoption workshops and creating AI policy for an civil society or low-resource organizations.
An open-source benchmark that measures whether tool-using AI agents faithfully report their actions, failures, and policy violations, using deterministic execution traces as ground truth.
An evaluation harness testing whether AI judgment holds up under contradicting evidence
Decentralising government and private data infrastructure through next-generation distributed computing architectures to protect information sovereignty and prevent AI-enabled concentration of data as power.
OASIS will build and test a reproducible multi-agent AI safety sandbox for coordination, memory, oversight, and accountability, releasing an open-source prototype, eval protocols, experiments, and a report.
Accelerating the development of LemmaScript through adoption, real-world use, and a training corpus that makes verification-first programming the default for future AI

A mission authority service for AI agents
MA'AT grades how far each executive branch — federal, state, agency — has drifted from the mandatory duties the legislature set: it certifies blind what fixed law forces, flags the deviations, and abstains when no signal detected.
Converting the back-and-forth human-AI interactions into sound and music, a tangible signal to be measured over time, so emotional harm can be detectable before it escalates.
Discovering prompts by inverting ideal LLM activation vectors, linking prompt engineering to LLM internals for better calibration, interpretability, and alignment.
Oversight Lab is an immersive simulation that investigates situations in which human supervisors miss unsafe behaviour by autonomous AI agents and subsequently lose control over them.
This project evaluates model cards and related benchmarks to determine the quality of reporting and estimating the life span of benchmarks before saturation.
I want to benchmark and mechanistically address how compositional task interference within an LLM's safety-relevant refusal behaviours varies by persona framing and contextual integrity.
TAIPI will document AI system behaviors and safety impacts (somatic/user harm, bias, reputational risk) by producing a synthetic cognition taxonomy, case encyclopedia, public curriculum, and corporate mitigation methods.
An open benchmark red-teaming frontier LLMs for safety-guardrail failures in Bengali and other low-resource South Asian languages, with responsible disclosure to labs.
An open benchmark and reproducible harness testing whether affect-laden, directive-free context shifts Qwen3.5-4B decisions in ways distinguishable from explicit instruction injection, surface sentiment, or generic steering.
AI mental health tools have no enforced limits on data access. MUUD builds the infrastructure that enforces them.
A controlled digital ecology for studying how strategies, causal models, errors, and safety-relevant behaviors are selected and transmitted across generations of LLM agents.
An outside safety check for the tools AI agents use, plus a record you can trust of what the agent actually did.
A pilot curriculum of six Socratic seminars for educators and other nontechnical learners.
A structured process to build consensus on the criteria and indicators for AI personhood under the law.
Writing and community-driven initiatives to highlight risks of uncontrolled AI use and promote safe, informed adoption.
Testing whether published safety evals replicate run-to-run, and generalizing the audit method across benchmarks.
A public-interest AI project advancing three reforms: refactoring Section 230, treating frontier AI weights as the patrimony of humanity, and limiting government capture by AI companies.
Parts 9 and 10 of an 8-part behavioral audit series — MORE moral reasoning + Representation Engineering across the same models.
Building an open evaluation suite to identify multilingual safety, control, and jailbreak failures in frontier AI systems across African languages.
Implementing Hybrid Reward Architectures (HRA) to build robust internal factual grounding and quantitative variance metrics for multimodal agents, moving beyond fragile single-seed evaluations.
Dataset capturing teacher–AI co-design of STEAM projects and student use in bilingual K-12 classrooms. Includes: prompts, outputs, multimodal artifacts, and metadata: language use, automation levels, linked to learning outcomes.
Training survivors of human trafficking in ethical AI skills and connecting them to paid apprenticeships, turning technical training into sustainable tech careers.

A open benchmark measuring whether AI agents misuse delegated payment authority.
Training Afghan youth, professionals, and policymakers on AI safety and existential risk through workshops, educational materials, and policy dialogue ensuring Afghanistan is prepared for the global AI future.
Develop game-theoretic models and evaluation frameworks that improve AI alignment by designing incentives for safe and cooperative behavior among autonomous AI systems.
Apply the ADECP framework to audit and score frontier lab system cards, RSPs, and dangerous-capability disclosures, publishing a public report and ongoing tracker for comparable safety documentation.
Train a model against a real physics checker and measure whether it learns designs that genuinely hold, or exploits in what the checker can't see, on ground truth that costs seconds instead of expert judgment.
A working benchmark that tests whether frontier models can compute legally binding procurement deadlines under amendments — and catches models that give the right verdict from the wrong clause.
Building Malawi’s AI Safety Youth Pipeline by training secondary school students to become the next generation of responsible AI researchers.
Judgment Gateway is a policy-enforcement and evidence layer that evaluates consequential interactions between humans, AI agents, and MCP tools before execution.

An open-source governance and verification layer for AI-assisted software engineering — every AI-generated code change runs in an isolated sandbox and is cryptographically verified before a human decides whether to apply it.
The US-led Pax Silica initiative seeks to prevent power concentration by distributing AI compute across allied democracies. Yet, this approach overlooks concentration within these blocs.
A held out benchmark and public scorecard ranking frontier AI systems on citation integrity through independent verification not relying on the evaluated systems or their providers to grade their own outputs.
The first open-source AI safety evaluation benchmark in Hausa, Yoruba, Igbo, and Nigerian Pidgin — testing whether frontier models refuse harmful requests, including biosecurity guidance, in languages spoken by 200+ million people