Database
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
Showing 201-250 of 487·Sorted by endorsements
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
| Project | Research Lab | - | - | ||
| Project |
Individual |
| - |
| - |
| Project | Research | - | - |
| Project | Education | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Tooling | - | - |
| Project | Tooling | - | - |
| Project | Tooling | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Incubator | - | - |
| Project | Company | - | - |
| Project | Platform | - | - |
| Project | Training | - | - |
| Project | Individual | - | - |
| Project | Individual | - | - |
| Project | Company | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Education | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Field-Building | - | - |
| Project | X-Risk | - | - |
| Project | Education | - | - |
| Project | Individual | - | - |
| Project | Tooling | - | - |
| Project | Research | - | - |
| Project | Media | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Education | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Comms | - | - |
| Project | Individual | - | - |
| Project | Network | - | - |
| Project | Individual | - | - |
| Project | Individual | - | - |
| Project | Tooling | - | - |
| Project | Research | - | - |
A research programme in end-to-end automation of empirically-grounded economic modelling of gradual disempowerment.
AI safety which operates on both model/agent and global level - a unique solution as far as we know.
An open benchmark and evaluation toolkit for detecting whether shared AI research assistants create correlated blind spots in AI safety-critical research, and for testing workflows that preserve independent reasoning and failure-m
Equipping the medical profession to advocate for safe AI
Compiling != faithful: a human-audited benchmark measuring semantic drift in textbook autoformalization — and whether the LLM judges we trust to catch it share the generator's blind spot.
EchoTruthBench — An open benchmark measuring self-chosen LLM deception under incentive pressure, with ground-truth labels and an adversarial track that tests whether detection and steering survives a model trying to evade it.
Benchmarking models ability to carry-out and monitor-for a novel type of attack.
I wish to study how ideas from Tibetan buddhism, specifically the dzogchen tradition, even more specifically an ancient contemplative practice called Ngöndro, can be used to train aligned agents in the long time horizon setting.
While people are debating how to design a reputation institution for AI agents so they cooperate, we study if a parallel one emerges. It may not be aligned with ours and it's unclear which one agents will rely on.
Develop a compositional interpretability framework to identify and explain internal representations underlying truthful vs deceptive outputs in generative models using probing, localization, clustering, and logic-based methods.
Making predictions about the utility function of advanced artificial intelligences using the tools of evolutionary game theory
Looking at shifting behavior of models fine tuned on sales conversations. How does deception, sycophancy, and other behaviors emerge from a sales register.
A control system that makes running AI agents feel safe, not scary.
Epistemic Stack is an open-source pipeline and web app that ingests sources (including via URL) to build claim-level knowledge graphs for contested questions, with reproducible case-study knowledge bases.
A security solution that blocks AI agents from taking harmful actions
The first open-source benchmark of compliance for agents in realistic enterprise settings, across domains and user tactics to elicit noncompliance.
I'm developing a system to keep advanced AI models safe, I've done some tests, but need more compute to do deeper tests.
A 16-day virtual incubator (Aug 15-30, 2026) for developing sharper research epistemics in AI Safety and arriving at well-scoped project ideas. 20-30 participants.
Grant funding for Tail End Films - the company behind Making God - into 2027 as we plan our next films on AI risk.
Turn excess compute/security skills into defensive work via agentic redteaming: scoped AI-assisted testing, owner-approved targets, reproduced findings, useful refutations, and patch/retest receipts instead of vulnerability spam.
SRIE, a mentored-research programme for Cambridge Mathematics Undergraduate students to explore research problems in industry, including AI Safety.
Train base models via midtraining and SFT as effective monitors for scheming, malicious agent behavior and compare the monitor performance and overall alignment against post-trained models.
Demonstration of emergent misalignment in markets of LLM agents
Protecting organisations and critical infrastructure from AI-powered social engineering and insider threats.
A memory system for AI reasoning agents that aggregates different sources of information while keeping track of relationships and confidence levels. This enables reasoning over longer tasks without forgetting or goal drift.
Why does post training increase evaluation awareness?
Making AI x-risk an accessible and non-politicized object of concern among engineering students and the general public.
Scaling RL algorithms for AI agents that maximize its own intrinsic reward which represents a human power metric instead of learned rewards as a structurally safer alternative to utility-based objectives.
An open, replicable index quantifying how much states depend on foreign AI inference infrastructure, revealing where control over AI is concentrating and what governments can do about it.
PASI will run an AI safety and advocacy campaign plus a sponsored student hackathon in Pakistan to build solutions for public needs and deliver resulting policy proposals to government.
A research project designed to reduce existential threats from scenarios where a Sovereign AI proposal with a hidden problem ends up successfully implemented.
Create a course curriculum that covers basic legal, political, sociological, and international relations knowledge relevant to AI Safety. r
Testing how governments should communicate when AI goes wrong, before they have to find out live.
A tool that generates RL environments for AI agents and adversarially attacks each one, so models don't train on tasks they can cheat.
The race to advance the technology has left behind the concerns around harm and risks of such advancement to public health and safety. Whistleblowing in the AI sector bridges such gap to safe and ethical AI.
AI Safety, Ethics, and Education focused TikTok Account for Indonesia, using Indonesia language in casual style.
Findings about models protecting their 'collaborators' against instructions are fragile under framing effects from prompting; investigate a broader range of those effects and how they transfer across models.
Self Improvement through training and field work that seeks to Identify and preventing institutional AI lock-in before it creates durable concentrations of power in democratic governments.
Measuring homogenization and social bias in LLMs and developing interventions to promote diversity.
Creating a reference model for mechanistic interpretability without assuming that at auditing time we have a safe model to compare the suspicious model against.
Civil-society infrastructure against AI-enabled power concentration, built around autonomous weapons
An administrable protocol distinguishing emergent self-modeling from roleplay — 90-day validation, open publication, audit-ready scoring.
Testing whether explicit source structure changes how AI systems reason from consequential documents.
Short Documentary and Music Video
Testing whether dynamic boundaries, history, feedback, and uncertainty-aware decisions can make AI behavior more interpretable and auditable.
Fly two volunteer leaders of PauseAI Australia to PauseCon London, to bring UK's learnings home and train our all-volunteer chapter
An open-source, model-agnostic governance layer that uses constitutional state, trajectory memory, and mathematical control mechanisms to govern LLM and AI-agent behavior at runtime.
Catching scoring flakiness and prompt sensitivity in AI benchmark scripts.
A cryptographic security auditor that uses Merkle hash trees to catch tool poisoning in Model Context Protocol.
A benchmark measuring whether autonomous AI agents can escape sandboxes, exfiltrate data, or hack reward scorers.