A constantly-updated aggregation of AI safety and ethics evaluations, statistically combining sparse literature results and self-run evals into a global ranking of models.
A constantly-updated aggregation of AI safety and ethics evaluations, statistically combining sparse literature results and self-run evals into a global ranking of models.
Project Details
Updated 07/19/26 · Provided via application · VerifiedThe project splits into two lines of work:
1) Gathering and aggregating existing AI safety/ethics evals.
We will produce a live, constantly-updated website which:
- Collects all existing evals of LLM safety and ethics and categorizes them.
- Combines results into a global score and category scores. Similar to how Arena AI and Artificial Analysis plug sparse head-to-head comparisons into a Bradley-Terry model to create global scores, we plug sparse model results on evals into a statistical model to effectively aggregate results across different evals.
2) Filling in gaps and keeping evals updated.
The result of (1) is that gaps in evals become readily apparent. For example, few models have been subject to evals of ethical treatment of nonhuman subjects, and few recent Chinese models appear in safety evals. Unfortunately, most evals are one-and-done affairs, and are not continually updated.
- For existing models, it will be clear which evals should be run on which models to reduce uncertainties and improve our awareness of their safety performance and ethics, and we will run these evals.
- For new model releases, it will be clear which evals should be run to cost-effectively give us a good understanding of their safety performance and ethics. As soon as possible upon model release, we will run these evals and publish the results online.
- We will design and run new evals in areas where existing evals are too sparse to get a good signal.
Here is a live MVP of the gathering and aggregation: aisafety.aozerov.com
As of writing, this MVP statistically aggregates results from ~400 models and ~50 source evals.
Tangential research questions we will answer using this project's products are:
- How valid are safety and ethics evals anyway? Which ones are more or less noisy? Are eval results predictable using data from other evals?
- How do safety and ethics relate to model capabilities?
- Aside from a global score, what are the primary dimensions across which models differ in their safety and ethics performance?
Team: me, a statistics PhD student at UC Berkeley.
Theory of Impact
Updated 07/19/26 · By grantmaking.aiIndividuals and businesses now have easy ways to choose LLMs based on their capabilities as soon as they are released (see Artificial Analysis and Arena AI). But safety and model values are rarely considered, because (a) they are not measured in a timely manner and (b) there are only sparse results here and there, and no global aggregation. Open models (Deepseek, GLM, etc.) achieve wide adoption due to great capabilities at low costs despite often poor safety performance and questionable ethical values.
The most immediate risk-reduction from this project is from people choosing to use and deploy safer models. But long-term a prominent, independent, and trusted source of AI safety/ethics data will:
- Create market pressure on labs to make AI models safer. We perhaps should not put all of our eggs in the regulation basket.
- Provide timely observability into possible x-risks from new models, particularly for near-frontier-level models released without an accompanying extensive safety evaluation. Evals in our aggregation include concern for human and nonhuman welfare, propensity for deception, and responsible use of power. If a powerful new model that performs poorly on these evals is publicly released, we want to know quickly.
People
Updated 07/19/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.