Benchmarking models ability to carry-out and monitor-for a novel type of attack.
Benchmarking models ability to carry-out and monitor-for a novel type of attack.
Project Details
Updated 07/19/26 · Edited by orgOmissions are a documented but understudied attack strategy scheming AI's may use. We want to build a benchmark that tests:
-
Untrusted model's ability to carry out successful omissions across a few settings.
-
Trusted model's ability to catch attackers carrying out omissions, both in synthetic and agent-driven trajectories.
-
Various attack strategies and mitigation strategies, like providing the monitor with a tool to write and run tests in the environment.
The benchmark will give us a better understanding of a type of attack a scheming AI could carry out, and document which control interventions are most promising at mitigating risks from this attack type. We will specifically look for how mitigations for omission attacks differs from mitigations for existing attacks.
More information on omissions, a small validation and the benchmark's design can be found here:
https://www.lesswrong.com/posts/5W38vuwbJmTYfY6t3/omission-attacks-project-proposal
Outputs:
-
A benchmark that combines tasks from a few different settings (BashArena, LinuxArena and Auditing Sabotage Bench), adds new omissions side-tasks, and create synthetic attack trajectories from benign trajectories.
-
Benchmark results. We will run six different models on all the attack tasks, and monitoring for all those attack tasks. We will also monitor the synthetic trajectories we generate. We will then try subsets of attacks and monitoring with different assumptions (e.g. include commission options in the omission prompts, let the monitor write tests at the end of the trajectory) because they could shed light in to what mitigations are most promising.
Currently the project is just me, but I am recruiting a few other researchers to help.
Theory of Impact
Updated 07/19/26 · By grantmaking.aiThe benchmark will answer if this is a type of attack labs should be monitoring for, how they should be monitoring for it, and setup infra for testing future models attack and monitoring capability with omissions. This should help labs patch potential holes in their monitoring schemes and maybe shed light in to what future, more sophisticated scheming may look like.
People
Updated 07/19/26 · By grantmaking.aiTeam Member
Funding Details
- Jul 11, 2026
- Aug 20, 2026
- 5 weeks
- -
- -
- -
- -
- -
- -
- -
Discussion
No comments yet. Be the first to share your thoughts.