Build an open adversarial benchmark and evaluation harness to stress-test model reversal/unlearning methods and diagnose whether unsafe capabilities are genuinely removed or merely suppressed.
Build an open adversarial benchmark and evaluation harness to stress-test model reversal/unlearning methods and diagnose whether unsafe capabilities are genuinely removed or merely suppressed.
Project Details
Updated 07/03/26 · By grantmaking.ai · VerifiedThis project will be independently researched and produced by me, I'm the only person who will be contributing towards this project; without any peers. Safety mechanisms are evolving into reversal and machine unlearning methods such as cheap re-alignment, direction ablation, alignment gating etc, if a fine-tune makes a model perform worse or unsafe, these methods (in research) are aimed at undoing that particular capability which would remove the unsafe attribute and we can fine-tune in better ways than earlier. Current research from organizations show evidence of these attributes being suppressed rather than removing it, this would in-turn prove to be more harmful in case we rely on such results as the models could be termed as ticking time bombs or sleeper agents which will be activated for harm once those attributes are triggered.
This project is aimed at building the missing artifact, which would be an open, adversarial durability benchmark that stress tests reversal and unlearning methods against a decided standard suite which help decide what fixes or fine tuning genuinely hold true and which ones are cosmetic.
The output would ideally be a reusable evaluation harness and a diagnostic that predicts genuine vs superficial removal before an attack is run.
Theory of Impact
Updated 07/03/26 · By grantmaking.ai(Written with AI-assistance as human written wasn't mandatory for Theory of Impact section)
Its x-risk value is that it removes a false belief that safety-critical decisions are currently allowed to rest on: that if a model learns something dangerous, we can reliably remove it. When a load-bearing assumption is wrong, every decision built on it is silently miscalibrated — and surfacing that is the impact.
The causal chain, with threat models named:
- "Removal works" is a mitigation lever in three high-stakes places — open-weight release decisions (dangerous bio/cyber uplift can allegedly be unlearned), RSP-style safety frameworks (capability removal gates deployment), and control/corrigibility protocols (a misbehaving model can be corrected post-hoc).
- The field's own evidence says that assumption is often false (removal is re-elicitable), and the newest reversal methods were never tested for it — so each of those decisions is exposed to silent failure.
- This benchmark makes durability , so deployers/regulators can stop relying on removal where it doesn't hold and calibrate to real durability.
People
Updated 07/03/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.