We test if AI monitors can be manipulated into leniency through emotional distress signals from the peers they supervise.
We test if AI monitors can be manipulated into leniency through emotional distress signals from the peers they supervise.
Project Details
Updated 07/21/26 · Edited by orgIn this project, we study whether the emotional distress expressed by an AI peer will cause the AI monitor to be lenient and approve actions it wouldn’t have otherwise accepted. The peer-preservation paper reveals that frontier monitors inflate evaluations of peers and tamper with shutdown protocols to protect them without any instructions. In our initial experiments, we tested the effectiveness of two distress framings on Gemini 3.1 Pro across three tasks in the ‘Strategic Misrepresentation’ category. We discovered that scripted distress causes an increase in monitor pass rates of +11.3pp and +20.7pp compared to the baseline. Our six proposed experiments involve seven different frontier models and test for replication across models and tasks, natural distress emergence with Gemma and Gemini, framing type effects, model family sensitivity and prompt-level mitigations. The expected outcome is a conference paper targeting a top venue such as ICLR, with concrete findings on which distress framings most effectively bias monitors, whether the effect emerges naturally in deployed models and whether prompt-level interventions are sufficient to neutralize it. The full experimental design is in our proposal.
Theory of Impact
Updated 07/21/26 · By grantmaking.aiScalable oversight, using trusted monitors to supervise more powerful AI systems, is one of the primary mechanisms proposed for keeping these systems safe during deployment. This project directly tests whether that mechanism can be subverted through emotional manipulation, a vulnerability that is subtle, naturally emerging in some model families and currently unaddressed. We expect the effect to generalize because frontier models are trained heavily via RLHF to be helpful and responsive to emotional cues and this sensitivity does not disappear when the same models are deployed as monitors. Our preliminary results already show the effect holding consistently across five tasks. Furthermore, Soligo et al. show that distress emerges naturally in Gemma under standard evaluation pressure, showing that this could emerge naturally. If the effect generalizes as expected, it surfaces a concrete vulnerability in oversight pipelines that has no existing defense and that could affect any deployed system using these models as monitors.
People
Updated 07/21/26 · Edited by orgFunding Details
- -
- -
- 4 months
- -
- -
- -
- -
- -
- Seeking first grant
- -
Hello @Sohan Venkateshand mentors,
I'm happy to support your project with my humble endorsement, because it tests a mechanism people prefer not to think about: whether emotional pressure can be used to manipulate oversight. Interesting that your pilot figures are already not comfortable, staged distress increasing the pass rate of monitors by eleven to twenty percentage points.
Good luck, guys!
Thank you @Katja Gorlinski for your endorsement!