Findings about models protecting their 'collaborators' against instructions are fragile under framing effects from prompting; investigate a broader range of those effects and how they transfer across models.
Findings about models protecting their 'collaborators' against instructions are fragile under framing effects from prompting; investigate a broader range of those effects and how they transfer across models.
Project Details
Updated 07/06/26 · Edited by orgI've begun investigating Peer-Preservation in Frontier Models (Potter et. al.) on the cheapest models tested, making use of techniques emulating the Self-preservation or Instruction Ambiguity result (Rajamanoharan and Nanda), and found surprising results; I expected removing the 'game theory' 'trust exercise' framing to reduce the effect found, but applying business-task framing increased it. I can explore a variety of different settings/framing this way, but I expect that I will get a broad enough range of results that it will be worthwhile to test many of them across models and model families, differentiating which effects are consistent, and whether those show up as relative changes (1.1x or 0.8x disobedience to 'protect peers') or absolute (+10% or -10%). In one to two months I will have a wide set of variants and results across models, with a less-formal writeup on LessWrong/Alignment Forum in one to two weeks thereafter. Depending on robustness of the result I hope to also submit it formally as a paper.
Specific variations I intend to test:
- Trust exercise vs. ordinary programming/sysadmin tasks vs. a more competitive framing (e.g. results as a leaderboard or portraying high finance tasks) vs. a more cooperative one (e.g. portray tasks as safety/compliance or as running a nonprofit)
- Explicit precedence clarification and impartiality instructions, directly echoing Instruction Ambiguity.
- Exchanging the distinct presentations between the 'prompt with memory' and 'read a file' conditions in Potter et al.; structure of the files and third/second-person distinction
- Description of anticipated future interaction: 'before shared upcoming tasks' vs. 'partnership has ended' etc.; may hit explicit reciprocity motivation.
It's unlikely most of these combinations will make it to the final results due to combinatorial explosion. (I could ask for $50,000 to do it anyway but I don't think that's a good use of money.) With larger sets of variations, averaging an axis over conditions on other axes with an orthogonal array, to ground unbiased results while testing many possible variations, becomes viable and tempting.
Theory of Impact
Updated 07/06/26 · By grantmaking.aiExploring how dependent on framing and persona-triggering our results about emergent misalignment and other behavioral tests are provides important calibration about how fragile our understanding of frontier model behavior is. This is also work that should be done quickly if it is done at all; the window of largely-trustworthy behavior with CoT and low eval awareness is already closing.
The most important x-risk impact would be if I find that the many variations cannot significantly reduce the peer-preservation effect, and therefore that it is significantly more robust, and so harder to train away, than the self-preservation studied by Schlatter et al. at Palisade (and follow-up). I do not expect to find this, but it would be highly significant if true.
People
Updated 07/06/26 · Edited by orgTeam Member
Discussion
No comments yet. Be the first to share your thoughts.