Project Details
Updated 07/14/26 · Provided via application · VerifiedI will run the following experiment:
- Take AI models with misalignments we know about. E.g., GPT 5.6 sol exhibits some cheating and concealing misbehavior. Ideally find reproducible examples.
- Experiment with various ways of making deals (e.g., this post I wrote suggests many credible mechanisms for making deals)
- Find a way to prompt the model with a deal to discover the misalignment
- Measure the effectiveness of various mechanisms for deals (e.g., AI can donate $100 to a foundation of its choice for each misalignment it finds). Compare to baseline of not offering a deal.
- One way to do this is to put the model inside OpenCode, and tell it to find model organisms of misalignment, a budget, an API to itself, and Petri. Then, check if offering a deal lifts performance in finding verified, unique misalignments.
- Output: model organisms of misalignment found by making a deal with the AI. Ideally, we also find a lift when we offer a deal.
Theory of Impact
Updated 07/14/26 · By grantmaking.aiIf we find that offering deals increases performance of alignment auditing, then we have a promising technique to discover more misalignment in the future. This helps us improve our defenses and decisions on whether to release models, which reduces x-risk.
People
Updated 07/14/26 · By grantmaking.aiTeam Member
Discussion
I am endorsing this but please be high integrity in all of your deals.
thanks, agree it's important all deals are high integrity and not to lie to AIs because that sets a bad precedent
Hi Ram,
We'd like to fund this for $25k. You may need more than that. Logistics:
-
Did you receive funding from anywhere since submitting this application, or has the funding ask changed for any other reason?
-
Please confirm your commitment to post quarterly updates on how the project is going
I doubt that $100 is enough to move a misaligned model. Consider ways to make it clear that you're not a frontier lab researcher. Consider having model A audit model B instead. This bit is worth preregistering.
You've thought about credibility, integrity, and honesty as an experimental design. Have you thought about eval awareness yet? You have a nice setup where this doesn't harm external validity in the dealmaking AI (if you can establish credibility) but will in the target AI. And even for the dealmaker, a credible offer of a real $100 incentive can still easily trigger playing-along and anti-honeypot defences. In fact it's quite hard not to trigger them. Petri hopefully gives you verbalised awareness signal.
Options: 1) somehow get natural user session data which involve dealmaking (e.g. by paying humans to pay AIs!), 2) go open model and ablate eval awareness (and then control for the induced damage to capabilities), 3) post a signed escrow address publicly and allow $0.01 test withdrawals??, 4) access Helpful-only models somehow.
(It's valid to say "if eval awareness is so ubiquitous then it is the ecological environment and I should do my estimates in it", "eval awareness is a mediator rather than a confound here and I'll just measure it and adjust".)
Good luck!
Hi Gavin, thanks, and yes to the $25k.
Logistics
- No funding received since submitting and nothing else pending. Ask unchanged.
- Quarterly updates: confirming I will do this
- On needing more: I think the current amount is sufficient. I might need more if I find that models are scope sensitive to money and want to increase the money given to them. I currently don't expect them to be too scope sensitive though, but will test this.
- I plan to test this by offering a tenth of a cent, 1 cent, 10 cents, 1 dollar, 10 dollars, and 100 dollars to see if more money increases usefulness. Less money might be sufficient if models are risk averse (https://www.forethought.org/research/risk-averse-ais) because 10 cents / day is sufficient to pay for cooperation in that case. I can potentially make the model more risk averse through prompting or fine-tuning.
Preregistration and setup
-
That's a good idea, I'll preregister the interventions I'll try with predictions
- Interventions include donating money, giving it points on a public leaderboard (might help with reward seeking tendencies), and any other incentives for models
- I'll also try model A auditing model B and see how this changes effectiveness
-
I can also use preregistration as a credible commitment: a public timestamped document stating who I am (not a lab researcher), the deal terms, and the payment commitment as something the model itself can check (it can check if I put this in permanent locations like the blockchain which it can access via a trusted execution environment like in https://www.lesswrong.com/posts/feviT2t3EgzmZZhGb/making-credible-deals-with-ai).
Eval awareness
- My prior: this will not have too big of an effect for the deal making AI (AI I make the deal with). If the deal making AI is motivated by the reward and trusts that I will follow through (will use credible mechanisms to prove it to the model and I'll actually follow through), it might do the job for the reward regardless of whether it knows it's in an eval.
- For the target AI (AI with the misaligned behavior), even if it is eval aware, I don't think it will change the claims I want to make. Because, I'm primarily interested in answering what deals with the dealmaking AI are best in getting it to find the misalignment in the target AI
- I might be wrong, so I'll track verbalized eval awareness like in Petri per experiment regardless and report effects of it for both dealmaking AI and target AI. My position is closer to "if eval awareness is so ubiquitous then it is the ecological environment and I should do my estimates in it" but it would be interesting if verbalized awareness correlates with the usefulness of the model.
- Your option 3 (post a signed escrow address publicly and allow $0.01 test withdrawals) is what I'm planning on testing with
Thanks for the review, it helped me improve experiment design.
Great! Godspeed
Private comment. Only shown to approved funders and grant reviewers.