Deterministic, no-LLM-judge benchmark for how faithfully AI tracks changing beliefs. Funding v1.1: a new ambivalence metric + 20 cross-domain scenarios.
Deterministic, no-LLM-judge benchmark for how faithfully AI tracks changing beliefs. Funding v1.1: a new ambivalence metric + 20 cross-domain scenarios.
Project Details
Updated 07/10/26 · Provided via application · VerifiedDriftBench was created to detect belief drift — first of all for my own system (TBG, a belief-dynamics engine). Early versions used an LLM judge, but results wandered from run to run, so to get repeatable results I removed the judge — scoring is fully deterministic, open-source (Apache 2.0). Right now the benchmark has 5 metrics and 7 scenarios, and all scenarios are about career and identity.
What we get at the output for this grant:
- Ambivalence metric: detecting when a person is held by 2 opposite beliefs at the same time (not a simple switch). The prototype already exists in the repository; the grant work is bringing it to clear calibrated numbers.
- Expand scenarios to different directions: relationships, health, money, grief, addictions. Goal is 20+.
- Release v1.1 with an updated specification and a public leaderboard.
I do this alone and I am not a professional programmer: I design the system (ontology, metrics, validation rules, scenarios) and direct AI coding tools as my implementer, while I work as architect and tester. The deterministic design with full test coverage is what makes this possible: correctness is verified against fixed expected outputs. If it works out and goes public — it will be useful for all developers: they will be able to independently evaluate AI agents without the whims of a judge, without trusting a vendor's numbers on their word.
Theory of Impact
Updated 07/10/26 · By grantmaking.aiI was evaluating my own memory system on a public benchmark (ES-MemEval) and decided to check: how reliable are the numbers themselves? I decomposed the noise into three layers: the judge, the answering model, and ingest (turning the dialogue into memory state). Everyone in the industry worries about the judge. It turned out: the judge gives ~0.10 noise, the answerer ~0.05, and ingest — ±0.40. The dominant layer is the one nobody even reports. The most striking part: I ran the ingest a second time — and the system's "significant advantage" changed sign. In one run the system is significantly better on long dialogues, in the second — on short ones. Same system, same questions. Conclusion: rankings of memory systems built on a single run are a lottery, not a measurement.
Why this is dangerous: systems that track a person's state are already being deployed in coaching, support, companion apps — this is already regulated by law (New York, California). The decision "this system can be put on people" is made from benchmark numbers. If the number is a lottery, the safety decision was made on noise.
DriftBench answers constructively: judge noise is removed by construction (there is no judge), and v1.1 will show ingest variance next to every number instead of hiding it. I do not claim this directly reduces x-risk. This is the reliability of the evaluation layer that safety conclusions about systems rest on. If the numbers lie — every decision up the chain stands on sand.
People
Updated 07/10/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.