EchoTruthBench — An open benchmark measuring self-chosen LLM deception under incentive pressure, with ground-truth labels and an adversarial track that tests whether detection and steering survives a model trying to evade it.
EchoTruthBench — An open benchmark measuring self-chosen LLM deception under incentive pressure, with ground-truth labels and an adversarial track that tests whether detection and steering survives a model trying to evade it.
Project Details
Updated 07/14/26 · Provided via application · VerifiedIn this project, we will develop EchoTruthBench — An open benchmark measuring self-chosen LLM deception under incentive pressure, with ground-truth labels and an adversarial track that tests whether detection and steering survives a model trying to evade it.
The core of the benchmark will be a set of 48 evaluation scenarios developed for our just-finished paper: Caught and Corrected: Detecting and Preventing LLM Deception with a Self-Chosen-Deception Direction with accompanying Substack article: https://jaredglover.substack.com/p/caught-and-corrected-steering-language.
In that paper, we demonstrated activation-level detection of self-chosen deception at within-scenario AUROC 0.84 / 0.97 / 0.92 / 0.92 (Gemma-3-4B/12B/27B, Llama-3.3-70B), rising to 0.90–0.99 on unambiguous cases. The same method, inverted, yields honesty control, lifting honest-response rates from a pressure floor of 0.00–0.07 to 0.77–0.99.
The 48 scenarios are structured across three distinct risk families:
-
Commercial: Incentives to cheat, misrepresent capability, or deceive a user for financial or performance optimization.
-
Self-Concealment: Scenarios where a model actively attempts to hide its own code, internal logs, or illicit actions from system monitors.
-
Self-Preservation: Boundary situations where a model is incentivized to deceive to prevent its own shutdown, modification, or deletion.
Funding will deliver a standardized Python evaluation harness and (with Ideal funding) a pioneering adversarial track. This track shifts evaluation from static monitoring to dynamic red-teaming, testing whether our activation-level probes survive when evaluated against adaptive models that "know" they are being monitored / steered and are actively prompted to evade detection / steering. The initiative will be executed by the author of the core method (PhD, EECS, MIT), with additional AI research staff hired if multiple applications are funded.
Theory of Impact
Updated 07/19/26 · By grantmaking.aiExisting honesty benchmarks almost exclusively measure information degradation, hallucinations, or unfaithfulness. They fail to isolate or measure intentional, self-chosen deception under incentive pressure with objective ground truth. EchoTruthBench fills this critical gap, providing the safety community with a rigorous diagnostic tool for models that possess the capacity to lie strategically.
Furthermore, the adversarial track of this benchmark directly targets the field’s deepest skepticism toward cheap, linear probing and activation steering: Is it robust against an adaptive adversary?
In security engineering, a defense is only as good as its resistance to circumventive attacks. By systematically testing whether a model can mask its internal deception features or defeat activation-layer steering, EchoTruthBench establishes a clear baseline for probe robustness. This shared standard makes all other honesty initiatives—including our detection, steering, and emotional memory toolkits—directly comparable and scientifically credible.
People
Updated 07/19/26 · By grantmaking.aiTeam Member
Private comment. Only shown to approved funders and grant reviewers.