Using dynamical systems theory approach to identify dishonesty in LLMs
Identifying reasoning pathologies/disingenuous behavior in reasoning traces based on activation dynamics rather than apparent semantics
Identifying reasoning pathologies/disingenuous behavior in reasoning traces based on activation dynamics rather than apparent semantics
Project Details
Updated 07/13/26 · Provided via application · VerifiedProject Background and Plan
-
A significant amount of effort has gone into analyzing and interpreting intermediate token generation, i.e. “Chain of Thought” (CoT) traces (see Guo et al. 2025 for example), which has led to widespread anthropomorphization of the token generation process (Kambhampati et al. 2026). An implicit assumption underpinning this approach is that these traces are faithful representations of the underlying computational process. However, a growing body of work has shown that CoT analysis is brittle and unreliable, exemplified by results showing models can arrive at correct solutions without producing human-interpretable traces and/or illogical reasoning steps (Valmeekam et al. 2026). This is of course problematic from an explainability point of view, and some have labeled this observed behavior as an emergent form of “dishonesty”.
-
As simple CoT monitoring is not sufficient to understand the internal processes of LLMs, we propose an entirely different approach based on monitoring the internal activations. Instead of focusing on step-by-step level interpretability, we will identify dynamical signatures of these traces which enable us to identify reasoning pathologies such as incorrect reasoning and/or deceptive behavior. This has the potential to make LLMs both more efficient and safe, as the identification of reasoning pathologies early in trace generation enables early termination/resampling before the end user is given incorrect and/or sensitive information.
-
To do this, we will harvest activations across prompts from a series of benchmark dishonesty questions on frontier open-source models. Some trajectories will give wrong/undesirable answers, while other generations will give correct/desirable answers. We will analyze the dynamics in activation space of these traces to identify shared dynamical signatures across tasks which are predictive of whether or not a model will "lie", or if its answer is incorrect.
Deliverables:
- The main deliverable will be a full conference paper submitted for publication to the International Conference on Learning Representations (ICLR) or related conference. ICLR is a premier AI conference which both contributors have previously attended and presented at. Presenting will enable rapid dissemination of our work to a broad technical audience. Furthermore, we believe that public education and engagement is critical in shaping the development of safe AI, and so also plan on reaching out to several forums to communicate our results to a broader non-technical audience. We have a track record of engaging with the general public with our previous work (some of my previous work was featured in Popular Mechanics).
Team
- I am currently an independent postdoctoral fellow at Georgia Institute of Technology. I will be collaborating with Kanishk Jain, a postdoctoral fellow at Emory University (both universities are in Atlanta). I completed my PhD at Carnegie Mellon, where I studied how bacteria adapt to fluctuating environments (I was essentially doing mechanistic interpretability on the computations underlying resource allocation decisions in bacteria). See here and here as examples of our past work. Together we have extensive dynamical systems theory development expertise, extensive experience working with high dimensional dynamical data, and extensive experience studying LLMs.
Theory of Impact
Updated 08/04/26 · By grantmaking.aiModel interpretability in general is critical in reducing x-risk, as it allows to better understand and control these systems. Our work in particular has direct ramifications to reducing x-risk as it has the potential to allow us to monitor and identify when an LLM is lying and/or reasoning incorrectly in a robust way.
People
Updated 08/04/26 · Edited by orgTeam Member
Funding Details
- Jul 1, 2026
- -
- 4 months
- -
- -
- -
- -
- -
- seeking first grant
- -
Track Record
High impact papers in both ML/AI and theoretical biophysics:
- https://arxiv.org/abs/2606.15521
- https://arxiv.org/abs/2410.08439 (ICLR spotlight award)
- https://journals.aps.org/prxlife/pdf/10.1103/5zbg-8vll (featured by Popular Mechanics)
A novel idea. Interesting in combining fields and will be impacteful.