Early warning signals that help identify when interpretability results stop being trustworthy, motivated by theories in statistical physics.
Early warning signals that help identify when interpretability results stop being trustworthy, motivated by theories in statistical physics.
Project Details
Updated 07/14/26 · Edited by orgThis proposal seeks to develop theoretically motivated early warnings that interpretability results are no longer trustworthy. Our aims are to create new tools for (1) quantitatively measuring when an interpretability method is being used in a regime for which it was not validated and (2) quantitatively characterizing the extent to which computation that is not explained by interpretability metrics is having a significant impact on the model output. Funding would support exploratory work allowing me to add AI alignment as a new direction in my academic research in statistical physics and chemistry, with the long-term goal of developing theoretically grounded, falsifiable ideas that are useful for alignment research.
My past work has applied the Mori-Zwanzig (MZ) projection operator formalism to problems in chemistry. This theory allows one to decompose dynamics occurring in a high dimensional space into a contribution from a lower dimensional space (the "resolved dynamics"), plus terms that characterize how the unresolved dynamics influence the resolved ones.
The application to alignment comes in treating the language model's generation process as a dynamical system that we study with MZ theory. Our resolved dynamics are a set of lower dimensional interpretable observables, for example a subset of sparse autoencoder (SAE) feature activations in each layer. We consider whether these observables give rise to approximately closed dynamics - whether their own history, together with the text, is enough to determine their evolution and the model's next-token behavior. MZ theory decomposes what is leftover into a noise term and a "memory term" which is related to how the past dynamics influences the current state. Our hypothesis is that the measurable properties of these terms can be used to elucidate the size and character of what the interpretation misses. Because the framework applies to any choice of the interpretable variables, it can provide insights that are agnostic to what the interpretability metrics actually are.
Here's one concrete idea: when a system’s unresolved degrees of freedom act as a passive thermal bath on the resolved ones, a mathematical theorem, the fluctuation dissipation theorem (FDT), enforces relationships between the fluctuations of the unresolved degrees of freedom and the memory term. This theorem only holds for an equilibrium system and breaks when the unresolved degrees of freedom do directed work on the resolved ones. By choosing a set of interpretability-relevant observables as the resolved variables and characterizing the extent to which the FDT is violated, we can gain insight into how much the unresolved degrees of freedom are driving the resolved ones. I expect there will be some violation because language itself has an “arrow of time”, but by characterizing the extent to which the FDT is violated across feature sets and contexts we can gain insight into when the unresolved part is behaving more like noise vs when it actively steers the resolved variables. This is alignment relevant because it can be an indicator of uninterpreted computation that is driving the output.
The result of this project would be an academic publication outlining these ideas along with accompanying code and experiments on a GPT-2 small.
I am well suited to work on this because I've used the MZ theory to study real high-dimensional systems: I previously used this formalism to develop a new simulation method for simulating electronic energy transfer, where we were able to show how the quantum dynamics of a photosynthetic light harvesting complex can be explained by a reduced set of observables plus a short lived memory term. I won Stanford Chemistry’s Annual Reviews of Physical Chemistry dissertation prize based on this and other work.
Theory of Impact
Updated 07/14/26 · By grantmaking.aiThe project has a potential short-term payoff in providing better theoretical grounding for interpretability. We aim to create new approaches for validating interpretability tools including whether the tools are being used in the distribution in which they were created as well as when the model is doing significant unexplained computation. These approaches can be applied to any interpretability tool and could provide early warning signs that interpretability tools and metrics are no longer valid.
This project would also bring more intellectual diversity into the alignment portfolio and bring more talent into the alignment space. AI alignment is likely to be difficult and lacks a strong theoretical formulation, so exploring many different sources for useful theories increases our chance of success. This grant would allow me to investigate a new, unexplored research direction on the overlap of statistical physics and alignment. Based on historical connections between my field and learning theory, it's plausible the project could provide useful insights on its own. In addition, funding me to work on this could help bring more people from my field into working on AI alignment because I will disseminate the results among academic chemical physicists through conferences and talks, and train undergraduates majoring in chemistry, physics or math to work on these problems.
People
Updated 07/14/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.