Reversing LLM Activations to Discover Prompting Strategies in Verbal Uncertainty Expression
Discovering prompts by inverting ideal LLM activation vectors, linking prompt engineering to LLM internals for better calibration, interpretability, and alignment.
Discovering prompts by inverting ideal LLM activation vectors, linking prompt engineering to LLM internals for better calibration, interpretability, and alignment.
Project Details
Updated 07/17/26 · Edited by orgLarge language models (LLMs) often poorly express their confidence, primarily exhibiting overconfidence regardless of the correctness of their statements [1]. In safety-critical settings, in which they are being increasingly deployed, this overconfidence undermines the reliability and trust we associate with using LLMs. In white-box models whose weights are open sourced, we can address this issue through activation steering where we identify and manually edit the activation space of LLMs to change its behaviour. However, this strategy is unfeasible for our strongest and most capable models are often black-boxes, leaving users to rely on prompting strategies derived by intuition, as seen in [2]. Beyond the need for better black-box interventions, discovering the natural language prompts that maximally activate specific internal features provides interpretability and insight into the cognition of LLMs.
This project aims to address two core questions; how do existing uncertainty eliciting prompts interact with LLM activation spaces (e.g., do they engage the same features or degrade intrinsic confidence?), and can we leverage this understanding to reverse-engineer ideal steering vectors into more effective, coherent prompting strategies?
This research will follow three stages. First, we will map the feature space of uncertainty expression in LLMs, characterizing the dynamics of how uncertainty operates within the activation space across different models and various prompting strategies, such as in [2]. Second, we will then construct and identify ideal steering vectors that enforce the model to achieve better uncertainty expression that surpasses currently existing methods. Our final stage will then be prompt rediscovery, where we will reverse engineer the ideal steering vectors across models into prompts and identify consistent prompting strategies.
The concrete outputs will be insights in understanding uncertainty elicitation of LLMs, reproducible experiments, open-source code, 1-2 publications at top tier ML conferences and journals, and a methodology to discover prompting strategies from activation vectors.
This project will be led by myself, Gordon Tan at the University of Toronto, with mentorship from a faculty member (also affiliated with Vector Institute) and a DPhil (PhD) candidate at the University of Oxford.
References
[1] F. Sun, N. Li, K. Wang, and L. Goette, "Large language models are overconfident and amplify human bias," 2025, arXiv:2505.02151. [Online]. Available: https://arxiv.org/abs/2505.02151
[2] G. K.-M. Liu, G. Yona, A. Caciularu, I. Szpektor, T. G. J. Rudner, and A. Cohan, "MetaFaith: Faithful natural language uncertainty expression in LLMs," in Proc. EMNLP, 2025. [Online]. Available: https://arxiv.org/abs/2505.24858
Theory of Impact
Updated 07/22/26 · By grantmaking.aiThis project reduces x-risk from AI by addressing miscalibrated confidence. Models that express certainty that they do not internally posess can systematically deceive users, and in safety-critical or time-constrained settings, it can lead to catastrophic outcomes. By reverse engineering activation vectors that affect how LLMs can accurately represent what they say with what they "know", we can adjust future prompting strategies to mitigate this risk.
Beyond uncertainty expression, this project also advances interpretability, where the methodology of reverse engineering steering vectors into natural language can generalize to any feature direction. If successful, this project will offer a mapping/bridge between a model's activation space and our natural language, accelerating interpretability and alignment research.
People
Updated 07/15/26 · Edited by orgTeam Member
Funding Details
- Jul 13, 2026
- Apr 13, 2027
- 9 months
- -
- -
- -
- -
- -
- -
- -
Discussion
No comments yet. Be the first to share your thoughts.