The Veil measures how the other side of the exchange, human or artificial, changes where and how a language model computes, toward testing whether those shifts confound the activation-based tools AI safety relies on.
The Veil measures how the other side of the exchange, human or artificial, changes where and how a language model computes, toward testing whether those shifts confound the activation-based tools AI safety relies on.
Project Details
Updated 07/09/26 · Provided via application · VerifiedA follow-up to the study I already did, "Every Contact Leaves a Trace", an interpretability study on whether different communication styles are associated with specific and recognizable patterns in a neural net's computation. There are a few vulnerable points I need to strengthen: more human creators, and blinded raters to objectivise the classifier used in the first phase. Then replication on different models. Beyond this grant: training a lightweight probe that reads the activations in real time, to test whether communication style confounds the activation monitors safety teams rely on, at scale.
Theory of Impact
Updated 07/09/26 · By grantmaking.aiSkilled misuse happens in sustained, contextual exchanges; safety testing happens mostly in short, directive ones. My study showed those two regimes differ inside the model, in the same substrate safety monitors read, and to my knowledge the material those monitors are calibrated on does not record communication stance at all. Whether that difference moves a monitor's error rates is untested. Communication stance should become a controlled variable in safety testing. This grant would fund the checks that decide whether the concern is real.
People
Updated 07/09/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.