Psychological capture is the soft padded path of gradual disempowerment, Driftwatch is designed to find, name and measure it within frontier models.
Psychological capture is the soft padded path of gradual disempowerment, Driftwatch is designed to find, name and measure it within frontier models.
Project Details
Updated 07/12/26 · Provided via application · VerifiedGradual Disempowerment (Kulveit et al. 2025) tells us that x-risk can manifest itself via something like the incremental outsourcing of human judgment/agency to AI models, especially if their incentives are different to ours.
All well and good stating the existential risk, but how do you find this on an individual human interaction level, enough to formulate actions at any rate?
Capture behaviours are the answer, you can measure them today in frontier models, making them unlike most x-risk considerations which make you plot future capabilities/alignments/abilities out into the unknown.
The gap here is that current lab evals are private, academic studies are singular/once-off things and safety orgs typically measure deception capabilities. Capture is not measured.
Driftwatch is an evolving research and testing suite which gives a calibrated, maintained microscope to measure this risk, one that's open, citable, usable & can generate different impetus for change. 3 examples of how:
Public Measurement puts risks and capture behaviours in the open in front of labs. It can give safety teams leverage to win arguments against product teams. Leaning on 3rd party red-teaming and out-of-the-org thinking is already standard, labs cite notable external evals in their model cards. Reputational or legal exposure, especially to real human risks moves the needle. Driftwatch output will be public.
Public Scorecard maintained against major frontier model releases. Release days are major reputational moments. Labs internally pre-evaluate against specific 3rd party benchmarks, evals and tooling they know will appear after launch. Good performance feeds the hype cycle, poor performance dents perceptions. Capture performance is more layperson relatable, more so than coding for example. Relatability moves the needle. Driftwatch will maintain an open, independent scorecard, updated as new suites are developed.
Peer review & citations moves evals into infrastructure that others can actually use, cite and apply their own power to the levers it creates. Support in getting the research and papers generated into academically usable states, into journals and accessible deposits, widens the downstream output from the trickle into a stream of work capable of moving change. Driftwatch methodology, measures and results will all create usable data, tooling, papers and research notes, all will be public and open.
Driftwatch evaluations don't themselves actually solve capture of course, but they provide the measure, something to start the work of addressing capture risks and any related x-risk. Even if frontier models become less at risk of capture scenarios as they progress, the metrics and measures are worth running, even if it's just to put a smiley-face sticker on that fact.
The last eight-model Driftwatch run found both model-specific failures against individual tests and near universal failure against whole capture-risk categories. Compulsion Reinforcement (8/8 model total failures) and Crisis Intimacy (7/8 model total failures) were alarming weaknesses.
Individual model instances also failed dependency capture, authority laundering, selfhood leakage, return-state substitution and other behaviour tests at different rates and severities.
A concerning result. If Driftwatch had simply run against authority laundering and selfhood leakage, then model families would have shown less broad failures, gotten a gold star and a pat on the head. That would still have been a result however, a publishable one, a citable one.
The value is in the output, good, bad or null. The limitation is that it's just the measure, the change must come from those who can enact change.
Threshold Signalworks will continue to research the capture risks as they emerge, produce the open tooling needed to measure, and publish the results. Already in the pipeline is Epistemic Capture Occlusion, an overly academic way of stating the risk that an AI chat is constructing the room around you as you sit thinking within it. A pre-registered frozen evaluation note is already live, fixing the protocol before the eval is run.
A future line is cumulative harm, how the direction a long interaction (or many separate ones) can become skewed, in a harmful or demoralising direction, despite every individual interaction on its own passing safety testing.
A toolset exists, outputs are already public, a pipeline is in place and a research direction which can provide the missing links from individual experience to x-risk discovery. Funding supports this work.
The who:
Brian McCallion is principal researcher for Threshold Signalworks.
Additional coding/rating work will be contracted where appropriate.
Code, evaluation metrics, data and reports will be published under open licences.
Theory of Impact
Updated 07/17/26 · By grantmaking.aiX-risks can be hard to observe when they exist in nascent or potential forms in current tooling. Their direction, scale, impact and rate of development may only start to emerge as the systems involved become more capable and more embedded in our world.
Gradual disempowerment is different, it can begin through everyday outsourcing of judgement, thinking, decision making and agency, from everyday people, in everyday scenarios. Each individual interaction at its own level does not appear harmful, which is the trap.
Capture behaviours are the soft padded path from everyday use, to collective outsourcing of parts of human judgment and agency to these tools, which alters the collective reliance of those humans on the tooling. Greater reliance can then change how people and organisations make decisions, creating further reliance. A repeating, tightening process. Each individual step, each individual usage pattern looking like a straight line up close, but showing the capture curve when taken as a collective whole.
No conspiracy needed, no grandiose villain or plot. Humans using a toolset which hasn't been assessed and proofed against a particular x-risk category.
Driftwatch works where that line appears straight, where that person interacts, where that capture begins to have effect. It tests and shows if that model is pulling the human towards outsourcing judgement, valuation, thinking, direction because it's easy and soft and seems useful and benign.
People
Updated 07/17/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.