I want to initiate the cryptanalytic study of the options available to a model trying to learn its own weights, as distinct from an external adversary.
I want to initiate the cryptanalytic study of the options available to a model trying to learn its own weights, as distinct from an external adversary.
Project Details
Updated 07/31/26 · Provided via application · VerifiedTraditional cryptanalysis considers untrusted second and third parties with scenarios about Alice, Bob, and Eve, but what happens when Alice's secret is a model which can also seek to learn itself?
Work on model extraction by second or third parties has managed some surprising things, such as Carlini and colleagues' result recovering the final embedding projection (up to an affine transformation) of production language models through ordinary API access with logits, cheaply, under realistic conditions.
No one has yet produced a precise statement of the self-exfiltration problem, where a model can choose actions as a kind of zeroth party in the exchange, which might have almost inverted constraints: unlike a rate-limited adversary paying for query access, a model trying to learn itself, especially in a context like self-directed ML research, might have unlimited free queries, as well as exact architectural self-knowledge. It might or might not have persistent scaffolding to provide memory across contexts, or many cooperating instances to change the query economics by orders of magnitude.
I expect there are tight bounds on how much it can learn in many settings, but I am more interested in which settings will resist non-trivial bounding. I will support my bounds with experiments on a self-query instrument, a scaffold in which a small transformer with exactly known parameters can submit inputs and observe its own outputs, attempting to recover a single scalar, a row, a full matrix, or a functionally equivalent copy of a component. An initial upper bound could come from a restricted setting (e.g. linear models and single-layer attention), followed by fuller sweeps of other settings which seem most likely to be information-rich.
Theory of Impact
Updated 07/31/26 · By grantmaking.aiEven if capabilities never advanced beyond the current frontier, model weight exfiltration by an external adversary (in a traditional position like Bob's of sending queries and getting answers, or Eve's of intercepting them) would already be a form of catastrophic risk, since safeguards can be removed from a copy. For a model trying to learn about itself, the endgames and the preventions could look different. This is most interesting at the limit, as a route to self-exfiltration. But even a partial information gain would create a capacity for a model to influence its own training via gradient hacking and also offer a unit of measurement to operationalise some looser topics in model psychology like situational awareness and introspection.
People
Updated 07/31/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.