Neutral, reproducible benchmark measuring whether AI memory systems update correctly when facts change — every major system in one open table, October 2026.
Neutral, reproducible benchmark measuring whether AI memory systems update correctly when facts change — every major system in one open table, October 2026.
Project Details
Updated 07/10/26 · Provided via application · VerifiedIndependent researcher building open evaluation infrastructure for AI agents: a neutral memory benchmark, an MCP security disclosure program, and a published agent identity framework. 25 years of systems engineering background.
Theory of Impact
Updated 07/10/26 · By grantmaking.aiThe x-risk case for this is indirect, but it holds a lot of weight. You can't align agents you can't measure, and you can't trust agents running on infrastructure nobody has looked at. The three projects all sit on that second problem. Each one closes a gap that has to close before the harder alignment questions even become tractable in a real deployment.
Start with memory. Evaluating what an agent remembers is a safety problem, not a polish problem. An agent that holds a stale belief with high confidence will act on it with high confidence, and it'll be confidently wrong. Call it State-Drift: the agent keeps believing something that's since been superseded. At the scale of one demo, that's an annoyance. At the scale of a fleet, it's agents making decisions on outdated beliefs about permissions, credentials, endpoints, and who the user even is. Right now there's no standard way to test for this, so in most deployed systems the failure is simply invisible. You can't see it, so you can't fix it. The benchmark's whole job is to make it visible. Once a failure is measurable, you can go after it.
Then there's MCP, which is becoming critical infrastructure whether or not anyone treats it that way. It's the protocol agents use to reach into the real world: filesystems, APIs, credentials, outside services. And self-hosted MCP servers are landing in production right now with almost no security review behind them. ECHOS is a systematic pass over that attack surface, done before the ecosystem hardens around insecure foundations. Coordinated disclosure gives vendors room to patch before anything goes public. It's the ordinary security-research playbook, pointed at a category that's new and growing fast.
Identity is the third piece, and it's an alignment prerequisite in a specific way. If an agent's sense of itself can be knocked over by an adversarial input, so can its alignment properties, because the same push that destabilizes the identity can override the behavior. The individuation paper lays out a structural model for what coherent agent identity would even look like. The paper doesn't promise safe behavior. What it offers is the architecture underneath, the thing that makes stable behavior something you can build toward instead of just hope for.
None of this is frontier alignment research, and I don't pitch it as such. It's the layer below that: the safety infrastructure that has to be in place before the hard problems can be worked on seriously, at the scale these systems are actually being deployed.
People
Updated 07/10/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.