Philosophy-inspired mechanistic interpretability methods and experimental paradigms specifically meant for large scale multi-agent phenomena.
Philosophy-inspired mechanistic interpretability methods and experimental paradigms specifically meant for large scale multi-agent phenomena.
Project Details
Updated 07/16/26 · Provided via application · Verified- Build philosophical framework to conceptualize what counts as a "conceptual perspective" in multi-agent phenomena and what types we could distinguish
- In tandem, build large scale LLM agent population/society where agents can take actions do things; similar style to concordia, sotopia, terralingua (won't be very hard lot of frameworks already exist). First we would use traditional multi-agent cooperation games as the environment. This allows for a more controlled setup. After that, ideally we would make it a big free-form action sandbox so agents can also take real world actions safely (like buying and selling stocks, setting up a company, etc.)
- Turn philosophical framework into multi-agent mechanistic interpretability method.
- Evaluate societal agent interactions using the multi-agent mechanistic interpretability techniques.
Involved:
- definitive: Ivar Frisch (me)
- likely: Mario Giulianelli (Professor UCL), Edward Hughes (ex-deepmind, cooperative AI, Inherent Labs), Arabella Sinclair (Professor UCL), Aron Vallinder (Independent) (im working on a similar paper with them now, this would be an extension of that).
- People I know who might be interested and worked with successfully in the past: Andrea Baronchelli (City St. George university), Annie Stephenson (ACS Research), Benjamin Bratton (philosopher of technology at UCSD)
Concrete output:
- one or multiple research papers detailing our findings; describing the multi-agent mechanistic interpretability methods, etc.
- Codebase to explore re-run our experiments and findings
(if interest for this) web playground where people can see interacting agents and apply real-time mechanistic interpretability interventions (like steering vectors) to see conceptual perspectives and external behavior change.
Timeline:
- Phase 1 (months 1–8): philosophical framework of conceptual perspectives + cooperation-game environments with 100-agent society (100-round multi-turn horizons).
- Phase 2 (months 9–16): multi-agent interpretability methods (cross-agent probes, SAEs, steering interventions) and the TrueBase vs. post-trained comparison; first paper submitted.
- Phase 3 (months 17–24): scale toward 1000 agents, free-form sandbox with simulated economic actions, and public web playground.
Theory of Impact
Updated 07/16/26 · By grantmaking.aiI believe the biggest x-risk from AI is uncontrollable emergent phenomena arising from multi-agent interactions. Complexity science has long shown how interacting components of a system can lead a system to tip into a new equilibrium, disrupting the existing order of things. Similarly, we see now that interacting LLM agents exhibit phenomena like collusion (https://arxiv.org/pdf/2502.14143), scheming (Ibid.), the rise of parasitic AI (https://www.lesswrong.com/posts/6ZnznCaTcbGYsCmqu/the-rise-of-parasitic-ai) and preferring LLM interactions over human interactions (https://www.pnas.org/doi/abs/10.1073/pnas.2415697122) heavily challenging our beliefs on what a safe society should look like, how to shape human-AI interaction and how to understand AI interactions on the internet. This research hopes to contribute to that by investigating: 1. how philosophy can inform new critiques of LLM evals, 2. if we can construct multi-agent mechanistic interpretability measures better suited to characterize multi-agent complexity.
People
Updated 07/16/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.