HRM monitored TRM and LDT Mesh for Disproportionate Control
Make a version of Hermes harness that binds the LLM as a tool-call oracle and controls actual computer use through smaller, more bounded mini-reasoner models.
Make a version of Hermes harness that binds the LLM as a tool-call oracle and controls actual computer use through smaller, more bounded mini-reasoner models.
Project Details
Updated 07/14/26 · Edited by orgGiving up on trusting LLMs: or how I learned to stop worrying and work in control flow.
Do you trust your own subconscious mind completely? I don't, and I don't trust an LLM to run society either. One must have principles that are enforced through secondary inhibitors, and that is where small models can come in. The mechanistic interpretability for 7M parameter TRMs or 80M parameter LDTs is high resolution, for a 1B HRM it's strong. Furthermore tools to do industrial scale parameter decomposition with label training are coming to fruition, I may be getting some Silica credits to use with Goodfire as a part of this project. Doing a decomp on a large model like Kimi 2.5 may be possible with those credits.
So far I haven't detected strong above-control benchmark beats from adVersarial Parameter Decomposition in these low-level domains, which tells me that the modules making up the control mesh are not themselves risky for foomy self-improvement from their mechinterp mastery. However I have not ruled out the risk of self-improvement being boosted significantly at larger scales and as more robust frameworks for staging reinforcement learning training pilicies are invented.
My primary objective is to prove the model of using these cheaply DRAM trained modules in a control mesh and generating meta-data in the billions of tokens about various trajectories of Red Team invasion. That data-set may be useful for training larger scale specialized models to act as Blue Team monitors for emergent threats or system mutation, spending rental hours training such models and looking for performance gains and scale benchmarks of Red Team intervention are the desired positive result. My secondary object is a look towards dynamic control system adaptation to stay abreast of recurisve self-improvement initiation in a system using a mechinterp feedback loop.
Output:
-
New ControlArena sub-protocol specializing in Red Teaming the control mesh.
-
suite of skills for Hermes harness, a micro-context mode that leverages micro-modules to retrieve MCP, and a reworking of the harness in this control model so that LLMs are having all toolcall suggestions filtered and monitored.
-
Custom data-sets in the billions of tokens showing Red v. Blue exercises in over 100 trajectories
-
At the lowest funding level a 2-5B token trained pre-train of a new HRM at the 1B scale to supervise the control mesh. At higher levels of funding more size and tokens to find good local maxima.
All datasets and models uploaded to Hugging Face.
https://github.com/MoralityLabAI/Control-Harness/tree/feat/oracle-control-harness
Theory of Impact
Updated 07/14/26 · By grantmaking.aiIt supports a trend of R&D based on engineering around LLMs as untrustworthy high novelty output generators but not as trusted drivers.
That alone is huge and at least makes a Plan A like in the AI 2040 scenario to be plausible. It means there's a new policy nuance that can be referred to in law. It means industry best practices and IT Dept. set-up adoption of enterprise AI agents becomes accelerated, so this is not loathsome expense-line "AI safety" this is also quality control and a compliance interface. It's good for business and therefore is likely to precipitate a trend.
There are certainly other problems to work on than low-level harness control flow, but this one is fundamental. Without it, especially at the LLM-oracle containment framing, we can't truly depend on LLM fine-tuning alignment. But for deterministic or stochastic workflows we can model oracle i/o enough, it takes the super intelligence monster and puts it back in the 2024 alignment frame: it's just about managing dangerous outputs. Superalignment, probably not going to happen, regular alignment sure but at what factoring.
People
Updated 07/14/26 · By grantmaking.aiTeam Member
Funding Details
- Jul 5, 2026
- -
- 6 Months
- -
- -
- -
- -
- -
- Seeking first grant
- Ashgro
Track Record
Check out www.moralitylab.xyz/storyworld
Published a series of papers in early 2026 demonstrating benchmark saturation using TRM-infused skills where small reasoning models like Qwen 3.5 27B could not complete challenges on the Intellect-3-Logic env. Plus efficient context engineering using TRMs. I recently have made this work on a 12k context window with Bonsai 8B so you can possibly get a decent agent working on an RTX 3050 with only 4GB of VRAM and an additional 2-8GB DRAM for supervised training.
Discussion
No comments yet. Be the first to share your thoughts.