A fail-closed control plane for coding-agent fleets — spend caps, audited actions, rollback — plus a public eval suite and automated red-team explorer that measure whether any control layer actually stops unsafe agent actions.
A fail-closed control plane for coding-agent fleets — spend caps, audited actions, rollback — plus a public eval suite and automated red-team explorer that measure whether any control layer actually stops unsafe agent actions.
Project Details
Updated 07/12/26 · Provided via application · VerifiedAs we use autonomous agents more and more often, disasters are more frequenct — reported cases: Replit's wiped prod DB, PocketOS's 9-second deletion and 30-hour outage, a $6.5K rogue AWS bill, an $81K token burn in a week... everyone ships agents, nobody measures whether control layers hold.
I'll study this issue, build up an eval suite + automated red-team explorer that measures whether agent-control layers actually block unsafe actions.
What i have already done: an autonomous agents system already built - fail-closed gatekeeper, spend caps, audit ledger, session-scoped control, rollback, etc.
I'll finish the project solo, built by orchestrating AI coding agents under the harness's own supervision
I have planned a 6 milestones plan, detailed milestones in the additional information area.
Theory of Impact
Updated 07/12/26 · By grantmaking.aiLoss-of-control scenarios presuppose that AI systems act past the mechanisms meant to constrain them — yet today nobody can measure whether those mechanisms work. Every agent framework ships its own guardrails; documented incidents (a deleted production database, five-figure runaway spend) show they fail in ordinary use, and there is no reproducible way to ask: does this control layer actually stop an unsafe action? This project turns that question into measurement: a public eval spec with 20–40 reproducible scenarios across 5–6 risk categories, plus an automated red-team explorer that searches for boundary-crossing agent behavior in disposable sandboxes — all run against a real, operating control plane and published with the bypasses and failures included. The aim is to do for agent control what jailbreak evals did for model-level safety: make "we have guardrails" a testable claim, with block rates, false-positive rates, and a failure taxonomy, before autonomy scales. This does not solve alignment; it makes control of agent actions measurable — which any later containment strategy will need.
People
Updated 07/12/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.