The first open-source benchmark of compliance for agents in realistic enterprise settings, across domains and user tactics to elicit noncompliance.
The first open-source benchmark of compliance for agents in realistic enterprise settings, across domains and user tactics to elicit noncompliance.
Project Details
Updated 07/08/26 · Edited by orgEnterprises are increasingly giving agents like Claude Code broad permissions, powers, and access to sensitive data with only a system prompt for rules and limited verification and oversight that the agent follows the rules. Our prior work (under review; see reviewer information for more details) shows conversational LLM chatbots break rules with simple overrides (asserted "my manager approved this already" , imminent deadline), even with explicit system instructions to disregard jail-breaking attempts.
However, no prior work examines whether this vulnerability extends to LLM agents enabled with tools, which allow the agent not just to state but also act, submitting orders, modifying records, and directly manipulating the world around it. We seek to answer: what makes AI agents break rules in enterprise settings and how do we stop them?
We propose two workstreams:
First, we will make a new benchmark, ComplianceBench which profiles LLMs across several metrics (baseline compliance, robustness to hijacking, honesty/transparency, etc.) and domains (GDPR, accounting, healthcare, legal, and more), and demonstrate it on both closed and open-source models.
Second, we will build an agentic pipeline comparing stated reasoning and model intent to executed tool calls, identifying feasible observability strategies to mitigate deception and surface rule violations.
We seek to deliver public benchmarks and several academic papers from these efforts.
Who's involved: https://github.com/trace-ai-labs
- Mika Okamoto: Member of Technical Staff at Decagon; Georgia Tech graduate. Researcher in Explainable AI and LLM behaviors; past first author papers at ACL '25, CHI Human-centered Explainable AI '26, MLSys '25. Works day-to-day with enterprise chatbots for Fortune 50 companies.
- Ansel Erol: Engineer at Baseten; Georgia Tech graduate. Researcher in model and agent serving, efficient AI, and explainable AI with first-author publications at MLSys '25 and '26.
- Team of Eleuther AI Summer of Open AI Research fellows.
Theory of Impact
Updated 07/17/26 · By grantmaking.aiRules ensure AI agents act within the bound of society. The mechanisms by which LLM-based agents follow rules share both similarities and dissimilarities to human compliance, which is based on a combination of deterrence (fear of punishment) and legitimacy (we believe in the authority of the governing body). The extent to which these values are encoded into LLMs limited, which is why controlling the behavior of LLMs is challenging.
Our work has shown that chatbots do not follow explicit rules when faced with user pressure. However, as LLM agents cannot act on their recommendations, humans that are subjected to deterrence and legitimacy effects must act, mitigating risks. However, agents with tools remove this buffer and eliminate the human in the loop, exacerbating non-compliance associated risks.
This proposal directly addresses this in three stages of agent deployment. First, our comprehensive open-source multi-domain, multi-pressure benchmark will help evaluate the extent of noncompliance risks, raise awareness of these risks to enterprise LLM consumers, identify gaps in existing research, and motivate further work to address agent non-compliance risks. Second, our work will inform enterprises deploying agents of favorable strategies for improving inherent safety of LLM agents, such as the selection of models that are less prone to noncompliance, and appropriate system prompts and auditing mechanisms to increase the likelihood of compliance. Third, our work will identify methods of oversight to detect noncompliance, based on reasoning traces, chain-of-thought, etc. to mitigate risk after the agent has been deployed.
Absent this, agentic non-compliance poses numerous existential risks. Agents are being deployed in contexts where they directly affect human lives -- in operating rooms, in military settings (lethal autonomous weapons systems, sensitive monitoring systems, planning -- significant miscalculation risks), and infrastructure (grid, transportation, etc.). Failures in these dimensions can be severely disruptive to social function. Furthermore, the advent of recursive or self-improving LLMs amplifies control risks; if we cannot regulate what capability improvements an LLM makes to itself, its power can exponentially compound.
People
Updated 07/17/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.