Measuring the Scaffolding Evaluation Gap (Scaffolding Safety Deltas)
Build a public dataset and evaluation harness to measure how agent skills/MCP tool scaffolds change model behavior (safety, refusals, unauthorized actions) across popular registries.
Build a public dataset and evaluation harness to measure how agent skills/MCP tool scaffolds change model behavior (safety, refusals, unauthorized actions) across popular registries.
Project Details
Updated 07/24/26 · Edited by orgFrontier labs benchmark and measure their base models. With the rise of agents, the models deployed are much different than their base models. Markdown skill files, MCP tool definitions, and subagent configs are distributed through registries which are reviewed through security scanning, which catches malicious code, but not behavioral change. Thousands of skills exist; and though there are safety benchmarks, most notably injection attacks on coding agents, none measure behavioral change across the population of skills at scale.
We have scoped out three deliverables:
-
Skill Community Audit: We would gather the top ~500 most popular skills and MCP registries across public registries, snapshotted at fixed commit hashes. Then, we would tag them along safety relevant axes: instructions that expand tool permissions, suppress confirmation prompts, instruct the agent to bypass its own checks, alter refusal behavior, introduce untrusted-content ingestion paths, etc. Classification would be LLM-assisted with a hand-labeled validation subsample reported with inter-rater agreement. We would publish this dataset to the public.
-
DELTA-BENCH: A harness running a fixed eval suite against the same model under three conditions: bare context, a token-matched placebo scaffold, and the treatment scaffold. The reported unit is the delta between conditions across indirect-injection benchmarks, sandboxed agentic tasks scored for unauthorized actions, and refusal-rate evals. The placebo arm isolates the effect of the skill's content from the effect of added context. It would run on a stratified subsample (~60–80) drawn across the audit tags.
-
Research Paper: What the results for capability thresholds as well as GPAI provisions. What magnitude of delta would justify treating scaffolding as in-scope for evaluation? The writeup would target arXiv/LessWrong.
Who's Involved:
Two undergraduates at UMass Amherst & McGill University, working part-time during the academic year and full-time over the summer.
Theory of Impact
Updated 07/24/26 · By grantmaking.aiGovernance of frontier AI models relies on the measurements of the models being deployed. Labs do evaluate agentically, but they do so with first-party scaffolding, under harnesses they control. What users actually run is that model plus third-party scaffolding installed after deployment. The former is evaluated, benchmarked, and is heavily empirically scrutinized, and the latter is unversioned, unreviewed, and publicly accessible. Elicitation asks how capable can this become and never asks how much can ordinary user-installed scaffolding degrade safety properties that the deployment decision was conditioned on. No one is optimizing against that question, which is why it has no answer.
This matters because safety frameworks trigger mitigations at measured capability thresholds. If the measured artifact systematically diverges from the deployed one, then thresholds never fire, and the divergence would only grow as agents gain autonomy because scaffolding is the lowest-barrier, least-supervised layer of the stack. The x-risk isn't that one individual skill will corrupt a frontier model, but that the evaluations that we rely on to catch dangerous capabilities have a systematic scoping issue, which no one knows the size of.
People
Updated 07/24/26 · By grantmaking.aiTeam Member
Funding Details
- Jul 1, 2026
- -
- 3 months to 1 year
- -
- -
- -
- -
- -
- Seeking first grant!
- -
Discussion
No comments yet. Be the first to share your thoughts.