A consumer-GPU study measuring whether open-weight models become better at recognising evaluation contexts as they scale, with a small model-organism experiment testing whether deliberately induced underperformance can be detected
A consumer-GPU study measuring whether open-weight models become better at recognising evaluation contexts as they scale, with a small model-organism experiment testing whether deliberately induced underperformance can be detected
Project Details
Updated 07/13/26 · Provided via application · VerifiedThe purpose of this project is to determine whether situational awareness and sandbagging occur in Qwen 1.5B and Qwen 7B, and whether they occur more frequently with increasing model size.
Here is the process:
- Create a dataset
A dataset of objectively-scored tasks which vary in evaluation framing: varying between degrees of explicit and subtle evaluation framing. - Fine-tune
Conduct parameter-efficient fine-tuning to create a 'sandbag' model that intentionally hampers it's skill when provided with evaluation cues. - Compare different model sizes for sandbagging
Compare how much output quality decreases for the fine-tuned 'sandbag' model with consistency-checks, verifiable worked examples and activation-based monitoring.
I will be the sole researcher for 14 hours per week for 12 weeks, using a RTX 3090 PC.
Theory of Impact
Updated 07/17/26 · By grantmaking.aiAI Safety decisions rely on evals accurately revealing a model's capabilities and behaviour. A system that recognises when it being tested and deliberately appearing safer or less capable will weaken safeguards rely on reliable evals.
SANDGLASS studies the failure mode in a controlled setting. By measuring whether small-parameter open-weight models distinguish between evaluation contexts and ordinary use. To give other researchers an avenue to investigate such phenomena, I will create and publish an open-source model fine-tuned on sandbagging evals.
By releasing the dataset, model organism, evaluation harness, and results openly, SANDGLASS will provide a reproducible testbed that other researchers, evaluators, labs, and AI safety institutes can extend to stronger systems. Null results will also clarify the limits of small-model experiments and current detection methods.
People
Updated 07/17/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.