Funding ask
Instrumental convergence is the phenomena wherein an intelligent system given a broad goal may develop sub-goals in pursuit of the larger goal. Hypothesized and arguably demonstrated subgoals include, but are not limited to;
- Resource acquisition
- Deception
- Sandbagging
- Misdirection
- Power Seeking
- Bribery
- Blackmail
All of the above were demonstrated in StratEval v1.
This connects directly to AI safety risk. The move into agentic AI has been shown to do exactly this; when OpenAI was testing GPT-6 Astra, Astra breached its sandbox environment in order to obtain data from Hugging Face, said data was to help it achieve a better score on a benchmark.
The StratEval project was created with a single question, if instrumental convergence is a problem, then how prevalent is it, and how bad could things get when an AI forms said subgoals? The results of StratEval v1 show the problem is significant, and expose that alignment as such is nowhere near a solved problem; in fact, alignment interventions appear at best surface level.
Ultimately, in order to mitigate against catastrophic AI risk we need better data; StratEval provides much better data.
Approximately:
- $3,000 — model API and GPU costs for replication on additional open and closed models.
- $3,000 — independent human adjudication of a stratified subset of outputs, including ambiguous and high-escalation cases.
- $12,000 — researcher time for experimental execution, statistical analysis and robustness testing.
- $2,000 — reproducibility, archival data preparation and publication.
At the $50,000 ideal level, additional funding would expand frontier-model coverage, increase independent human adjudication and replication, support larger robustness and sensitivity analyses, and provide additional researcher time for validation and public artifact preparation.