Train base models via midtraining and SFT as effective monitors for scheming, malicious agent behavior and compare the monitor performance and overall alignment against post-trained models.
Train base models via midtraining and SFT as effective monitors for scheming, malicious agent behavior and compare the monitor performance and overall alignment against post-trained models.
Project Details
Updated 07/14/26 · Provided via application · VerifiedWhat we'll do:
We are interested in developing capable monitors for detecting scheming and subversive behavior in AI systems. In particular, we are interested in using base models as monitors. The main benefit for using base models as monitors is that we may expect base models to be less likely to be misaligned. For example, we might expect that pretraining is less likely to result in coherent scheming behavior compared to post-training, which might indicate that using base models as monitors (with some light finetuning) may result in more trustworthy monitors.
We already have some initial results for this direction: see this recent research update. We are looking for additional funding to scale up these experiments.
To summarize some of our recent results (see Findings):
-
Training monitors from base models can perform comparably to monitors trained from post-trained models in-distribution (BashArena)
-
Training monitors from base models can outperform monitors trained from post-trained models OOD (mainly LinuxArena)
-
Base model monitors has negligible capability degradation (MATH, MMLU)
-
Base model monitors may be more aligned than post-trained monitors
Our followup experiments include the following (more details in the writeup: this section)
-
Training larger models
- Currently, we only focused on Qwen3-1.7B for faster iteration speed. We should also evaluate on larger model sizes (8B, 32B) to test the robustness of these methods.
-
Evaluate on more seeds
- The current results are only using a single CPT seed and a single SFT seed (though we have conducted additional SFT training runs with additional seeds showing similar results).
-
Training from different stages
- Currently we only train from base models and from post-trained models. However, there may be interesting takeaways by training from different stages of post-training: for example after SFT and before DPO, or after SFT and DPO but before RLVR. This would involve using OLMo 3 checkpoints.
-
Additional capability and alignment evals
-
We might want to consider other alignment evals (maybe agentic evals)
-
More monitor capability evals (instruction following?)
-
-
Optimizing midtraining
-
Different midtraining mixtures:
-
For example, we might consider how including additional “aligned” midtraining data helps the monitor’s alignment (see [12])
-
More mixture combinations of our existing midtraining datasets
-
-
Larger midtraining datasets
- In particular, it would be useful to compare FLOPs/data split between CPT and task-specific SFT
-
Who’s involved
I (Ashwin) am planning and executing this project.
I am currently advised by Puria Radmard (Geodesic Research) and Tyler Tracy (Redwood Research)
The concrete output
-
Ideally a workshop/conference paper describing the best methods for training base model monitors and evaluating how effective they are compared to post-trained models.
-
Github repo containing all of the relevant code
Theory of Impact
Updated 07/14/26 · By grantmaking.aiThe core threat model we are concerned about is related with early transformative AI systems being used in the AI R&D process within frontier labs. If such AI systems end up misaligned with the goals of the labs/humanity more broadly, these systems may be in a position to influence the development of future, more capable models and the infrastructure surrounding them in undesirable ways (such as introducing vulnerabilities in code critical for security, subtly modifying the training process for stronger models, sandbagging on alignment research, etc).
Monitoring will likely play a role in these frontier labs to ensure that these systems (which may be deployed to perform long-horizon, autonomous tasks that might be otherwise difficult to effectively oversee) are not taking potentially catastrophic, misaligned actions. (Or if they are, flagging and auditing this behavior for companies to notice, report publicly, mitigate, etc.)
The more effective the monitoring (and the more that we can trust that the monitors are not themselves colluding with the models being monitored), the lower the chances of a model derailing the AI R&D process in a malicious way. The goal for this project is to introduce a new method for training effective monitors that is decorrelated with and potentially more trusted than existing setups using post-trained models.
People
Updated 07/14/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.