Project Details
Updated 07/28/26 · Edited by orgCaML researches how to use midtraining (/Synthetic Document Finetuning) to make AI systems morally open-minded and compassionate toward all sentient beings. In this project, we’ll be researching how our midtraining data can be scaled so that the desired effects are not degraded by later fine-tuning.
Our previous work includes:
-
Releasing benchmarks on Inspect (including TAC, MCB, MORU and ANIMA) evaluating moral reasoning under uncertainty and compassion in frontier models
-
Conducting research showing how compassion can be effectively instilled during midtraining, but that fine-tuning will sometimes erode desired values
-
Finding that broad compassion for non-humans generalises to humans
Midtraining may produce deeper values (e.g. here) without the risks of RL, suggesting it is likely to generalize better than finetuning to transformative AI and is already used in frontier model pipelines. But we lack research into how different aspects of midtraining are affected by scaling, or how self-fulfilling alignment can be used. Some specific questions we’re looking to answer in this project include:
-
Does scaling midtraining data affect performance on our semi-agentic TAC benchmark for given post-training?
-
What mechanisms degrade compassion in finetuning?
-
How does scaling the fine-tuning (SFT, RLAIF, RLVR) and (compassion) midtraining affect persistence of the chosen value?
-
How do these effects scale with larger open-weights models?
-
Investigate how this self-fulfilling alignment interacts with self-other overlap techniques (in collaboration with Geodesic)
We’ll publish the results of these findings through arXiv. If the work demonstrates promising mechanisms for scaling compassionate midtraining, we’ll seek to present it directly to labs, and release a scaled dataset for the alignment community.
Theory of Impact
Updated 07/28/26 · By grantmaking.aiWe want to ensure our work is as useful as possible to frontier labs. We have heard consistently from our contacts in labs that they have a good chance of being able to integrate our research into their pipelines if we can demonstrate that our techniques work, are low-cost, and don’t undermine other priorities.
CaML’s focus in this project is demonstrating to them:
-
How midtraining can produce self-fulfilling alignment
-
How compassion can be more robustly integrated at scale
-
Why these interventions are good for alignment, and don’t harm capabilities or user friendliness
Our contacts at multiple major labs have told us that the biggest barrier to integrating promising research is uncertainty on whether it will hold at scale. The project we are seeking funding for seeks to answer that for self-fulfilling alignment, using compassionate midtraining.
Compassion to nonhumans is an excellent test case for robustly instilling broad and positive values, because it encourages moral circle expansion, and promotes care to vulnerable beings in ways that seem to generalise to the human case. It is also orthogonal to conventional finetuning, allowing cleaner interpretations of the results of midtraining.
People
Updated 07/28/26 · Edited by orgTeam Member
Funding Details
Only visible to verified funders, reviewers, and admins.
Email hi@grantmaking.ai to get verifiedTrack Record
CaML has published 7 papers on Arxiv this year and have been features in DailyPapers.io. We were speakers at Sentient Futures San Fransisco and Sentient Futures London, where wide audiences attended our talks. We have been the leading organization producing nonhuman welfare benchmarks on the UK AISI's Inspect and have now produced two agentic benchmarks to address previous feedback from researchers in frontier labs and maintain leaderboards for these benchmarks. We have also co-lead the Hyperstition compatition which has caused us to enter negotioations with senior AI research organizations for our synthetic data generation pipelines. We have been mentors to over 15 researchers in the Sentient Futures incubators, resulting in ML papers released and some mentees receiving their own funding for technical alignment projects. We hope to continue guiding new technical talent into the field.
Discussion
I've funded this through the Falcon Fund and I think this is a good (though low probability of success) bet
Just confirmed with Marcus that he was referring to having also funded (through the Falcon Fund) similar work building pipelines focused on scaling compassionate midtraining data for animal welfare, specifically this one led by Sentient Futures
Hi, Jasmine
I saw you recently added a 72k Grant, did it cover this application's ask, or you are seeking additional 9-50k?
Hi Anton,
I don't think we added that recently I thought that was there from the start? I added the Goodfire compute grant recently. Yes, we already had the 72k grant before we began this submission (for other projects) and definitely require more funding.
Thanks,
Miles
Got it, thanks! I must have misread the dates
It seems we are digging at the same root from two sides. You instil character through midtraining data and ask whether it survives later fine-tuning; and I ask whether the structure of how data is selected leaves a fingerprint in character at all. Both of us are betting that the data stage.
The "self-fulfilling-alignment-surviving-fine-tuning" part is the one I find most alive, because the failure direction is documented (a narrow careless fine-tune can break a model, Jan Betley.), and you are testing the constructive direction of the same lever. Genuinely glad this project exists. 👏 Best of luck!
Thanks Katja, great to hear you're working on character data attribution, it seems really important!
Thanks for your kind reply, Miles. Character data attribution is probably the quiet core of this: whatever character we want models to keep, someone has to trace what shaped it.
And since I believe projects in a round like this should hold each other up: any comment, critique or support from you under my project would be genuinely valued. Critique maybe most of all. 😊
Quick clarification, since my earlier comment may have muddied this. Sentient Futures does field building. We do the technical work, and we advise them on the compassion scaling pipeline. The training runs, the benchmarks, the papers, and the research we've done on whether instilled values survive later fine-tuning are ours.
On fit for this round: the general question we are testing is whether values installed at the mid-training stage persist through subsequent fine-tuning or wash out. Compassion is our test domain because we have benchmarks for it and a measured result to build on, but the failure mode we are probing, values that look installed and then vanish under later training, applies to anything a lab tries to instil.
CaML does rigorous, empirical technical research in order to better understand AI behavior and motivations. Their research is currently outside the Overton window of most AI safety researchers, but I believe that the questions they are asking will become very relevant in the next few years as AI systems continue to improve in complex behavioral capabilities and autonomous deployment. I'm excited to help contribute to push this research forward!
CaML has pioneered a lot of the novel research around technical AI safety for nonhuman welfare, and I'm looking forward to seeing what new outputs they produce.