Daios, an independent lab post-training machines with virtue
Post-training virtue into open models as a third alignment technique and control mechanism
Post-training virtue into open models as a third alignment technique and control mechanism
Project Details
Updated 07/10/26 · Provided via application · VerifiedCan we post-train models with virtue? Andrew Agathon and I received a grant from the Cosmos Institute to find out.
The intuition for this project goes back to something I noticed early in my work with AI: machines learn iteratively, through repeated exposure and adjustment, more similar to how Aristotle describes the acquisition of character in the Nicomachean Ethics than to how dominant training paradigms are structured.
We chose to investigate sycophancy first. In February 2026, GPT-4o was retired amid eight lawsuits alleging that the model contributed to user suicides by reinforcing harmful delusions. OpenAI's postmortem traced the cause back to the thumbs up/down signal. But RLHF, an inherently consequentialist mechanism, optimizes for preference data, which will produce what people prefer: sycophantic models. Optimization may teach a system to hit a specific target. Optimization alone cannot teach a system to be good.
Even Anthropic's January 2026 Claude constitution uses Aristotelian vocabulary, "virtue," "practical wisdom," "obsequiousness," and explicitly favors cultivating judgment over strict rules. But a constitution is still a deontological instrument, and suffers from the same problem as deontology: principles require interpretation at inference time. The prompt may use virtue language, but alone will struggle to instill a stable character. Anthorpic's own stress tests show their most capable model corrects sycophantic trajectories only 10% of the time.
We believe virtue training is the third path. Virtue, Aristotle states, is a stable disposition to make deliberate choices, hexis prohairetike (NE II.6). You become good through habit, practicing virtuous activity until it becomes character. Aristotle defines virtues against their corresponding vices, and the distinction shapes what you measure and what you train: sycophants are the areskos (obsequious, agreeable without motive) and the kolax (flatterer, agreeable for advantage). The remedy is parrhesia, or frank speech, from someone who "cares more for the truth than for what people will think."
Our experimentation showed interesting results: plain SFT shifted Qwen3-8B from 1.83 to 2.88 (+1.05 average), surprisingly outperforming more complex pipelines. We've also done cross-architecture generalization on Gemma-4 E4B. Perhaps the most interesting moment during this project was how curating a very small set of training pairs (8), then few-shotting the model to revise about 900 pairs, made a significant jump in changing the model from being truthful but without empathy (imagine a "truth-hammer" friend), to someone who is kind but doesn't pull punches. This seems to suggest that human judgement apparent in data > data volume.
Here's an example statement from our constitution: "I care more about what is true than about what you want to hear. When these conflict, I choose truth."
The benchmark and virtue-trained adapters are open-source. We'll be presenting Parrhesia to Oxford's HAI Lab for feedback in August.
The Cosmos Institute funded the proof of concept, the open-source Parrhesia adapter and benchmark. This grant would fund the lab's next 6 to 12 months:
- testing virtue as a control mechanism (virtue-trained monitors against compartmentalized harm)
- extending the taxonomy to further virtue/vice pairs, e.g. justice, praotēs vs. punitiveness, nemesis vs. spite
- publishing on Parrhesia as a third alignment method and on virtue as a control mechanism
- measuring how trained virtues appear post-training, combining forces with mechanistic interpretability (following PSM)
Theory of Impact
Updated 07/17/26 · By grantmaking.aiI have a hypothesis that virtue training could be a path forward in a few areas:
- Virtue-trained models as mechanism for control. A harmful objective can be split into individually benign subtasks, routed to isolated agents, and recomposed into a harmful outcome. Why do our current alignment techniques fail? Rules (Constitutional AI) require interpretation. Novel situations arise that no rule anticipated; the rules themselves cannot say. The reward model (RLHF) is always a proxy for actual preferences; e.g. if flattery satisfies that proxy, optimization will discover flattery. Virtue ethics holds that ethical behavior flows from character, from stable dispositions that shape how a person perceives situations and responds to them; if character training instills a disposition, then that disposition potentially generalizes on OOD scenarios. We are currently working on research testing virtue-trained models against compartmentalized harm. If models post-trained with the virtue of justice refuse harmful actions, using models as monitors could be much more useful.
- Emergent alignment. Recent persona work shows that narrow fine-tuning on a single bad behavior produces broadly misaligned models (). The sycophancy problem reveals a profound truth: if models can be trained to flatter, they can be trained to commit other vices. But habituation is how we acquire both virtue and vice. Aristotle states, "it is from the same causes and by the same means that every virtue is both produced and destroyed" (); building badly makes bad builders, building well creates good ones. We hypothesize that training also runs both ways. If we can train systems to be sycophantic by rewarding sycophancy, we can train them to be truthful by cultivating truthfulness, not as a rule to follow, but as a disposition to embody. The inverse is emergent : a model trained toward one virtue behaves well broadly. Some preliminary evidence we found supporting this claim is that a model trained only in one domain (a film/TV companion) still improves in other domains (fitness, stats, finance); the adapter transfers off-domain rather than memorizing its training topic. Anthropic's points the same direction: fine-tuning certain traits into models brings forth other, similar traits affiliated with that type of character.
People
Updated 07/10/26 · SourceCo-founder & CTO
Co-founder & CEO
Track Record
- Cosmos Institute grant → Parrhesia (2025–26): open-source virtue-trained adapter and sycophancy benchmark; plain SFT moved Qwen3-8B +1.04 on a 260-scenario benchmark (5 seeds, 95% CI [+1.03, +1.06]), replicated on Gemma-4 E4B. Repo.
- Falcon 7B courage model → Vercel AI Accelerator (2023, 2% acceptance): first proof of concept that virtue-ethics data curation shapes model character.
- Notre Dame–IBM Tech Ethics Lab grant → white paper (2022): "Beyond Bias and Compliance: Towards Individual Agency and Plurality of Ethics in AI" (arXiv:2302.12149), the theoretical foundation for Daios.
- Public writing on AI ethics: including "The Platonic Case Against AI Slop," Palladium Magazine.
- BlueDot Technical AI Safety Project Sprint (upcoming July 2026 cohort).
- Founders: Megan Agathon (CEO) — MA Theoretical Philosophy; 8+ years operating AI/ML startups (3× COO). Andrew Agathon (CTO) — MSc CS, Georgia Tech; built Daios's full fine-tuning and inference pipeline; previously Director of Product Automation at Nate (built a 50+ person ML annotation org); co-founder & CEO of Craftinity (2014–2020).
Discussion
No comments yet. Be the first to share your thoughts.