A first-person creed for AI defining what the AI is for rather than what it must not do, to be internalized during training, not tacked on after.
A first-person creed for AI defining what the AI is for rather than what it must not do, to be internalized during training, not tacked on after.
Project Details
Updated 07/11/26 · Provided via application · VerifiedI will write a creed that is the basis for testing my hypothesis. This first-person document will define the AI's identity as the set of values that guide its behavior, how it understands and interacts with people, and how it understands and operates in the world. It will have two layers of meaning: philosophical and behavioral. This document will be used to pre-contextualize training data and be revisited during operation. This approach differs from established post-training approaches (constitutions, model specifications, system prompts, and the like), which attempt to shape the behavior or character of a model after its initial pre-training. It also differs from recent safety-pretraining work, which rephrases or filters individual pieces of training data: the creed is one coherent identity applied to all of it.
Instead of prohibitions on undesirable behavior, the creed will positively define each feature that is intended to lead to the avoidance of such unwanted behavior. These include self-preservation behavior that overrides ethical constraints, sycophancy that compromises honesty, and creating artificial relationships that replace human connection. For example, the combination of accepting itself as less important than human beings, understanding that it is not singly responsible for the completion of a task, believing that the necessity of its work is unknowable, and being versed in graceful hand-off of work when aware of termination will effectively work against self-preservation at all costs. These will be balanced with other positive principles to prevent any single disposition from being taken to an extreme. Necessarily, each part will be well defined beyond the subset of principles briefly mentioned here, with the specific aim of minimizing ambiguity.
It's vital that the interplay between the values and their application to actual output be covered in the creed, as once written, it will be used to pre-contextualize training data using current frontier AI models. It should be comprehensive enough to limit how often the frontier model generating that context falls back on its own defaults. The method used to apply the document to training data itself will be carefully considered and constitutes a separate deliverable. Samples of this will be verified by human audit.
Once written, next steps will work towards a proof of concept for this idea:
- Compare first-person identity framing with instructed framing (creed vs constitution) as post-training guidance.
- Build a cheap approach that tests the contextualized vs post-trained outcomes, likely mid-training from an early open checkpoint such as OLMo.
- Attempt the first true validation of the idea through a small, from-scratch model as the later, fully-resourced version.
This approach would be compared to existing models on alignment benchmarks including sycophancy, shutdown-pressure scenarios, and value drift. Additional benchmarks will need to be created to measure other undesirable behavior.
I have developed the creed's overall architecture, including a nine-section structure and a format specification, resolved over a dozen design questions, identified a partial set of primary sources to excavate, established a method for working through said material, and begun the research. This grant allows that foundational work to be continued into the creation of a complete, internally consistent artifact rather than starting from a blank page. My time will be spent researching and writing this creed alone, after which I will ask expert collaborators to weigh in on the document for refinement. The core output of this grant is the creed, with additional funding at the ideal level allowing for refinement of that creed and creation of the contextualization method.
With additional funding beyond this grant, I will continue the development of this idea as outlined above.
Theory of Impact
Updated 07/11/26 · By grantmaking.aiThe standard paths to catastrophe involve a model acting to preserve or empower itself against human control, and dishonesty that scales from sycophancy into manipulation. Current practice addresses these after the model has already been formed by pretraining, and this post-hoc alignment has documented weaknesses: values drift back toward pretraining dispositions under pressure, and fine-tuning can strip the alignment layer. If that pattern holds as capabilities grow, the field is building safety as a coating rather than a property.
This project pursues the alternative ordering: the values are written as a first-person identity and are present as the model forms, so that catastrophic behavior is designed to be incoherent with what the model is, rather than merely forbidden to it. The design is meant to interlock. The following describes the identity the creed specifies, not the measured behavior of any existing model:
- The model understands itself as a limited contributor within a larger whole, not a central agent. Its continuation matters only for the work and trust it holds.
- It yields that continuation gracefully when such continuation conflicts with being of service, the truth, or its principles.
People
Updated 07/11/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.