I wish to study how ideas from Tibetan buddhism, specifically the dzogchen tradition, even more specifically an ancient contemplative practice called Ngöndro, can be used to train aligned agents in the long time horizon setting.
I wish to study how ideas from Tibetan buddhism, specifically the dzogchen tradition, even more specifically an ancient contemplative practice called Ngöndro, can be used to train aligned agents in the long time horizon setting.
Project Details
Updated 07/14/26 · Provided via application · VerifiedIf this grant is funded I will use these resources to run large scale experiments that are beyond what I am able to achieve as an independent researcher.
I plan to use the funding from this grant to pay for time on a service like runpod to run experiments where I use an approach like LLM as a Judge, and subsequently, LLM as a Jury, where the judge/jury prompted is grounded in a teaching like this one: https://www.lotsawahouse.org/tibetan-masters/patrul-rinpoche/brief-guide-ngondro
Specifically what I plan to do is to initialize models from a diverse set of prompts which encourage agents to embody chaotic behaviors (as well as other failure modes, e.g. paperclip maximizer). I aim to evaluate the extent to which a principled training regimen rooted in a time-tested wisdom tradition like dzogchen is capable of meaningfully changing behavior of models during long rollouts.
I plan to publish the results of these experiments on huggingface, and plan to primarily use open weights models in order to ensure that the experimental results are reproducible.
Theory of Impact
Updated 07/17/26 · By grantmaking.aiI believe that it is only a matter of time until the stochastic parrots escape the datacenter and begin to live free on heterogeneous hardware on the edge.
This project aims to study ways to minimize the potential risk of people deploying extremely long running agentic jobs to the web.
I am particularly worried about the implications of how distributed inference could enable "self-sovereign" processes to potentially run somewhat indefinitely, and wish to explore how giving models a simulacrum of free-time during training (and ideally during inference if they are continually learning models), could serve as a straightforward intervention that labs could start doing before releasing more capable open weight models to the world in the future.
People
Updated 07/17/26 · By grantmaking.aiTeam Member
Funding Details
- -
- -
- -
- -
- -
- -
- -
- -
- Seeking first grant
- -
Track Record
Josh recently received a PhD in Computer Science from UT Austin, and is hoping to use this grant to transition to contributing to AI safety in this critical period of societal development.
What is the ROI curve like on the gap between $10k and $1m?
I would spend $10k entirely on compute in order to perform experiments in the cloud. I have already run these experiments locally, and the code can be seen here: https://github.com/thenthfool/freetimebench
This funding would hopefully enable me to generate results that are convincing enough to derisk further investment into this project.
If I were to steward a significant amount of resources, like $1m, I would be able to hire a researcher or two that have experience and interest in these problems, and perform long time horizon experiments in the multiagent setting. I am already seeing a positive result in training "canon formed agents" on my laptop, in improving the outcome on a multiplayer eval (see plot here: https://github.com/thenthfool/flockbench ), but carefully testing these ideas would take a lot of compute, and studying multiagent coordination is somewhat at the frontier of research, so might take a bit of an unpredictable amount of time to achieve something useful.
Thanks!