Implementing evals for RL and LLM agents' ability to learn and properly apply biologically and economically aligned pluralistic utility functions and with that to avoid runaway conditions.
Implementing evals for RL and LLM agents' ability to learn and properly apply biologically and economically aligned pluralistic utility functions and with that to avoid runaway conditions.
Project Details
Updated 07/23/26 · Edited by orgWe investigate how more aligned, safer, corrigible, and interruptible AI systems can be built following the principles of:
-
homeostatic bounded objectives;
-
multi-objective balancing of both bounded ultimate and unbounded instrumental objectives;
-
pluralistic universal human values that oppose each other by design (Schwartz Value Circumplex);
-
and proactive horizon scanning for unexpected side effects.
Concurrently, we have been implementing various long-horizon evals that elicit “runaway failure modes” contradicting the above listed principles. We believe these principles are partially neglected and need much more attention.
Theory of Impact
Updated 07/23/26 · By grantmaking.aiOur research agenda is summarised here — https://threelaws.net/#research-agenda — covering highlights of two LessWrong posts:
A further trimmed down summary of the above posts is the following:
People
Updated 07/14/26 · Edited by orgResearch lead
Funding Details
- -
- -
- -
- -
- -
- -
- -
- -
- unfunded
- -
Track Record
- A brief review of the reasons multi-objective RL could be important in AI Safety Research (LessWrong and Alignment Forum)
- Using soft maximin for risk averse multi-objective decision-making (AAMAS journal)
- Why modelling multi-objective homeostasis is essential for AI alignment (and how it helps with AI safety as well) (LessWrong and Alignment Forum)
- Systematic runaway-optimiser-like LLM failure modes on Biologically and Economically aligned AI safety benchmarks for LLMs with simplified observation format (BioBlue) (LessWrong)
- From homeostasis to resource sharing: Biologically and economically aligned multi-objective multi-agent gridworld-based AI safety benchmarks (arXiv)
Discussion
No comments yet. Be the first to share your thoughts.