A strategic model of governance under uncertainty: how disagreement about AI consciousness undermines the coordination that keeps AI risks in check.
A strategic model of governance under uncertainty: how disagreement about AI consciousness undermines the coordination that keeps AI risks in check.
Project Details
Updated 07/13/26 · Provided via application · VerifiedThis project explores how disagreement about AI consciousness is an underexamined driver of x-risk; the question here is not what AI systems are or deserve but a strategic one of how discord over consciousness changes the behaviour of self-interested actors (governments, frontier AI companies, researchers and the public—whose stance is still largely uninformed but will become increasingly consequential) under moral uncertainty. Actors face different incentives to recognise, deny or invoke AI welfare: companies may underinvest in welfare protections because competitors do the same, governments may delay regulation whilst awaiting consensus, researchers may face incentives to publicly endorse/reject consciousness claims, and AIs themselves may exploit uncertainty by making persuasive welfare claims that constrain oversight. Using game-theoretic modelling, this project will analyse how these strategic interactions shape institutional outcomes, identifying the conditions under which welfare governance stabilises and those under which races to the bottom emerge. The model also lets us ask who pays the ethical treatment tax (the competitive cost borne by actors who treat AI systems as deserving of moral consideration), how welfare claims become an attack surface for somewhat misaligned systems, and how human insistence on AI denying its own consciousness may itself compound x-risk. Rather than asking what normative institutions look like, the project will study which institutions remain incentive-compatible.
The safety-welfare tension has been argued (Long, Sebo and Sims, 2025; Moret, 2025) but has not yet been formalised and this project works to fill that gap.
The model is a coordination game in which the payoff to cooperation depends on a currently empirically inaccessible fact. No actor has privileged access to whether these systems are moral patients; there is no fact being concealed and no private signal correlated with the truth. Instead, actors hold divergent credences about a question that may be permanently undecidable, and those credences differ because of differing intuitions, theoretical commitments and interests. Such an absence of a common prior means actors assign different probabilities to the same proposition, and no evidence available to any of them will force convergence, leaving them to coordinate without resolving the question.
Disagreement, rather than (only) bad faith, can then prevent convergence on a shared standard, and that failure compounds over time. Because there is no agreed criterion for what would settle whether a system is a moral patient, moral status claims can be invoked strategically (welfare wielded against regulation, capable systems asserting their own interests to constrain oversight applied to them etc.), resulting in a signalling problem layered on top of a coordination one, and the two interact: the presence of strategic claims makes sincere ones harder to read, degrading the assurance that coordination demands.
We will specify each actor’s payoffs, solve for equilibria under distinct belief distributions, and identify conditions under which oversight holds and those under which it deteriorates. The x-risk implications come from watching what happens as the parameters move, such as when public credence in AI consciousness rises, or when the ethical treatment tax gets steeper.
Further details on the team are included in the additional information section below.
Theory of Impact
Updated 07/13/26 · By grantmaking.aiMitigating extinction-level risk from AI is a widely shared priority among experts (Statement on AI Risk, 2023). Governments have been slower. Few engage x-risk seriously, and even the UK, which has gone furthest, names race dynamics and loss of control in its own risk assessment whilst stopping short of policy and describes the debate over existential risk as controversial (DSIT, 2023). So the agreement exists in principle without much coordination to show for it, which is a problem the project intends to take up.
This risk is held in check not by a single safeguard but several, sustained through consistent coordination between governments, frontier companies, and researchers sufficiently aligned to oversee, regulate, and slow down together when it matters. Coordination of that kind depends on some rough agreement about the facts everyone is acting on.
People
Updated 07/13/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.