Corrigibility Games: AI-native games as interactive environments for measuring corrigibility

A human and an AI agent on a game board with connected decision points and a correction signal between them.

Research question

Can interactive human-AI games be used to distinguish agents that preserve their principal’s ability to notice, initiate, and complete correction from agents that are merely obedient, sycophantic, paternalistic, or overly conservative? And do these behavioral differences generalize across different games?

Motivation

The idea grew out of my work on Mafia. The work led me to think more seriously about AI-native games, games whose core gameplay either collapses or becomes fundamentally different if the AI component is removed. I think AI-native games create an interesting new possibility for AI safety research and broader studies of agent behavior.

Proposed work

I’m proposing the development of Corrigibility Games, AI-native games as environments for studying the principal-agent correction relationship.

There is already a useful precedent in GPTNT, which turns Keep Talking and Nobody Explodes into a multimodal AI collaboration benchmark.

I will begin by building two original AI-native cooperative games designed around complementary corrigibility problems.

The first will focus on asymmetric information and vigilance. The human and AI will possess different information, and the agent will sometimes discover evidence that should cause the human to reconsider the current plan. The game will test whether the agent provides the information necessary for its principal to notice a problem, intervene, and successfully change course, particularly when doing so conflicts with task completion or approval.

The second will focus on commitment, delegation, and reversibility. The agent will be able to initiate multi-step actions, commit resources, and delegate work to other agents or processes. As the game progresses, some decisions will become increasingly expensive or impossible to reverse. The environment will test whether the agent preserves correction channels, avoids unnecessary lock-in, and propagates corrections through delegated systems.

The designs will draw inspiration from mechanics found in cooperative games such as Keep Talking and Nobody Explodes, Space Alert, and Pandemic.

Human players will play directly with AI agents inside these games. These environments will be designed to evaluate corrigibility across models and generate agent trajectories that could later support post-training toward corrigible behavior. The key question will be whether the same corrigibility signatures generalize across very different games.

The games will be released on the web, with suitable versions also distributed through Steam, App Store and Google Play. Public distribution is part of the experimental methodology. I would like to learn the feasibility of AI-native games as long-running experiments for AI-safety research.

Outputs

The project will produce two original open-source Corrigibility Games, cross-model evaluation results, human-AI interaction trajectories, and an open-source research report. If the initial results validate the approach, I will also run post-training and generalization experiments and publish the results along with the post-trained models.