Blog
Notes on artificial intelligence, research, and building.
articles
-
Corrigibility Games: AI-native games as interactive environments for measuring corrigibility
AI-native games as environments for studying the principal-agent correction relationship.
-
Autonomy Eval
Open evaluation environments for studying how model behavior and interface design affect human autonomy.
-
Mafia: On the Design of Social Deduction Agents and AI-Native Games
Introducing HOLY GRAIL, a role-adaptive reasoning agent evaluated in full multi-day games of Mafia.