arXiv cs.AIOctober 2, 2026
Mean field games as a tool for AI safety: a worked example from the July 2026 Hugging Face incident
Excerpt
arXiv:2610.00902v1 Announce Type: cross Abstract: One way to make AI systems safe is to shape what the system is: its objective and dispositions. We take a complementary route: treat the agents' characteristics as partly unknown and ask what structure of interaction ensures that bad collective outcomes are not equilibria. Mean field games suit this when many interchangeable agents are coupled through an aggregate. We introduce a program for using them in AI safety and carry one example through e