Mark Riedl
Glenn Matlin
August 8, 2026
Publication
No LM-enabled wargame should inform planning, doctrine, policy, or crisis response without an auditable safety case; the proper use of open-ended wargames today is to stress-test decision-influencing LM agents.
Accepted
August 2026
Authors
Mark Riedl†, Glenn Matlin† († equal contribution)
Venue
Social Simulation with LLMs: Fidelity in Applications Workshop at COLM 2026
Open-ended strategic wargames are high-stakes LM-based social simulations: they model adversaries, institutions, escalation, plan brittleness, doctrine, and crisis response. Language models (LMs) are attractive because they can play agents, generate scenario branches, adjudicate ambiguous actions, and summarize lessons, but the same affordances make open-ended roles dangerous: model language determines both what an actor attempts and what becomes simulated reality. This position paper argues that no LM-enabled wargame should inform planning, doctrine, policy, or crisis response without an auditable safety case, and that the proper use of open-ended wargames today is to stress-test decision-influencing LM agents. We identify five failure modes: decision laundering, adjudication opacity, role collapse, escalation-through-adjudication, and failure of strategic imagination. Ordinary benchmarks cannot establish safety for these settings. Wargames can expose failures as stress tests; they are not themselves safety cases for consequential use.
Continue exploring
Return to the publication archive or step back to the broader research agenda.