Research
From Win Rate to Worst Case: Auditing Language-Model Policies in Imperfect-Information Games
Exact best-response audits reveal worst-case vulnerabilities that selected-opponent evaluations can miss.
Central Question
How vulnerable are elicited language-model policies to an optimal strategic opponent, even when they perform well against selected opponents?
Main Contribution
The paper instantiates an exact best-response audit for complete policies elicited within fixed finite-game schemas. Sequence-form scoring separates policy loss from representation loss and removes rollout noise.
The study audits 25 models across poker, simultaneous-move, negotiation, and bluffing testbeds. Every scored river and Leduc policy is exploitable, and 88% of policies in the primary draw are strictly security-dominated.
The paper also defines the stated-opponent security gap. This quantity measures inconsistency between an elicited policy and its separately elicited model of an opponent.
Citation
[1] J. Guo, “From Win Rate to Worst Case: Auditing Language-Model Policies in Imperfect-Information Games,” unpublished manuscript, Imperial College London, London, U.K., 2026.