← All project work

Project

LLM exploitability audit

A Python toolkit for exact worst-case audits of language-model policies.

Python / 2026

Evaluation · Game theory

Read research View code

What it contains

The toolkit supports experiments that evaluate language-model policies against strategic counterplay.

Its structure keeps policy generation, game evaluation, and audit outputs separate.

Why it exists

The project makes worst-case strategic evaluation easier to inspect and reproduce.