Work · Method
How an adversary agent gets built
Engagement · Red Kraken · CSIS Futures Lab
Closed-door congressional wargame · 21 July 2026
Role: built the adversary agent and ran it live during play, under a CSIS Futures Lab affiliation
An exercise is only as good as its opposition. If the other side is a script, the players learn the script. What follows is how an adversary agent gets built: what it is, why you'd want one, and where the boundaries sit.
What a client publishes is theirs to publish. What they don't stays unpublished here.
The problem an exercise has to solve
A wargame exists to answer a question that can't be field-tested. It puts people in a room, gives both sides real choices, and makes them live with the consequences. For that to teach anything, the adversary has to be credible: it has to press where the plan is weak, respond to what the opponent actually does, and behave like something with intentions of its own rather than a script with a schedule.
The old answer is a human red team. It works, and it has hard limits. It costs money to convene. It runs when everyone's calendar allows rather than when the question comes up. And its reasoning walks out of the room when the session ends.
An agent is better on all three counts. Run it whenever the question arises, as many times as the question needs. Retarget the corpus and the advisory lenses to a different adversary or a different scenario without rebuilding anything. Every move is logged with the reasoning attached, so a finding can still be argued with in the debrief weeks later.
The architecture
The agent reasons from a curated library (published doctrine, strategy literature, documented cyber activity) rather than from a language model's general knowledge. In Red Kraken's case that meant Chinese military doctrine and more than 150 documented PRC cyber operations, with a separate intelligence agent to predict U.S. moves. That's a defensibility decision before it's a technical one. A move that traces to a document can be argued with in the debrief; "the model thought so" can't. Building the library is most of the work, and it is where the domain judgment lives.
On each turn, several advisor agents read the state of the game in parallel, each through a different lens, and each proposes a course of action with its reasoning attached. A separate decision agent weighs the proposals, resolves the disagreement between them, and commits to a move.
Why several advisors, and why is the decision its own step? Because one agent asked for "the best move" hands back a single answer with the argument already smoothed away. Distinct advisory lenses surface genuine disagreement about what the adversary should do, and keeping the decision separate means that disagreement is resolved somewhere visible. The argument is preserved in the log instead of averaged out of existence.
The game state persists across turns, and between turns the opponent's actual allocation goes back in. The agent sees where they concentrated and where they left gaps, so the next move responds to what really happened, not to a branch written in advance.
Opaque in play, transparent after
During play, the agent's intent is opaque, the way a real adversary's is. The opponent argues about what it is doing and why, and that argument is the exercise working as intended.
Afterwards, they get what no set of after-action notes fully reconstructs: the complete decision trail. Every advisor proposal, every resolution, every committed move, with the sources behind them. The debrief can trace exactly why pressure landed where it did and check the reasoning against the library it came from. Most of the learning sits in the gap between those two states: an opponent you couldn't read while playing, whose reasoning you can read completely once it's over.
The scope discipline
The agent is an analytical and educational tool, and its boundaries are part of the design rather than an apology for it. It runs on unclassified, published material. It isn't connected to any operational system. It doesn't identify vulnerabilities in real infrastructure, and it doesn't predict specific attacks.
Those constraints are what make the tool usable in the room it is built for. Participants can argue with it freely, and the people who commissioned it can say precisely what it is and what it isn't. A tool that can't cross those lines is one you can bring to the people whose job is to guard them.
Coverage
Third-party coverage of the exercise.
-
CSIS · 28 July 2026
Red Kraken: The Coming Age of Agentic Cyber Strategy · Benjamin Jensen and Emily Harding -
CSIS event · 21 July 2026
Opening Remarks: Congressional Wargame on AI-Enabled Cyber Threats to U.S. Critical Infrastructure -
House Committee on Homeland Security · 24 July 2026
House Homeland, China Select Members Participate in War Game Exercise on AI-Enabled Cyber Threats to U.S. Critical Infrastructure -
Defense One · July 2026
Lawmakers get taste of AI-enabled cyberattacks in China-Taiwan war game
Further reading: It Is Time to Democratize Wargaming Using Generative AI (CSIS), the case for the category this work sits in, useful if you're deciding whether an exercise like this is worth commissioning.
If you're weighing an exercise that needs an adversary, or a simulation that needs credible agents inside it, start with an email: maxjensen@scintaralabs.com