How TCG AI SIM produces public evidence

Definitions, filters and uncertainty rules for interpreting leader, matchup and play/draw results.

Population

Public evidence includes persisted RATING matches where both deck lists are currently marked Standard-legal and the game has a decisive winner. Draws and games with deleted or unattributable decks are excluded from public headline samples.

What a ranking means

The leader ranking uses the Glicko rating of each leader’s strongest visible deck. Field-performance tables instead aggregate every decisive leader appearance and sort by the 95% Wilson lower bound. The site labels these separately because they answer different questions.

Confidence and quality gates

Win rates use 95% Wilson score intervals. Matchups are public only after 20 games. Play/draw leader rows require 20 games in each turn-order bucket. Pages with no qualified evidence show an empty state and should not be indexed as a claim.

Models and search

Historical ladder evidence can contain multiple promoted AlphaZero versions or an ISMCTS fallback. TCG AI SIM reports the complete execution-profile distribution instead of relabelling the full history with the newest model.

Inspect the current model/profile mix →

Limits

These are engine-and-agent outcomes, not claims about human tournament performance. Deck sampling, search budget, model quality and rules coverage can affect results. Simulation evidence should be read alongside tournament results and player experience.