Loading runs…
Explore the 1,200 episodes across eight agents and 50 task–policy pairs used for the paper's mechanism analysis. Each run shows the agent's calls, the monitor's decisions, and the scored outcome. This corpus is distinct from the complete ten-agent leaderboard evaluation.
Loading runs…