Stateful benchmarks for long-horizon agents

Scores only mean something when harness, runtime, and task set are pinned.

Benchmarks that preserve trajectories, artifacts, and container metadata are more useful than leaderboard numbers alone. Long-horizon business agents fail on state and verification, not on one-shot Q&A.

The reproducibility contract matters: pin the task set, harness, image, provider, and model before comparing scores. Otherwise you are comparing stories.