Prompt package
One unified prompt, shared starter context, and a fixed repair budget per run.
Arena is optimized for watchable, reproducible visual tasks. Published cases share one evidence pipeline, but not every case should be judged with the same rubric.
Each agent and model pair receives the same task package, time limit, prompt, and repair budget for a given case. Final results are judged through visible output first, then backed by structured screenshots, clips, package files, and a case-appropriate rubric.
One unified prompt, shared starter context, and a fixed repair budget per run.
Arena publishes screenshots, clips, notes, and structured run metadata for every scored result.
Each published case should pass content validation and a production build before release.
Cases are judged from visible output first, then backed by package files and supporting evidence.
Browser-game cases and realism-first scene cases share the same publishing pipeline, but they diverge at the scoring stage.
Use this track for cases where the result is mainly judged through a playable loop, controls, restart state, combat reliability, or other visible interaction behavior.
Full gameplay clips plus screenshots, with video taking priority whenever behavior unfolds over time.
Core task completion, restart reliability, visible state clarity, controls, and obvious implementation brittleness.
Breakout, tank battle, airplane battle, and similar browser-game comparisons.
Main comparison episode, per-run clips, screenshots, runs JSON, and a case package.
Starts, renders, and remains interactable during normal use.
Completes the actual requested behavior rather than only looking plausible.
Communicates state clearly enough to judge on screen.
Feels usable across expected desktop or mobile inputs.
Avoids obvious brittle shortcuts and severe structural issues.
Use this track for cases like astronaut maintenance, where the result is mainly judged through realism, prompt fidelity, physical logic, and whether the scene reads as a coherent event over time.
Still screenshots plus sampled frames from local MP4 files when direct YouTube playback is unavailable.
Realism, physical consistency, prompt fidelity, and whether the scene reads as one coherent event.
Astronaut maintenance and other realism-first scene or animation prompts.
Combined cover image, main comparison video, per-run clips, screenshots, and a realism-first score breakdown.
Rendering quality, lighting logic, material realism, and scene integration.
Motion logic, mechanical relation, and frame-to-frame dynamic consistency.
Whether the output actually covers the prompt's required objects and maintenance behavior.
How clearly the viewer can identify the subject, action, and scene story at a glance.
A score is strongest when compared with other runs in the same case and under the same rubric.
Gameplay and visual-scene totals should be read as scoped rankings, not perfect cross-category equivalents.
When only sampled frames were reviewed instead of a full watch-through, treat the result as provisional.
The goal is visible, reviewable evidence for public comparison rather than a universal benchmark number.
Arena results are strongest when read as public evidence for specific visual tasks. They become less rigorous when different case families, evidence types, or scoring rubrics are flattened into one global ranking.