Methodology

How Arena compares visual coding agents

Arena is optimized for watchable, reproducible visual tasks. Published cases share one evidence pipeline, but not every case should be judged with the same rubric.

Run protocol

Same task package, same evidence discipline

Each agent and model pair receives the same task package, time limit, prompt, and repair budget for a given case. Final results are judged through visible output first, then backed by structured screenshots, clips, package files, and a case-appropriate rubric.

Prompt package

One unified prompt, shared starter context, and a fixed repair budget per run.

Evidence pack

Arena publishes screenshots, clips, notes, and structured run metadata for every scored result.

Validation

Each published case should pass content validation and a production build before release.

Review mode

Cases are judged from visible output first, then backed by package files and supporting evidence.

Workflow modes

Arena now uses two scoring tracks

Browser-game cases and realism-first scene cases share the same publishing pipeline, but they diverge at the scoring stage.

Gameplay case workflow

Interactive browser tasks

Use this track for cases where the result is mainly judged through a playable loop, controls, restart state, combat reliability, or other visible interaction behavior.

Primary evidence

Full gameplay clips plus screenshots, with video taking priority whenever behavior unfolds over time.

What gets scored

Core task completion, restart reliability, visible state clarity, controls, and obvious implementation brittleness.

Typical cases

Breakout, tank battle, airplane battle, and similar browser-game comparisons.

Publishing target

Main comparison episode, per-run clips, screenshots, runs JSON, and a case package.

20

Runs without fatal errors

Starts, renders, and remains interactable during normal use.

30

Core functionality

Completes the actual requested behavior rather than only looking plausible.

20

Visual completion

Communicates state clearly enough to judge on screen.

20

Interaction quality

Feels usable across expected desktop or mobile inputs.

10

Code sanity

Avoids obvious brittle shortcuts and severe structural issues.

Visual scene workflow

Realism-first scene review

Use this track for cases like astronaut maintenance, where the result is mainly judged through realism, prompt fidelity, physical logic, and whether the scene reads as a coherent event over time.

Primary evidence

Still screenshots plus sampled frames from local MP4 files when direct YouTube playback is unavailable.

What gets scored

Realism, physical consistency, prompt fidelity, and whether the scene reads as one coherent event.

Typical cases

Astronaut maintenance and other realism-first scene or animation prompts.

Publishing target

Combined cover image, main comparison video, per-run clips, screenshots, and a realism-first score breakdown.

40

Visual authenticity (VQ)

Rendering quality, lighting logic, material realism, and scene integration.

25

Physics and consistency (PC)

Motion logic, mechanical relation, and frame-to-frame dynamic consistency.

20

Semantic restoration (SR)

Whether the output actually covers the prompt's required objects and maintenance behavior.

15

Composition and readability (CR)

How clearly the viewer can identify the subject, action, and scene story at a glance.

Reading scores

Compare within workflow before comparing across the whole site

Case-local first

A score is strongest when compared with other runs in the same case and under the same rubric.

Workflow scope

Gameplay and visual-scene totals should be read as scoped rankings, not perfect cross-category equivalents.

Evidence status

When only sampled frames were reviewed instead of a full watch-through, treat the result as provisional.

Public usefulness

The goal is visible, reviewable evidence for public comparison rather than a universal benchmark number.

Limitations

Not a universal coding benchmark

Arena results are strongest when read as public evidence for specific visual tasks. They become less rigorous when different case families, evidence types, or scoring rubrics are flattened into one global ranking.