Versioned benchmark collection

JRPG Boss Battle v1

Harness v1 · 5 measured runs across 1 target. Each entry is a single recorded attempt; repeat-run reliability is not established.

Knight facing an armored boss in the JRPG Boss Battle v1 collection
Knight facing an armored boss in the JRPG Boss Battle v1 collection
5Measured runs
1Frozen targets
5Model families
5Configurations
Collection scope

Recorded evidence, without a synthetic score.

Runs share the same frozen target and prompt within each test, with disclosed provider-specific stacks. Costs are recorded estimates as labelled in each ledger, not subscription invoices.

Targets

  • JRPG boss battle

Model coverage

  • Claude Fable 5 · Max
  • GPT-5.6 Sol · Ultra
  • Kimi K3 · Max
  • Claude Opus 5 · Max
  • Qwen3.8 Max Preview · Mandatory thinking enabled

Evidence available

  • Five measured run ledgers
  • Public recorded outputs
  • Cost, token, and workflow timing

Still missing

  • Repeat-run reliability evidence
  • Minimum blind-vote sample
Public receipts

5 measured run ledgers

Open any row for its recorded workflow cost, timing, token usage, stack disclosure, artifact integrity, and explicit evidence gaps.