Versioned benchmark collection
JRPG Boss Battle v1
Harness v1 · 5 measured runs across 1 target. Each entry is a single recorded attempt; repeat-run reliability is not established.

5Measured runs
1Frozen targets
5Model families
5Configurations
Collection scope
Recorded evidence, without a synthetic score.
Runs share the same frozen target and prompt within each test, with disclosed provider-specific stacks. Costs are recorded estimates as labelled in each ledger, not subscription invoices.
Targets
- JRPG boss battle
Model coverage
- Claude Fable 5 · Max
- GPT-5.6 Sol · Ultra
- Kimi K3 · Max
- Claude Opus 5 · Max
- Qwen3.8 Max Preview · Mandatory thinking enabled
Evidence available
- Five measured run ledgers
- Public recorded outputs
- Cost, token, and workflow timing
Still missing
- Repeat-run reliability evidence
- Minimum blind-vote sample
Public receipts
5 measured run ledgers
Open any row for its recorded workflow cost, timing, token usage, stack disclosure, artifact integrity, and explicit evidence gaps.
01Claude Fable 5Max reasoning · JRPG boss battle$123.42 API-equivalent estimate with exact cache-write TTL accounting, not a subscription invoicerecorded estimate50:48.1 wall-clockEnd-to-end workflowOpen 02GPT-5.6 SolUltra reasoning · JRPG boss battle$9.98 API-equivalent estimate, not a subscription invoicerecorded estimate25:51.8 wall-clockEnd-to-end workflowOpen 03Kimi K3Max reasoning · JRPG boss battle$2.86 API-equivalent usage accounting, not an itemized subscription cash chargerecorded estimate41:20.2 wall-clockEnd-to-end workflowOpen 04Claude Opus 5Max reasoning · JRPG boss battle$48.95 First-party Claude API list-price equivalent, not a Claude Code subscription invoicerecorded estimate1h 15m 49.9s wall-clockEnd-to-end workflowOpen 05Qwen3.8 Max PreviewMandatory thinking enabled reasoning · JRPG boss battle$15.15 Qualified third-party NanoGPT API-list-price scenario for the measured token mix; not an Alibaba Token Plan cash chargerecorded estimate58m 40.6s wall-clockEnd-to-end workflowOpen
