Versioned benchmark collection

MacBook Cinematic v1

Harness v1 · 5 measured runs across 1 target. Each entry is a single recorded attempt; repeat-run reliability is not established.

Silver laptop illuminated in the MacBook Cinematic v1 collection
Silver laptop illuminated in the MacBook Cinematic v1 collection
5Measured runs
1Frozen targets
5Model families
5Configurations
Collection scope

Recorded evidence, without a synthetic score.

Runs share the same frozen target and prompt within each test, with disclosed provider-specific stacks. Costs are recorded estimates as labelled in each ledger, not subscription invoices.

Targets

  • MacBook-class cinematic ad scene

Model coverage

  • Claude Fable 5 · Max
  • GPT-5.6 Sol · xhigh
  • Kimi K3 · Max
  • Claude Opus 5 · Max
  • Qwen3.8 Max Preview · Mandatory thinking enabled

Evidence available

  • Five measured run ledgers
  • Public recorded outputs
  • Cost, token, and workflow timing

Still missing

  • One reconciled Kimi API request
  • Repeat-run reliability evidence
  • Minimum blind-vote sample
Public receipts

5 measured run ledgers

Open any row for its recorded workflow cost, timing, token usage, stack disclosure, artifact integrity, and explicit evidence gaps.