Measured harness ledgerPublic result
Kimi K3

JRPG boss battle — Kimi K3 Max

Build an interactive Three.js JRPG boss battle with supplied character assets, required combat beats, animation, sound controls, effects, and a validator-ready game loop.

Max reasoningHeadline result
Workflow cost
$2.86
Wall-clock
41:20.2 wall-clock
Processed tokens
5.77M processed
Record state
partial_token_timing_artifact_validation_ledger
Public summary

Kimi K3 Max partial_token_timing_artifact_validation_ledger ledger: 41:20.2 wall-clock, 5.77M processed, and $2.86 API-equivalent usage accounting, not an itemized subscription cash charge.

Run identity and stack
  • Result ID: jrpg-boss-battle-kimi-k3-max
  • Technical model: kimi-k3
  • Provider: Kimi Code CLI / Moonshot AI
  • Stack: Kimi Code CLI / Moonshot AI
  • Stack: Technical model/configuration: kimi-k3
  • Stack: Three.js / Vite game
  • Stack: Supplied GLB fixture bundle
  • Stack: Harness v1 JRPG validator
  • Stack: Requested tool profile: codex-cli-imagegen
Cost basis
  • The Kimi Code CLI session used an Allegretto subscription. This is token-arithmetic equivalence, not a per-run cash charge.
Primary artifact integrity
  • Kind: interactive-threejs-jrpg-boss-battle
  • Path: artifacts/jrpg-boss-battle-kimi-k3-max/source/src/main.js
  • SHA-256: 7d40c7d6c5c51e0d1cfa7d64693365b460ce0c24766c6cf8ef5933859c328f4b
Validation evidence
  • Result: PASS, 41/41 checks
  • Verification: The supplied scorecard was JSON-validated during archival and records a complete validator pass; the archived project was not freshly rerun during this archival step.
  • Path: artifacts/jrpg-boss-battle-kimi-k3-max/validation/scorecard.json
  • SHA-256: caef8efe6ab7c3b1e8a3d6fe30209c216e16fa114bba32a81d80612ddf386f97
  • Validator SHA-256: b571ff1895fd6a61d98ccb4a9f7ca4ecefc961abfc3f6b77e23742bddd921ab8
Recorded caveats
  • Wall-clock includes tools, installs, browser checks, approval waits, and idle time.
  • One user turn produced 62 billable Kimi model requests.
  • Kimi Code CLI 0.27.0 exposes no separate reasoning-token field; output includes provider-accounted reasoning, visible prose/code, and tool-call JSON.
  • Cache reads are discounted, so total processed tokens overstate cost.
  • Allegretto does not itemize a per-run cash charge; the $39 monthly subscription is not divided by this run.
  • The exact browser version and local hardware identity were not supplied.
  • The scorecard and capture are supplied result evidence and were not freshly rerun during archival.
  • The requested Codex-CLI image-generation profile is recorded from the task and source notes, not independently attested by a retained tool-call transcript.
  • The validator establishes functional completion but does not replace blind visual-quality evaluation.
Visible evidence gaps
  • browser and local-hardware identity
  • independent validator rerun
  • blind-evaluation record
Builder test available

This result is part of a Builder test. Open it for the exact prompt and any released projects, RemakeBench Harness workflows and production skills. Public proof and known evidence gaps stay visible here.

  • Kimi K3 Launch 002 · v1
RemakeBenchResearch console