Measured harness ledgerPublic result
Kimi K3

Explorable space-flight game — Kimi K3 Max

Build a responsive browser space-flight game with flight controls, a coherent star-system environment, lighting, assets, and a playable game loop.

Max reasoningHeadline result
Workflow cost
$3.25
Wall-clock
1:04:11 wall-clock
Processed tokens
6.03M processed
Record state
partial_token_timing_and_source_ledger
Public summary

Kimi K3 Max partial_token_timing_and_source_ledger ledger: 1:04:11 wall-clock, 6.03M processed, and $3.25 API-equivalent usage accounting, not an itemized subscription cash charge.

Run identity and stack
  • Result ID: space-flight-game-kimi-k3-max
  • Technical model: kimi-k3
  • Provider: Kimi Code CLI / Moonshot AI
  • Stack: Kimi Code CLI / Moonshot AI
  • Stack: Technical model/configuration: kimi-k3
  • Stack: Three.js / Vite game
  • Stack: Generated spaceship assets
  • Stack: Harness v1 space-flight prompt
Cost basis
  • The Kimi Code CLI session used an Allegretto subscription. This is token-arithmetic equivalence, not a per-run cash charge.
Primary artifact integrity
  • Kind: interactive-threejs-space-flight-game-primary-component
  • Path: artifacts/space-flight-game-kimi-k3-max/source/src/main.js
  • SHA-256: a524d0a29174216290a5f38c2906c7636881c4b5617f4ad632830fe212533871
Recorded caveats
  • Wall-clock is end-to-end workflow latency, including tool execution, manual-approval waits, browser playtests, and idle time.
  • One user prompt produced 76 billable Kimi model requests.
  • Kimi's observed wire format does not expose a separate reasoning-token field; output tokens include provider-accounted reasoning, visible prose/code, and tool-call JSON.
  • Cache-read tokens are discounted, so total processed tokens overstate cost.
  • The raw metrics ledger records the session window but not output-file hashes; this result ledger supplies SHA-256 bindings to the archived source and runtime asset.
  • The supplied project README's approximate performance claim is not treated as measured benchmark FPS because no measurement receipt was supplied.
  • The /usage screenshot is contextual plan evidence only. Its rounded usage totals do not supersede the exact marker-window token receipt, and without before/after records it cannot measure per-run subscription quota consumption.
  • The metrics extraction is version-specific to Kimi Code CLI 0.26.0 until validated across more completed sessions.
Visible evidence gaps
  • browser and hardware environment
  • local FPS measurement receipt
  • final capture
  • blind-evaluation record
Builder test available

This result is part of a Builder test. Open it for the exact prompt and any released projects, RemakeBench Harness workflows and production skills. Public proof and known evidence gaps stay visible here.

  • Kimi K3 Launch 002 · v1
RemakeBenchResearch console