Measured harness ledgerPublic result
Kimi K3

MacBook-class cinematic ad scene — Kimi K3 Max

Create one polished MacBook-class product-ad shot in Blender with a modeled device, legible industrial detail, intentional materials, lighting, camera movement, and a validator-ready scene.

Max reasoningHeadline result
Workflow cost
≥$10.85
Wall-clock
1:26:58.2 wall-clock
Processed tokens
≥26.65M processed
Record state
partial_token_timing_artifact_validation_ledger
Public summary

Kimi K3 Max partial_token_timing_artifact_validation_ledger ledger: 1:26:58.2 wall-clock, ≥26.65M processed, and ≥$10.85 API-equivalent lower bound over paired usage records, not an itemized subscription cash charge (lower bound).

Run identity and stack
  • Result ID: macbook-cinematic-blender-kimi-k3-max
  • Technical model: kimi-k3
  • Provider: Kimi Code CLI
  • Stack: Kimi Code CLI
  • Stack: Technical model/configuration: kimi-k3
  • Stack: Blender MCP
  • Stack: Cinematic product scene
  • Stack: Harness v1 MacBook validator
  • Stack: Requested tool profile: blender-mcp
Cost basis
  • One API request lacks a paired usage record, so the actual total may be higher. The Kimi Code session used an Allegretto subscription; this is token-arithmetic equivalence, not a per-run cash charge.
  • Lower-bound usage coverage: 122 of 123 API requests; the actual total may be higher.
Primary artifact integrity
  • Kind: blender-cinematic-product-ad-scene
  • Path: artifacts/macbook-cinematic-blender-kimi-k3-max/macbook_cinematic_final.blend
  • SHA-256: 197da22b6c5d777a24b37888b9175ea17bf8ce71de249eff6b681b66b4d79896
Validation evidence
  • Result: PASS, 49/49 checks
  • Verification: Supplied result evidence. The current archival environment has no Blender executable, so this log was JSON-validated and matched to the canonical task fixture but was not freshly rerun headlessly.
  • Path: artifacts/macbook-cinematic-blender-kimi-k3-max/validation_log.json
  • SHA-256: 6941ac0076ded31a4e1755c315374921ae15115034bd15b8b345584d009e6583
  • Validator SHA-256: 54591de8d4ef2953c6f7e2fe3e4f4b3d9407b953049cc858b1012afaf9424d51
Recorded caveats
  • Wall-clock is end-to-end workflow latency and includes tool execution, Blender renders, approval waits, and idle time.
  • One user prompt produced 123 Kimi model requests, but only 122 have reconciled usage records.
  • Kimi Code CLI 0.26.0 exposed no authoritative separate reasoning-token field; output tokens include provider-accounted reasoning, visible prose/code, and tool-call JSON.
  • Cache reads are discounted, so total processed tokens overstate cost.
  • The supplied scene validator reports PASS 49/49, but it was not independently rerun during archival because Blender is unavailable in this environment.
  • The private session ZIP named in the receipt was not present in the supplied project folder and is intentionally not published.
  • No /usage screenshot, render-hardware identity, or blind-evaluation record was supplied.
Visible evidence gaps
  • metrics reconciliation for the one unpaired API request
  • independent headless validator rerun
  • render-hardware identity
  • Kimi tool-use transcript or disclosure
  • blind-evaluation record
Builder test available

This result is part of a Builder test. Open it for the exact prompt and any released projects, RemakeBench Harness workflows and production skills. Public proof and known evidence gaps stay visible here.

  • Kimi K3 Launch 002 · v1
RemakeBenchResearch console