Measured harness ledgerPublic result
GPT-5.6 LunaExplorable space-flight game — GPT-5.6 Luna xhigh
Build a responsive browser space-flight game with flight controls, a coherent star-system environment, lighting, assets, and a playable game loop.
xhigh reasoningHeadline result
- Workflow cost
- $0.33
- Wall-clock
- 10:55.9 wall-clock
- Processed tokens
- 1.20M processed
- Record state
- partial_token_timing_and_source_ledger
Public summary
GPT-5.6 Luna xhigh partial_token_timing_and_source_ledger ledger: 10:55.9 wall-clock, 1.20M processed, and $0.33 API-equivalent estimate, not a subscription invoice.
Run identity and stack
- Result ID: space-flight-game-gpt-5.6-luna-xhigh
- Technical model: gpt-5.6-luna
- Provider: OpenAI Codex
- Stack: OpenAI Codex
- Stack: Technical model/configuration: gpt-5.6-luna
- Stack: Three.js / Vite game
- Stack: Generated spaceship assets
- Stack: Harness v1 space-flight prompt
Cost basis
- Requests with more than 272,000 input tokens use long-context pricing; all 22 supplied calls were short-context.
- Separately priced tools and non-token services are excluded.
Primary artifact integrity
- Kind: threejs-space-flight-game-primary-source
- Path: artifacts/space-flight-game-gpt-5.6-luna-xhigh/source/src/main.js
- SHA-256: a5badbbaf62208d01644f8213313da39653d6021d9699e58617c6ced14d7e1f9
Recorded caveats
- Wall-clock is end-to-end latency, not model-only compute; it includes tool time and idle gaps between user turns.
- Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
- Cached input is deeply discounted, so total processed tokens overstate cost.
- This is an API-equivalent estimate rather than the actual charge for a subscription-backed Codex session.
- Separately priced tools and non-token services are excluded.
- Cache-creation tokens are zero because the Codex transcript schema does not expose a cache-write field.
- Subagent logs are excluded to avoid double-counting inherited parent context.
Visible evidence gaps
- browser and hardware environment
- local FPS
- final capture
- blind-evaluation record
Builder test available
This result is part of a Builder test. Open it for the exact prompt and any released projects, RemakeBench Harness workflows and production skills. Public proof and known evidence gaps stay visible here.
- Tech Review 001 · v1
