Measured harness ledgerPublic result
Qwen3.8 Max PreviewExplorable space-flight game — Qwen3.8 Max Preview Unreported
Build a responsive browser space-flight game with flight controls, a coherent star-system environment, lighting, assets, and a playable game loop.
Unreported reasoningHeadline result
- Workflow cost
- $6.95
- Wall-clock
- 38m 43.2s wall-clock
- Processed tokens
- 4.44M processed
- Record state
- partial_token_timing_runtime_artifact_ledger_orientation_failure
Public summary
Qwen3.8 Max Preview Unreported partial_token_timing_runtime_artifact_ledger_orientation_failure ledger: 38m 43.2s wall-clock, 4.44M processed, and $6.95 Third-party NanoGPT API-list-price equivalent for the measured token mix; not an Alibaba Token Plan cash charge.
Run identity and stack
- Result ID: space-flight-game-qwen3.8-max-preview
- Technical model: qwen3.8-max-preview
- Provider: Alibaba Cloud Model Studio
- Client: Qwen Code 0.20.0
- Stack: Alibaba Cloud Model Studio
- Stack: Qwen Code 0.20.0
- Stack: Technical model/configuration: qwen3.8-max-preview
- Stack: Three.js / Vite game
- Stack: Generated spaceship assets
- Stack: Harness v1 space-flight prompt
- Stack: Requested tool profile: base shell/file tools
Cost basis
- NanoGPT published Qwen3.8 Max Preview rates of $1.50/M input and $5.00/M output. It did not publish a separate cache-read discount, so all measured input tokens use the same input rate. The actual run used Alibaba Token Plan Personal Lite; $6.95 is a third-party API-list-price equivalent, not a first-party Alibaba price, an itemized provider charge, or provider cost.
- Qualified API-list-price equivalent.
Primary artifact integrity
- Kind: threejs-space-flight-game-source-entry-point
- Path: artifacts/space-flight-game-qwen3.8-max-preview/source/src/main.js
- SHA-256: 1307a0deeda5ea1b9e2cf39977dbea4bc1262eb440f2618ecf1d4dc97cd5cea0
Recorded caveats
- Wall-clock is end-to-end workflow latency, not model-only compute, and includes tool execution and waits within the recorded prompt window.
- Output tokens include Qwen-accounted thoughts/reasoning, visible prose/code, and tool-related output; 40,727 thought tokens are a subset of the 82,279 output tokens.
- The 120 FPS figure is a post-run HUD-reported measurement on an Apple M1 Pro / ANGLE Metal browser session, not a standardized GPU benchmark or model-supplied performance capture.
- The $6.95 estimate uses NanoGPT's published third-party Qwen3.8 Max Preview API rates, not an Alibaba first-party pay-as-you-go price or a task-level subscription charge.
- The task-isolated Credit estimate is provisional and after-only; its weekly meter cannot be allocated to this task alone.
- The result has no quality score because the ship orientation is confirmed wrong and the model did not produce its own final capture.
- No blind-evaluation record or standardized RTX render was supplied.
Visible evidence gaps
- a fresh infrastructure-corrected retry with a model-accessible hardware-backed capture path
- correct fighter orientation and improved gameplay framing
- standardized local FPS measurement receipt
- standardized RTX render
- blind-evaluation record
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.
