Measured harness ledgerPublic result
Qwen3.8 Max Preview

Explorable space-flight game — Qwen3.8 Max Preview Unreported

Build a responsive browser space-flight game with flight controls, a coherent star-system environment, lighting, assets, and a playable game loop.

Unreported reasoningHeadline result
Workflow cost
$6.95
Wall-clock
38m 43.2s wall-clock
Processed tokens
4.44M processed
Record state
partial_token_timing_runtime_artifact_ledger_orientation_failure
Public summary

Qwen3.8 Max Preview Unreported partial_token_timing_runtime_artifact_ledger_orientation_failure ledger: 38m 43.2s wall-clock, 4.44M processed, and $6.95 Third-party NanoGPT API-list-price equivalent for the measured token mix; not an Alibaba Token Plan cash charge.

Run identity and stack
  • Result ID: space-flight-game-qwen3.8-max-preview
  • Technical model: qwen3.8-max-preview
  • Provider: Alibaba Cloud Model Studio
  • Client: Qwen Code 0.20.0
  • Stack: Alibaba Cloud Model Studio
  • Stack: Qwen Code 0.20.0
  • Stack: Technical model/configuration: qwen3.8-max-preview
  • Stack: Three.js / Vite game
  • Stack: Generated spaceship assets
  • Stack: Harness v1 space-flight prompt
  • Stack: Requested tool profile: base shell/file tools
Cost basis
  • NanoGPT published Qwen3.8 Max Preview rates of $1.50/M input and $5.00/M output. It did not publish a separate cache-read discount, so all measured input tokens use the same input rate. The actual run used Alibaba Token Plan Personal Lite; $6.95 is a third-party API-list-price equivalent, not a first-party Alibaba price, an itemized provider charge, or provider cost.
  • Qualified API-list-price equivalent.
Primary artifact integrity
  • Kind: threejs-space-flight-game-source-entry-point
  • Path: artifacts/space-flight-game-qwen3.8-max-preview/source/src/main.js
  • SHA-256: 1307a0deeda5ea1b9e2cf39977dbea4bc1262eb440f2618ecf1d4dc97cd5cea0
Recorded caveats
  • Wall-clock is end-to-end workflow latency, not model-only compute, and includes tool execution and waits within the recorded prompt window.
  • Output tokens include Qwen-accounted thoughts/reasoning, visible prose/code, and tool-related output; 40,727 thought tokens are a subset of the 82,279 output tokens.
  • The 120 FPS figure is a post-run HUD-reported measurement on an Apple M1 Pro / ANGLE Metal browser session, not a standardized GPU benchmark or model-supplied performance capture.
  • The $6.95 estimate uses NanoGPT's published third-party Qwen3.8 Max Preview API rates, not an Alibaba first-party pay-as-you-go price or a task-level subscription charge.
  • The task-isolated Credit estimate is provisional and after-only; its weekly meter cannot be allocated to this task alone.
  • The result has no quality score because the ship orientation is confirmed wrong and the model did not produce its own final capture.
  • No blind-evaluation record or standardized RTX render was supplied.
Visible evidence gaps
  • a fresh infrastructure-corrected retry with a model-accessible hardware-backed capture path
  • correct fighter orientation and improved gameplay framing
  • standardized local FPS measurement receipt
  • standardized RTX render
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console