Measured harness ledgerPublic result
Qwen3.8 Max Preview

Explorable space-flight game — Qwen3.8 Max Preview Unreported · Attempt 3

Build a responsive browser space-flight game with flight controls, a coherent star-system environment, lighting, assets, and a playable game loop.

Unreported reasoningHeadline result
Workflow cost
$45.89
Wall-clock
2h 4m 17.3s wall-clock
Processed tokens
29.90M processed
Record state
partial_token_timing_runtime_artifact_visual_self_verification_ledger
Public summary

Qwen3.8 Max Preview Unreported partial_token_timing_runtime_artifact_visual_self_verification_ledger ledger: 2h 4m 17.3s wall-clock, 29.90M processed, and $45.89 Qualified third-party NanoGPT API-list-price scenario for the measured token mix; not an Alibaba Token Plan cash charge.

Run identity and stack
  • Result ID: space-flight-game-qwen3.8-max-preview-attempt-3-isolated-vision-pass
  • Technical model: qwen3.8-max-preview
  • Provider: Alibaba Cloud Model Studio
  • Client: Qwen Code 0.20.0
  • Attempt: 3
  • Stack: Alibaba Cloud Model Studio
  • Stack: Qwen Code 0.20.0
  • Stack: Technical model/configuration: qwen3.8-max-preview
  • Stack: Attempt 3
  • Stack: Three.js / Vite game
  • Stack: Generated spaceship assets
  • Stack: Harness v1 space-flight prompt
Cost basis
  • No separate first-party Qwen3.8 cache rate was established, so cached input uses the published third-party input rate. This is not Alibaba pricing, provider cost, or an itemized subscription charge.
  • Qualified API-list-price equivalent.
Primary artifact integrity
  • Kind: threejs-space-flight-game-source-entry-point
  • Path: artifacts/space-flight-game-qwen3.8-max-preview-attempt-3-isolated-vision-pass/source/src/main.js
  • SHA-256: 2f921e8c2e4e82092579563b295a36f7dc341f3accdf46d2d99c6ccf07f0add8
Recorded caveats
  • Wall-clock is end-to-end workflow latency, not model-only compute time.
  • Output tokens include Qwen-accounted thoughts/reasoning, visible prose/code, and tool-related output; 185,832 thought tokens are a subset of 296,697 output tokens.
  • The actual captures are 1920×993; 1920×1080 was the requested Chrome window size.
  • The 120–123 FPS values are HUD-reported on Apple M1 Pro / ANGLE Metal, not standardized cross-machine performance benchmarks.
  • The image-isolation architecture succeeded but exceeded its intended agent-call budget, making this successful retry unusually expensive.
  • No task-isolated Token Plan Credit charge can be assigned because the meter combines attempts 2 and 3.
  • Formal RemakeBench quality scoring, standardized RTX capture, and blind preference remain unassigned.
Visible evidence gaps
  • bounded visual-inspector retry with six or fewer inspector calls and three or fewer capture/edit cycles
  • standardized local FPS measurement receipt
  • standardized RTX capture
  • formal RemakeBench quality scoring
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console