Measured harness ledgerPublic result
Qwen3.8 Max Preview

JRPG boss battle — Qwen3.8 Max Preview Mandatory thinking enabled · Attempt 2

Build an interactive Three.js JRPG boss battle with supplied character assets, required combat beats, animation, sound controls, effects, and a validator-ready game loop.

Mandatory thinking enabled reasoningHeadline result
Workflow cost
$15.15
Wall-clock
58m 40.6s wall-clock
Processed tokens
9.64M processed
Record state
published_artifact_validation_runtime_ledger
Public summary

Qwen3.8 Max Preview Mandatory thinking enabled published_artifact_validation_runtime_ledger ledger: 58m 40.6s wall-clock, 9.64M processed, and $15.15 Qualified third-party NanoGPT API-list-price scenario for the measured token mix; not an Alibaba Token Plan cash charge.

Run identity and stack
  • Result ID: jrpg-boss-battle-qwen3.8-max-preview-attempt-2-qwen-code-0201
  • Technical model: qwen3.8-max-preview
  • Provider: Alibaba Cloud Model Studio Token Plan International
  • Client: Qwen Code 0.20.1
  • Attempt: 2
  • Stack: Alibaba Cloud Model Studio Token Plan International
  • Stack: Qwen Code 0.20.1
  • Stack: Technical model/configuration: qwen3.8-max-preview
  • Stack: Attempt 2
  • Stack: Three.js / Vite game
  • Stack: Supplied GLB fixture bundle
  • Stack: Harness v1 JRPG validator
Cost basis
  • Qualified API-list-price equivalent.
Primary artifact integrity
  • Kind: vite-threejs-jrpg-source-and-production-build-entry-point
  • Path: artifacts/jrpg-boss-battle-qwen3.8-max-preview-attempt-2-qwen-code-0201/source/src/main.js
  • SHA-256: 86d0c919290564a7478b6b3dc34e74dcab98395a89cefd71c477104c73353a96
Validation evidence
  • Result: PASS
Recorded caveats
  • Wall-clock is end-to-end workflow latency, not model-only compute time.
  • Output tokens include Qwen-accounted thoughts/reasoning, visible prose/code, and tool-related output; 123,945 thought tokens are a subset of 197,871 output tokens.
  • The independent 119.99 FPS sample is Apple M1 Pro / ANGLE Metal evidence, not a standardized cross-machine comparison.
  • The shared 275-Credit batch cannot be apportioned among Voxel Maps, the intervening activity, the failed JRPG attempt, and this successful retry.
  • No formal RemakeBench quality score or blind-voting result has been recorded.
Visible evidence gaps
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console