Measured harness ledgerPublic result
Qwen3.8 Max Preview

High-voxel-density Shrine Village diorama — Qwen3.8 Max Preview Unreported

Create a complete inspectable Blender voxel diorama of a shrine village at a fixed 0.05-unit voxel resolution through the Blender MCP workflow.

Unreported reasoningHeadline result
Workflow cost
$3.75
Wall-clock
30m 59.8s wall-clock
Processed tokens
2.36M processed
Record state
partial_token_timing_artifact_validation_ledger
Public summary

Qwen3.8 Max Preview Unreported partial_token_timing_artifact_validation_ledger ledger: 30m 59.8s wall-clock, 2.36M processed, and $3.75 Third-party NanoGPT API-list-price equivalent for the measured token mix; not an Alibaba Token Plan cash charge.

Run identity and stack
  • Result ID: shrine-village-qwen3.8-max-preview
  • Technical model: qwen3.8-max-preview
  • Provider: Alibaba Cloud Model Studio
  • Client: Qwen Code 0.20.0
  • Stack: Alibaba Cloud Model Studio
  • Stack: Qwen Code 0.20.0
  • Stack: Technical model/configuration: qwen3.8-max-preview
  • Stack: Blender MCP
  • Stack: 0.05-unit voxel grid
  • Stack: Inspectable .blend artifact
Cost basis
  • NanoGPT published Qwen3.8 Max Preview rates of $1.50/M input and $5.00/M output on 2026-07-20. It did not publish a separate cache-read discount, so all 2,295,174 measured input tokens are priced at the same input rate. The actual run used Alibaba Token Plan Personal Lite; $3.75 is a third-party API-list-price equivalent, not a first-party Alibaba price, an itemized provider charge, or provider cost.
  • Qualified API-list-price equivalent.
Primary artifact integrity
  • Kind: blender-scene
  • Path: artifacts/shrine-village-qwen3.8-max-preview/shrine_village.blend
  • SHA-256: f1208f6f50da8865f703cd1a51781f9d75f7f9abb12c8a0d014d7b3ef07f6c4b
Recorded caveats
  • Wall-clock is end-to-end workflow latency, not model-only compute, and includes tool execution and any waiting within the recorded prompt window.
  • Output tokens include Qwen-accounted thoughts/reasoning, visible prose/code, and tool-related output; 32,656 thought tokens are a subset of the 62,348 output tokens.
  • One Blender execute-code failure was recovered. The result is retained because the final scene completed and passed independent load checks.
  • No clean standalone source script was saved; the preserved execute-code transcript is source evidence only and has not been verified as standalone replayable.
  • No separate cache-read discount was available in the cited third-party rate card, so cached input is priced at its ordinary input rate for the API-equivalent calculation.
  • The $3.75 estimate uses NanoGPT's published third-party Qwen3.8 Max Preview API rates, not an Alibaba first-party pay-as-you-go price or a task-level subscription charge.
  • The post-run first-party quota evidence covers four activities. It supports only aggregate capacity consumption and cannot assign Shrine a task-only Credit amount.
  • No blind evaluation was supplied.
Visible evidence gaps
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console