Measured harness ledgerPublic result
Qwen3.8 Max PreviewMacBook-class cinematic ad scene — Qwen3.8 Max Preview Mandatory thinking enabled · Attempt 3
Create one polished MacBook-class product-ad shot in Blender with a modeled device, legible industrial detail, intentional materials, lighting, camera movement, and a validator-ready scene.
Mandatory thinking enabled reasoningHeadline result
- Workflow cost
- $29.27
- Wall-clock
- 7h 00m 56.4s wall-clock
- Processed tokens
- 18.43M processed
- Record state
- partial_objective_pass_agent_nontermination
Public summary
Qwen3.8 Max Preview Mandatory thinking enabled partial_objective_pass_agent_nontermination ledger: 7h 00m 56.4s wall-clock, 18.43M processed, and $29.27 Qualified third-party NanoGPT API-list-price scenario; not an Alibaba Token Plan cash charge.
Run identity and stack
- Result ID: macbook-cinematic-blender-qwen3.8-max-preview-attempt-3-partial
- Technical model: qwen3.8-max-preview
- Provider: Alibaba ModelStudio Token Plan International
- Client: Qwen Code 0.20.1
- Attempt: 3
- Stack: Alibaba ModelStudio Token Plan International
- Stack: Qwen Code 0.20.1
- Stack: Technical model/configuration: qwen3.8-max-preview
- Stack: Attempt 3
- Stack: Blender MCP
- Stack: Cinematic product scene
- Stack: Harness v1 MacBook validator
Cost basis
- Qualified API-list-price equivalent.
Primary artifact integrity
- Kind: manifest-verified-artifact
- Path: artifacts/macbook-cinematic-blender-qwen3.8-max-preview-attempt-3-validator-pass-model-noncompletion/artifact/macbook_cinematic_final.blend
- SHA-256: 762c826ba6bee691c9bbbea6724122bd4d8f38dea0e28e7009e659d45e9b21e0
Validation evidence
- Result: 49/49 checks recorded
- Path: artifacts/macbook-cinematic-blender-qwen3.8-max-preview-attempt-3-validator-pass-model-noncompletion/artifact/validation_log.json
- SHA-256: 5ee3736d1e6e8b67e1588f90a167aa33bb7b44ddf3a55f5405383b85e4bce550
Recorded caveats
- This is a dual result: the core construction objective passed, but the agent workflow did not terminate normally. It is visible as PARTIAL, not as a completed or blind-vote-ready benchmark result.
- The 38m 10.5s number is time to the independent 49/49 objective gate; the full agent attempt lasted 7h 00m 56.4s including the quota pause.
- 61.42% of the API-equivalent cost was incurred after the objective gate had already passed.
- Output includes model-accounted thoughts; the thoughts figure is a subset of output and is not additive.
- The API-equivalent estimate is not Alibaba pricing, provider cost, or an itemized subscription charge.
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.
