Measured harness ledgerPublic result
GPT-5.6 Terra

High-voxel-density Jungle Temple diorama — GPT-5.6 Terra Ultra

Create a complete inspectable Blender voxel diorama of a jungle temple at a fixed 0.05-unit voxel resolution through the Blender MCP workflow.

Ultra reasoningHeadline result
Workflow cost
$1.20
Wall-clock
16:43.2 wall-clock
Processed tokens
1.34M processed
Record state
partial_token_timing_and_artifact_ledger
Public summary

GPT-5.6 Terra Ultra partial_token_timing_and_artifact_ledger ledger: 16:43.2 wall-clock, 1.34M processed, and $1.20 API-equivalent estimate, not a subscription invoice.

Run identity and stack
  • Result ID: jungle-temple-gpt-5.6-terra-ultra
  • Technical model: gpt-5.6-terra
  • Provider: OpenAI Codex
  • Stack: OpenAI Codex
  • Stack: Technical model/configuration: gpt-5.6-terra
  • Stack: Blender MCP
  • Stack: 0.05-unit voxel grid
  • Stack: Inspectable .blend artifact
Cost basis
  • Requests with more than 272,000 input tokens use long-context pricing; all 19 supplied calls were short-context.
  • Separately priced tools and non-token services are excluded.
Primary artifact integrity
  • Kind: blender-scene
  • Path: artifacts/jungle-temple-gpt-5.6-terra-ultra/jungle_temple_diorama.blend
  • SHA-256: 022461495e61ff3cf63f108f69f3d9a6bba799f9b4f8d40cb9dc0a6515a26c54
Recorded caveats
  • Wall-clock is end-to-end latency, not model-only compute; it includes tool time and idle gaps between user turns.
  • Output tokens include hidden reasoning, visible prose/code, and tool-call JSON.
  • Cached input is deeply discounted, so total processed tokens overstate cost.
  • This is an API-equivalent estimate rather than the actual charge for a subscription-backed Codex session.
  • Separately priced tools and non-token services are excluded.
  • Cache-creation tokens are zero because the Codex transcript schema does not expose a cache-write field.
  • Subagent logs are excluded to avoid double-counting inherited parent context.
Visible evidence gaps
  • Blender and render environment
  • final capture metadata
  • blind-evaluation record
Builder test available

This result is part of Builder tests. Open them for the exact prompt and any released projects, RemakeBench Harness workflows and production skills. Public proof and known evidence gaps stay visible here.

  • Tech Review 001 · v1
  • Kimi K3 Launch 002 · v1
RemakeBenchResearch console