I tested 10 model/harness combinations on the same Three.js task
A developer benchmarked 10 model/harness combinations on a Three.js task; Qwen 3.8 27B on OpenCode scored 95.64% fastest at 8m48s.
The author ran an identical Three.js sci-fi hangar build prompt across 10 model/harness combinations and recorded score, tokens, durations, and tool errors. Qwen 3.8 27B x-high on OpenCode achieved 95.64% in 8m48s, the best fast result, while GLM 5.3 Flash Max on OpenCode scored highest at 96.89% in 20m28s. Other runs included GLM 5.3 Flash, Luna 5.6, SOL 5.6, and Astra 6.0 across Codex Open, OMP Open, OpenCode, DSH, and PTC harnesses, with scores ranging from 78.54% to 96.89%.
- Best score: GLM 5.3 Flash Max on OpenCode at 96.89% (20m28s)
- Fastest strong result: Qwen 3.8 27B x-high on OpenCode, 95.64% in 8m48s
- Worst result: GLM 5.3 Flash Max on OMP Open at 78.54%
- Methodology tracks tokens, reasoning tokens, tool errors, and durations
Full article315 words · extracted from alvins82.github.io · click to collapse
I've been testing a simple prompt with different model and harness combinations to work out which one produces best results. I do this in /goal mode.
Prompt: Build a single-page Three.js sci-fi hangar with hovering drones, animated warning lights, emissive runway strips, and subtle volumetric-style fog planes. Include drone formation toggle and cinematic camera path. Output one self-contained HTML file with inline JavaScript.
| GLM 5.3 Flash Max | Codex | Open | 9m 0.232s | 9.344s | 457,685 | 17,458 | 7,694 | 475,143 | $0.005061 | $0.005853 | $0.004364 | $0.015278 | 85.26% | 14 | 1 | Blocked | No |
| Luna 5.6 Max | Codex | Open | 9m 13.098s | 6.199s | 1,146,755 | 25,512 | 6,879 | 1,172,267 | $0.083075 | $0.106368 | $0.127560 | $0.317003 | 92.76% | 32 | 2 | Yes | Yes |
| SOL 5.6 Max | Codex | Open | 10m 48.765s | 8.460s | 1,069,163 | 28,278 | 7,515 | 1,097,441 | $0.057707 | $0.101146 | $0.141390 | $0.300243 | 94.60% | 21 | 5 | No | No |
| Astra 6.0 Max | Codex | Open | 37m 29.705s | 3.589s | 1,292,366 | 43,129 | 15,271 | 1,335,495 | $0.654860 | $1.226880 | $2.156450 | $4.038190 | 94.93% | 20 | 5 | Yes | No |
| GLM 5.3 Flash Max | OMP | Open | 30m 14.979s | 6.596s | 1,678,509 | 63,405 | — | 1,741,914 | $0.027013 | $0.019775 | $0.015851 | $0.062639 | 78.54% | 57 | 0 | Yes | Yes |
| Qwen 3.8 27B x-high | OMP | Open | 41m 25.836s | 15.275s | 3,407,451 | 71,935 | 49,131 | 3,479,386 | $0.187257 | $0.251736 | $0.215805 | $0.654798 | 86.92% | 89 | 0 | Yes | Yes |
| GLM 5.3 Flash Max | OpenCode | Open | 20m 28.948s | 6.284s | 4,305,447 | 50,468 | 33,414 | 4,355,915 | $0.010045 | $0.062573 | $0.012617 | $0.085234 | 96.89% | 67 | 0 | Yes | Yes |
| Qwen 3.8 27B x-high | OpenCode | Open | 8m 48.470s | 10.182s | 665,490 | 41,817 | 28,788 | 707,307 | $0.012184 | $0.054101 | $0.125451 | $0.191736 | 95.64% | 13 | 0 | Yes | Yes |
| Qwen 3.8 27B x-high | DSH / PTC | Open | 24m 32.929s | 7.724s | 1,012,499 | 89,894 | — | 1,102,393 | $0.035355 | $0.078907 | $0.269682 | $0.383944 | 91.69% | 23 | 5 | Yes | Yes |
| Qwen 3.8 27B x-high | DSH | Open | 18m 15.239s | 6.755s | 2,654,457 | 78,232 | — | 2,732,689 | $0.050424 | $0.215424 | $0.234696 | $0.500544 | 95.48% | 42 | 2 | No | No |
All GLM runs are labelled GLM 5.3 Flash Max; input tokens include cached input. Output tokens are the generated total, including reasoning; when a harness reports reasoning separately, the reasoning column shows that subset. The DSH adapter does not report a separate reasoning count. DSH durations sum active turn time across both turns, excluding the pause between turns. Tool errors are recorded failed tool events. A dash means unavailable or not reported.
Costs are OpenRouter-equivalent USD estimates using rates retrieved 2026-09-09: GLM 5.3 Flash, Qwen 3.8 27B, GPT-5.6 Luna, GPT-5.6 Sol, and GPT-6 Astra. Base rates in USD per million tokens (input / cache read / output) are GLM $0.075 / $0.015 / $0.250, Qwen $0.420 / $0.085 / $3.000, Luna $0.200 / $0.020 / $1.200, Sol $1.000 / $0.100 / $5.000, and Astra $10.000 / $1.000 / $50.000. Input cost is uncached input at the prompt rate; cache read cost uses the input-cache-read rate; output cost uses the completion rate and includes reasoning tokens. Cache-write tokens were zero in every session. The Codex row labelled Luna 5.6 Max records gpt-5.6-sol in its transcript and is therefore priced at the Sol rate. No individual request crossed OpenRouter's 272k-token long-context pricing threshold.
Text extracted automatically; images, tables and formatting may be missing. Original: https://alvins82.github.io/hangar-harness-model-tests/