ZeroHour
Hacker News · AIpublished ()ingested alvins82

I tested 10 model/harness combinations on the same Three.js task

infoAI researchimportance 30
AI summary · glm-5.3-flash

A developer benchmarked 10 model/harness combinations on a Three.js task; Qwen 3.8 27B on OpenCode scored 95.64% fastest at 8m48s.

The author ran an identical Three.js sci-fi hangar build prompt across 10 model/harness combinations and recorded score, tokens, durations, and tool errors. Qwen 3.8 27B x-high on OpenCode achieved 95.64% in 8m48s, the best fast result, while GLM 5.3 Flash Max on OpenCode scored highest at 96.89% in 20m28s. Other runs included GLM 5.3 Flash, Luna 5.6, SOL 5.6, and Astra 6.0 across Codex Open, OMP Open, OpenCode, DSH, and PTC harnesses, with scores ranging from 78.54% to 96.89%.

  • Best score: GLM 5.3 Flash Max on OpenCode at 96.89% (20m28s)
  • Fastest strong result: Qwen 3.8 27B x-high on OpenCode, 95.64% in 8m48s
  • Worst result: GLM 5.3 Flash Max on OMP Open at 78.54%
  • Methodology tracks tokens, reasoning tokens, tool errors, and durations
Full article315 words · extracted from alvins82.github.io · click to collapse

I've been testing a simple prompt with different model and harness combinations to work out which one produces best results. I do this in /goal mode.

Prompt: Build a single-page Three.js sci-fi hangar with hovering drones, animated warning lights, emissive runway strips, and subtle volumetric-style fog planes. Include drone formation toggle and cinematic camera path. Output one self-contained HTML file with inline JavaScript.

GLM 5.3 Flash Max Codex Open 9m 0.232s 9.344s 457,685 17,458 7,694 475,143 $0.005061 $0.005853 $0.004364 $0.015278 85.26% 14 1 Blocked No
Luna 5.6 MaxCodexOpen9m 13.098s6.199s1,146,75525,5126,8791,172,267$0.083075$0.106368$0.127560$0.31700392.76%322YesYes
SOL 5.6 MaxCodexOpen10m 48.765s8.460s1,069,16328,2787,5151,097,441$0.057707$0.101146$0.141390$0.30024394.60%215NoNo
Astra 6.0 MaxCodexOpen37m 29.705s3.589s1,292,36643,12915,2711,335,495$0.654860$1.226880$2.156450$4.03819094.93%205YesNo
GLM 5.3 Flash MaxOMPOpen30m 14.979s6.596s1,678,50963,4051,741,914$0.027013$0.019775$0.015851$0.06263978.54%570YesYes
Qwen 3.8 27B x-highOMPOpen41m 25.836s15.275s3,407,45171,93549,1313,479,386$0.187257$0.251736$0.215805$0.65479886.92%890YesYes
GLM 5.3 Flash MaxOpenCodeOpen20m 28.948s6.284s4,305,44750,46833,4144,355,915$0.010045$0.062573$0.012617$0.08523496.89%670YesYes
Qwen 3.8 27B x-highOpenCodeOpen8m 48.470s10.182s665,49041,81728,788707,307$0.012184$0.054101$0.125451$0.19173695.64%130YesYes
Qwen 3.8 27B x-highDSH / PTCOpen24m 32.929s7.724s1,012,49989,8941,102,393$0.035355$0.078907$0.269682$0.38394491.69%235YesYes
Qwen 3.8 27B x-highDSHOpen18m 15.239s6.755s2,654,45778,2322,732,689$0.050424$0.215424$0.234696$0.50054495.48%422NoNo

All GLM runs are labelled GLM 5.3 Flash Max; input tokens include cached input. Output tokens are the generated total, including reasoning; when a harness reports reasoning separately, the reasoning column shows that subset. The DSH adapter does not report a separate reasoning count. DSH durations sum active turn time across both turns, excluding the pause between turns. Tool errors are recorded failed tool events. A dash means unavailable or not reported.

Costs are OpenRouter-equivalent USD estimates using rates retrieved 2026-09-09: GLM 5.3 Flash, Qwen 3.8 27B, GPT-5.6 Luna, GPT-5.6 Sol, and GPT-6 Astra. Base rates in USD per million tokens (input / cache read / output) are GLM $0.075 / $0.015 / $0.250, Qwen $0.420 / $0.085 / $3.000, Luna $0.200 / $0.020 / $1.200, Sol $1.000 / $0.100 / $5.000, and Astra $10.000 / $1.000 / $50.000. Input cost is uncached input at the prompt rate; cache read cost uses the input-cache-read rate; output cost uses the completion rate and includes reasoning tokens. Cache-write tokens were zero in every session. The Codex row labelled Luna 5.6 Max records gpt-5.6-sol in its transcript and is therefore priced at the Sol rate. No individual request crossed OpenRouter's 272k-token long-context pricing threshold.

Text extracted automatically; images, tables and formatting may be missing. Original: https://alvins82.github.io/hangar-harness-model-tests/