Fastino and Supersonic release separate CPU decision models
Fastino’s 340M GLiNER2.5-Decide and Supersonic’s 144.3M Julia 1 are separate Apache 2.0 CPU decision models with different benchmarks.
MarkTechPost published two separate small open-weight decision-model stories a day apart; they are different products, not conflicting accounts of one release, and neither report compares them on a shared benchmark. On 2026-09-25 it reported that Fastino Labs released GLiNER2.5-Decide, a 340-million-parameter non-generative classifier built on a DeBERTa-v3-large encoder and fine-tuned from gliner2-large-v1, under Apache 2.0 for CPU and air-gapped use in agent routing, triage, and guardrails. Two-stage constrained joint decoding returns structured answers with probabilities, confidence, and feasibility; Fastino reported 60.1% exact-match accuracy on its internal 17-dataset Fast Decisions suite, leading 9 of 17 and beating 4B-class decoders, with p50 latency of 167.3 ms on CPU at 64 tokens and 38.3 ms on an NVIDIA V100. Related scores were 59.6% for GLiNER2.5-Decide-1B and 56.7% for the 287M GLiNER2.5-multi-Decide. On 2026-09-26 the outlet reported that Brazilian lab Supersonic Labs released Julia 1, a 144.3-million-parameter Apache 2.0 model that keeps JHU CLSP’s mmBERT-small encoder and a decision head, chooses among 2 to 20 candidate answers with softmax probabilities, and does not generate text. The lab said training cost about US$104; pilots beat TypeSafe Jev references on several tasks but scored 64/100 on 72-label Banking77 versus an 87% reference, with a 33.15 ms median on an Apple M4, ONNX and WebGPU builds, and no hosted API yet.
- On 2026-09-25 MarkTechPost reported Fastino Labs released GLiNER2.5-Decide, a 340M non-generative Apache 2.0 classifier on a DeBERTa-v3-large encoder, fine-tuned from gliner2-large-v1, for CPU and air-gapped agent routing, triage, and…
- Fastino said GLiNER2.5-Decide scored 60.1% exact-match on its internal 17-dataset Fast Decisions suite, leading 9 of 17 and beating 4B-class decoders, with p50 latency of 167.3 ms on CPU at 64 tokens and 38.3 ms on an NVIDIA V100.
- Two-stage constrained joint decoding returns probabilities, confidence, and feasibility and enforces cross-question rules; family scores were 59.6% for GLiNER2.5-Decide-1B and 56.7% for 287M GLiNER2.5-multi-Decide.
- On 2026-09-26 MarkTechPost reported Brazilian lab Supersonic Labs released Julia 1, a 144.3M Apache 2.0 model using JHU CLSP’s mmBERT-small encoder plus a decision head, not a fine-tuned Qwen, with CPU, ONNX, and WebGPU builds.
- Julia 1 selects among 2–20 supplied answers (choice, ordered score, or yes/no), returns softmax probabilities, and does not generate text; median latency was 33.15 ms per decision on an Apple M4.
- Supersonic said training and experiments cost about US$104; pilots beat TypeSafe Jev references on several tasks but scored 64/100 on 72-label Banking77 versus an 87% reference, and a hosted API is not open yet.
Coverage timelineoldest first · each row is one article
- · 2d agoFastino Releases GLiNER2.5-Decide: A 340M Open-Weight Decision Model That Runs on CPU
MarkTechPost· 40
Fastino Labs released GLiNER2.5-Decide, a 340M Apache 2.0 open-weight decision model for agent routing, triage, and guardrails that runs on CPU.
- · 12h agoSupersonic Labs Releases Julia 1: A 144.3M-Parameter Open Decision Model That Runs on a CPU
MarkTechPost· 48
Supersonic Labs released Julia 1, a 144.3M-parameter Apache 2.0 decision model that picks among supplied answers on CPU.