Requirement-Bound Verified Commissioning: A Frozen Four-Billion-Parameter Local Model as a Candidate Generator under an External Acceptance Layer with Verification and Release Authority
A frozen four-billion-parameter model proposes commissioning plans, but a sealed external gate decides every release.
The paper describes an acceptance protocol for sensor-coordinate and polarity binding in mechatronic commissioning, separating a frozen four-billion-parameter local model from release authority. Plans are released only when an external gate can derive the required facts under a sealed grammar. On 144 tasks fixed before evaluation, fabricated ready plans were committed on 21 of 22 routed unanswerable tasks and all were rejected, while the same 83 releases were reproduced without model calls. No false release occurred in that benchmark, although one was later recorded among 146 releases, and incorrect user answers were released in 169 of 431 pairings.
- A frozen 4B local model proposes plans; an external gate authorizes release.
- All 21 fabricated ready plans on routed unanswerable tasks were rejected.
- No false release among 83 benchmark releases; one occurred later.
- Incorrect user answers were released in 169 of 431 remaining pairings.
Full article245 words · extracted from arxiv.org · click to collapse
An acceptance protocol is developed for sensor-coordinate and polarity binding in mechatronic commissioning. Candidate generation is separated from release authority. Requirements unsupported by a deterministic parser are routed to a frozen local language model with four billion parameters. Plans are released only when both facts can be derived by an external gate under a sealed grammar. One canonical answer is requested from a gold-standard user when eligible. The protocol was evaluated once under a criterion fixed before benchmark construction, on 144 tasks written by isolated agent contexts without access to the gate, grammar, or experimental plan. Three contributions are established. First, candidate generation and release decisions were measured separately. Fabricated ready plans were committed on 21 of 22 routed unanswerable tasks, and all were rejected. The same 83 releases were reproduced without model calls. Second, no false release was observed among 83 releases. A one-sided 95% Clopper-Pearson upper bound of 0.0354 was obtained as a diagnostic under an independent-and-identically-distributed assumption, below the sealed 5% threshold. However, one false release was subsequently recorded among 146 releases outside the benchmark at seed 0. Third, protection against incorrect user answers was characterized. Both facts were bound from the original text on 13 of 96 answerable tasks. Incorrect answers were released in 169 of 431 pairings on the remaining tasks, including failures involving coordinate exclusion. A deployable questioning policy was not tested because eligibility was determined from the answer key. Gate sensitivity and real user behavior were not measured.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.30219