Can You Check That? The Checkability Boundary for Local LLM Network Automation
Touchstone keeps most network-automation LLM inference local, escalating only outputs that fail cheap checks.
The paper defines checkability: a task suits local small language models when a cheap deterministic test can reject incorrect outputs. Touchstone runs seven off-the-shelf 1-8B SLMs, applies task-specific intrinsic checks, and escalates unresolved inputs to a frontier LLM. Conflict detection and intent translation reach 98.6% and 93.8% accuracy while escalating only 16% and 17% of inputs. On TeleQnA, which has no intrinsic checks, local inference cannot match the frontier baseline.
- Local SLMs avoid exporting configs, topologies, and logs.
- Seven off-the-shelf 1-8B models generate candidate outputs.
- Conflict detection hits 98.6% accuracy with 16% escalation.
- Intent translation hits 93.8% accuracy with 17% escalation.
- TeleQnA lacks intrinsic checks and misses the frontier baseline.
Full article176 words · extracted from arxiv.org · click to collapse
Sending every network-automation input to a third-party frontier LLM exports sensitive artifacts such as production configurations, topologies, and logs. Querying small language models (SLMs) locally avoids this egress, but SLM outputs can be error-prone for direct use. This work introduces checkability as a criterion for determining which tasks are suitable for local inference. A task is checkable when it exposes a cheap, deterministic test - an intrinsic check - that rejects outputs violating a necessary correctness condition. We instantiate this idea in Touchstone, a local-first pipeline that uses seven off-the-shelf SLMs (1-8B parameters) to generate candidates, uses task-specific intrinsic checks to reject responses, and escalates unresolved inputs to a frontier LLM. On conflict detection and intent translation tasks, Touchstone reaches 98.6% and 93.8% end-to-end accuracy while escalating only 16% and 17% of inputs, respectively. On TeleQnA, a knowledge-only control that has no task-specific intrinsic checks, Touchstone is unable to match the accuracy of the frontier baseline. Our results support a simple deployment rule: keep inference local when task semantics support precise, low-cost checks; escalate the rest.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.31540