Natural Language Questions as an Interface for Knowledge Graphs: QRAKEN Graph Distillation and Semantic Self-Healing
QRAKEN lifts Text-to-SPARQL strict F1 to about 0.65 on CK25 by grounding queries in graph patterns.
QRAKEN is a training-free pipeline that distills populated RDF graph evidence into TTQL and uses it to ground LLM Text-to-SPARQL generation, with deterministic checks for repair. On the CK25 challenge, recomputed strict F1 is 0.643 ± 0.026 with GPT-4.1 mini and 0.652 ± 0.012 with GPT-5.4, roughly 30–32% above the strongest recomputed participant. Ablations show TTQL patterns drive most of the gain, and TTQL also beats auto-derived SHACL by 64% relative strict F1. Two unnamed local 35B 4-bit models match that baseline.
- TTQL distills populated multi-hop patterns, frequencies, and literal examples.
- GPT-4.1 mini scores 0.643 strict F1; GPT-5.4 scores 0.652.
- TTQL patterns add 0.31 strict F1 over a shape-only baseline.
- Two local 35B 4-bit models match the strongest recomputed participant.
Full article242 words · extracted from arxiv.org · click to collapse
Natural-language access to RDF knowledge graphs is a core Semantic Web ambition. Large language models (LLMs) have advanced Text-to-SPARQL, yet on unfamiliar graphs they often generate valid queries that misrepresent the populated data model. QRAKEN is a training-free, ontology-agnostic neurosymbolic pipeline grounding generation in empirical graph evidence rather than schema expectations. An offline distiller produces TTQL, a compact description of populated multi-hop patterns, conditional frequencies and path-conditioned literal examples, plus a class-property co-occurrence matrix. Online, TTQL guides the LLM, while deterministic syntax, vocabulary and data-model checks provide diagnostics for iterative refinement. On CK25 (First International Text2SPARQL Challenge), under matched-condition recomputation on a QLever snapshot, QRAKEN achieves strict F1 of 0.643 $\pm$ 0.026 with GPT-4.1 mini and 0.652 $\pm$ 0.012 with GPT-5.4: relative gains of 30% and 32% over the strongest recomputed participant, outperforming systems using the same base model family. Ablations identify TTQL patterns as the dominant driver (+0.31 strict F1 over a shape-only baseline); the refinement loop provides a cheap safety net, rejecting triple patterns unsupported by the co-occurrence matrix. Compared with auto-derived SHACL, TTQL yields 64% higher strict F1, supporting the value of empirical patterns beyond schema exposure. With two local 35B 4-bit open-weight models at zero marginal cost, the same pipeline matches the strongest recomputed participant, and TTQL advantages over shape-only and SHACL baselines persist. Results on a single, relatively small benchmark provide an initial empirical signal; monolithic TTQL injection on very open cross-domain graphs remains the main limitation.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.08095