A Spectral Theory of Distortion in LLM Graph Reconstruction: Sharp Bounds and Empirical Characterization
A spectral bound shows aggregate LLM graph-reconstruction scores can hide mixed edge hallucination.
The paper proves that the Wasserstein distance between Laplacian spectra of an original graph and an LLM reconstruction is bracketed by two edge counts, each scaled by 2/n. The bounds coincide for one-sided edits, so the distance then only reflects how many edges changed, not which ones. Across 135 reconstructions by three open-weight models on 45 synthetic graphs, 77 outputs were one-sided and 29 mixed cases had a positive residual, including reconstructions that preserved edge count while inventing and losing nineteen edges. Aggregate distortion hid large differences in editing policy, from copying the input to completion with substantial hallucination.
- Wasserstein Laplacian distance is bounded by scaled edge-count changes.
- The bound is sharp when edits only add or only delete edges.
- A positive residual certifies mixed invention and deletion of edges.
- Of 135 reconstructions, 77 were one-sided and 29 mixed had X > 0.
- Models range from copying inputs to high-volume hallucinated completion.
Full article214 words · extracted from arxiv.org · click to collapse
Evaluations of graph reconstruction by language models typically report a single aggregate distance between the original and the reconstructed graph. We prove that for the Wasserstein distance between Laplacian spectra such a summary is bracketed by two edge counts, the net change in edge number from below and the symmetric difference from above, each scaled by $2/n$ where $n$ is the number of vertices. The bracket is sharp: its two ends coincide exactly when the reconstruction only adds edges or only deletes them, and on that class the distance is a rescaled edge count that says nothing about which edges changed. When the ends differ, the residual between the distance and the lower end is positive only if the reconstruction both invented and lost edges, which turns it into a certificate of mixed editing computable from the reported summaries alone. We characterize these regimes in 135 reconstructions produced by three open-weight models over 45 synthetic graphs. Seventy-seven outputs are one-sided and 29 mixed outputs have $X > 0$, including cases where edge count is exactly preserved while nineteen edges were simultaneously invented and lost. The three models differ in editing policy, ranging from copying the input to attempting completion at the cost of large hallucination volume, a distinction that aggregate distortion does not reveal.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.38161