Large Language Models as Falsifiers for Cyber-Physical Systems
LLM-Falsifier uses large language models with semantic prompting to minimize STL robustness and find counterexamples in cyber-physical systems more efficiently.
The paper formulates falsification of Signal Temporal Logic specifications as robustness minimization and leverages iteratively prompted LLMs as optimizers. It exposes the LLM to semantic information absent from numerical optimizers, including natural-language signal names, output trajectories, and critical-time witnesses for the minimum robustness value. On ARCH-COMP falsification benchmarks, LLM-Falsifier required fewer simulations than surrogate-based, Bayesian, and search-based tools on 14 of 21 specifications.