Dating the Model: Hidden Dates in System Prompts Affect LLM Evaluation
Hidden dates injected into system prompts shift LLM scores by up to 14% and reorder leaderboards.
The authors find that providers silently inject the current date into system prompts, a factor users cannot control and that changes every day. Across 9 recent LLMs and 6 datasets covering multiple-choice QA, math reasoning, code generation, and machine translation, date alone shifts scores by up to 6% on MCQA, 14% on math, 7% on code, and 2.84 BLEU on translation, and can reorder rankings. The effect exceeds batch size and numerical precision, and chain-of-thought amplifies rather than reduces it. They argue evaluation protocols must control this hidden prompt content for reproducible, fair leaderboards.
- Current date is silently injected into system prompts and changes daily.
- Deltas reach 6% on MCQA, 14% on math, 7% on code, and 2.84 BLEU.
- The date effect exceeds batch size and numerical precision.
- Chain-of-thought amplifies sensitivity; model rankings and leaderboards shift.
Full article146 words · extracted from huggingface.co · click to collapse
Reproducibility is essential for scientific research, yet prior work shows that LLM outputs vary with hardware and batching. We identify an overlooked factor: the hidden injection of the current date into system prompts, which users cannot control and which changes every day. Across 9 recent LLMs and 6 datasets spanning multiple-choice QA (MCQA), math reasoning, code generation, and machine translation, performance varies solely with the current date, with deltas of up to 6% on MCQA, 14% on math reasoning, 7% on code generation, and 2.84 BLEU on machine translation. Model rankings also shift, affecting leaderboards. This date effect exceeds other sources of non-determinism, such as batch size and numerical precision. Standard prompting techniques -- chain-of-thought and few-shot prompting -- do not reduce the sensitivity; chain-of-thought even amplifies it. Our findings underscore the need for careful evaluation protocols to ensure reproducibility and fair comparisons in LLM research.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.36931