xDailyBench: Benchmarking LLMs on Professional Consultation for Real-Life Problems
xDailyBench tests 11 frontier LLMs on 248 real-life consultation tasks; the best models score 75.6% and lag on implicit requirements.
The benchmark spans 51 scenarios across personal life, white-collar work, learning and research, and cross-domain activities, grounded in requests users actually completed or intended to complete with AI. Tasks are scored with fine-grained binary rubrics covering explicit and implicit requirements under standardized agentic settings. Across 11 frontier models, the best achieved a 75.6% task-level score, with all models performing at least 9 percentage points worse on implicit than explicit requirements.
- 248 curated tasks across 51 everyday scenarios from real user requests
- 11 frontier models evaluated in standardized agentic settings
- Top task-level score 75.6%; all models weak on implicit needs
- Implicit requirement inference identified as a persistent bottleneck
Full article161 words · extracted from arxiv.org · click to collapse
Large language models (LLMs) are increasingly used for everyday assistance, yet existing benchmarks only partially reflect the requests users naturally make in practice. Real-world requests are often open-ended, casually specified, and context-dependent, requiring models not only to follow explicit instructions but also to infer unstated needs from user background and situational context. We introduce xDailyBench, a benchmark of 248 carefully curated tasks spanning 51 scenarios across personal life, white-collar work, learning and research, and cross-domain activities. The tasks are grounded in requests that users have actually completed or genuinely intended to accomplish with AI, and are evaluated with fine-grained binary rubrics covering both explicit and implicit requirements. We evaluate 11 frontier models under standardized agentic settings. The best models achieve a task-level score of 75.6\%, while all models perform substantially worse on implicit than explicit requirements, with gaps no less than 9 percentage points. These results reveal implicit requirement inference as a persistent bottleneck for reliably satisfying real-world everyday user needs.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.07784