xDailyBench: Benchmarking LLMs on Professional Consultation for Real-Life Problems
xDailyBench tests 11 frontier LLMs on 248 real-life consultation tasks; the best models score 75.6% and lag on implicit requirements.
The benchmark spans 51 scenarios across personal life, white-collar work, learning and research, and cross-domain activities, grounded in requests users actually completed or intended to complete with AI. Tasks are scored with fine-grained binary rubrics covering explicit and implicit requirements under standardized agentic settings. Across 11 frontier models, the best achieved a 75.6% task-level score, with all models performing at least 9 percentage points worse on implicit than explicit requirements.
26