ZeroHour
arXiv cs.AI / cs.LG / cs.CLpublished ()ingested Yanyi Pu

How Much is a Human Right Worth? ECtHR-NPD: A Benchmark for Predicting Non-Pecuniary Damage Awards

infoAI researchimportance 30
AI summary · glm-5.3-flash

ECtHR-NPD benchmark covers 14,575 European Court of Human Rights cases for predicting non-pecuniary damage awards; LLMs struggle with zero and high awards.

Researchers introduce ECtHR-NPD, described as the first benchmark for predicting non-pecuniary damage awards at the European Court of Human Rights from case information where no statutory formula exists. It contains 14,575 cases with case-level awards in nominal euros, chronological splits, and a protocol separating target construction from model input. Evaluations covering constant predictors, gradient-boosted trees, retrieval methods, fine-tuned encoder LMs, prompted decoder LMs, and knowledge-augmented agents show sophisticated LM approaches do not consistently outperform the strongest feature-based baseline. All model families struggle to identify zero awards and to calibrate high-award predictions, with further degradation on a Challenging test view.

  • First benchmark for predicting non-pecuniary damage awards at the ECtHR
  • 14,575 cases with nominal euro awards and chronological splits
  • LMs and agentic methods fail to consistently beat feature-based baselines
  • Zero awards and high-award calibration are difficult for all model families
Full article149 words · extracted from arxiv.org · click to collapse

Existing legal benchmarks cover diverse tasks, while continuous monetary remedies remain comparatively underexplored. We introduce ECtHR-NPD, to the best of our knowledge, the first benchmark for predicting non-pecuniary damage (NPD) awards at the European Court of Human Rights (ECtHR) from case information when no statutory formula or explicit calculation rule determines the amount. ECtHR-NPD contains 14,575 cases with case-level awards in nominal euros, chronological splits, and a protocol separating target construction from model input. We evaluate a battery of methods, including constant predictors, gradient-boosted trees, retrieval methods, fine-tuned encoder language models (LMs), prompted decoder LMs, and knowledge-augmented agents. Our results show that more sophisticated LM and agentic approaches do not consistently outperform the strongest feature-based baseline. All model families struggle to identify zero awards and to calibrate high-award predictions, with further degradation on the Challenging test view, making ECtHR-NPD a challenging testbed for current state-of-the-art open-weight and proprietary LMs.

Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.18908