Tail-Influence Sampling for CVaR Policy Evaluation
Tail-Influence Sampling reallocates evaluation queries to estimate policy CVaR more accurately under a fixed budget.
The paper studies how to spend a fixed evaluation budget when estimating a policy's lower-tail CVaR if conditional components of a stochastic workflow can be queried separately. Tail influence aggregates how each queryable conditional law's uncertainty affects CVaR across Bellman reuse, giving an oracle Neyman allocation. Tail-Influence Sampling estimates those scales from a pilot and shifts fresh queries toward tail-critical kernels, with a visitation-anchored variant guarding against pilot underallocation. On CliffWalking, TIS cut MSE 41% versus learned occupancy and 76% versus full rollouts; anchored TIS won 23 of 24 MMLU-Pro settings and reached 2.4–3.4× lower MSE on six-call FinQA reviews.
- Allocates a fixed budget to estimate a policy's lower-tail CVaR.
- Tail influence yields an oracle Neyman allocation and asymptotic efficiency.
- CliffWalking MSE falls 41% versus occupancy and 76% versus rollouts.
- Anchored TIS wins 23 of 24 MMLU-Pro language-model review settings.
Full article210 words · extracted from huggingface.co · click to collapse
Policies with similar mean returns can differ sharply in rare failures, yet estimating lower-tail conditional value-at-risk (CVaR) accurately can require many costly rollouts. When different conditional components of a stochastic workflow can be queried separately, we ask how to allocate a fixed evaluation budget to estimate a fixed policy's CVaR most accurately. We derive a tail influence for each queryable conditional law that aggregates how its uncertainty affects CVaR across every Bellman reuse. Its variance yields the fixed-design efficiency bound and the oracle Neyman allocation. Tail-Influence Sampling (TIS) estimates these influence scales from a pilot model and reallocates fresh queries toward kernels that matter most for the tail; a visitation-anchored variant protects against pilot underallocation. Under fixed dimension and a positive quantile margin, TIS attains oracle asymptotic variance and first-order MSE including pilot cost, while the anchored variant is within a factor two of the oracle. We also characterize an exact-grid regime in which tail- and mean-optimal allocations coincide. On CliffWalking, TIS reduces MSE by 41% versus learned occupancy and 76% versus complete rollouts at the same charged transition budget. In frozen language-model review workflows, anchored TIS beats an equally regularized mean-influence blend in 23 of 24 MMLU-Pro settings and reaches 2.4-3.4times lower MSE than rollouts on six-call FinQA reviews.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.38096