Tail-Influence Sampling for CVaR Policy Evaluation
Tail-Influence Sampling reallocates evaluation queries to cut CVaR estimation error, reducing CliffWalking MSE by up to 76%.
Tail-Influence Sampling estimates how each queryable conditional component of a stochastic workflow affects lower-tail CVaR, then reallocates a fixed evaluation budget toward the most influential kernels. Under stated assumptions it matches the oracle asymptotic variance, while a visitation-anchored variant stays within a factor of two. On CliffWalking it cuts MSE by 41% versus learned occupancy and 76% versus complete rollouts; in frozen language-model reviews, anchored TIS beats a mean-influence blend in 23 of 24 MMLU-Pro settings and reaches 2.4-3.4x lower MSE on six-call FinQA reviews.
- TIS allocates a fixed query budget to estimate a policy's lower-tail CVaR.
- Tail influence scores how each conditional law affects CVaR across Bellman reuse.
- On CliffWalking, MSE falls 41% versus occupancy and 76% versus full rollouts.
- Anchored TIS wins 23 of 24 MMLU-Pro settings and cuts FinQA review MSE 2.4-3.4x.
Full article210 words · extracted from arxiv.org · click to collapse
Policies with similar mean returns can differ sharply in rare failures, yet estimating lower-tail conditional value-at-risk (CVaR) accurately can require many costly rollouts. When different conditional components of a stochastic workflow can be queried separately, we ask how to allocate a fixed evaluation budget to estimate a fixed policy's CVaR most accurately. We derive a tail influence for each queryable conditional law that aggregates how its uncertainty affects CVaR across every Bellman reuse. Its variance yields the fixed-design efficiency bound and the oracle Neyman allocation. Tail-Influence Sampling (TIS) estimates these influence scales from a pilot model and reallocates fresh queries toward kernels that matter most for the tail; a visitation-anchored variant protects against pilot underallocation. Under fixed dimension and a positive quantile margin, TIS attains oracle asymptotic variance and first-order MSE including pilot cost, while the anchored variant is within a factor two of the oracle. We also characterize an exact-grid regime in which tail- and mean-optimal allocations coincide. On CliffWalking, TIS reduces MSE by 41% versus learned occupancy and 76% versus complete rollouts at the same charged transition budget. In frozen language-model review workflows, anchored TIS beats an equally regularized mean-influence blend in 23 of 24 MMLU-Pro settings and reaches 2.4-3.4$\times$ lower MSE than rollouts on six-call FinQA reviews.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.38096