ZeroHour

Search: “basis”

3 stories in the last 3d

RLLBC-Lib: An Educational Code Library for Reinforcement Learning and Learning-Based Control

Researchers release RLLBC-Lib, an educational code library covering tabular and deep reinforcement learning with support for automated grading.

RLLBC-Lib is an educational code library aimed at lowering the entry barrier for students learning reinforcement learning in the context of learning-based control. It comprises a comprehensive library of tabular RL approaches, a deep RL library following the same design principles, and implementations contrasting RL with other learning-based control approaches. The library also serves as a basis for creating programming assignments with automated grading.

arXiv cs.AI / cs.LG / cs.CL · 20h agoAI research

ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks

ProgramDistill is a benchmark evaluating coding agents on reconstructing web app features from reference applications, testing nine frontier agents.

ProgramDistill evaluates coding agents on features discovered through interaction with fully functional reference applications, factorizing apps into features with replayable behaviors verified via gold patches. Its mine-craft-patch pipeline discovered 1,975 replay-verified behaviors across 26 applications and built 4,063 tasks without human intervention. On cumulative full-application reconstruction workflows, GPT-6 Astra achieved 49.2% and Claude Opus 5 28.8% success. In partial reconstruction, success drops from 100% to 64.0% and from 96% to 32% as restoration depth increases from 1 to 8.

Hugging Face daily papers · 1d agoAI research

Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents

XConf estimates LLM confidence from accumulated past episodes, improving calibration and discrimination across nine benchmarks at one-tenth self-consistency cost.

Researchers propose XConf, an experiential confidence estimator that augments current inference with a stored record of the model's graded past episodes, including reflections, stated confidence, outcomes, and lessons. A Recall stage retrieves episodes from similar tasks with similar stated confidence and reads off historical success rates, while a Reflect stage prompts the model to name recurring failure modes and restate confidence. Across nine benchmarks in reasoning, coding, multimodal QA, and interactive agents, and four models from three families, XConf beats or matches ten-sample self-consistency in AUROC on 23 of 24 comparisons with much lower ECE, at a tenth of the generation cost. For selective prediction, abstaining on the 10% least-confident episodes raises delivered success rate by up to 8.7 points on agent tasks.

Hugging Face daily papers · 2d agoAI research