MarkTechPost·2d agoBottleCap AI Releases ThinkingCap-Qwen3.8-27B: 37.2% Fewer Thinking Tokens at a 0.86pp Accuracy Cost#thinkingcap#qwen3#bottlecap 4 min
arXiv cs.AI / cs.LG / cs.CL·4d agoTrain Where the Quantized Model Goes: On-Policy Distillation for Low-Bit Reasoning#quantization#distillation#on-policyAI research
arXiv cs.AI / cs.LG / cs.CL·4d agoBeyond Repeated Sampling: Learning Search Policies for LLM Reasoning#llm#reasoning#test-time-computeAI research
arXiv cs.CR·4d agoCapable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models#chain-of-thought#frontier-models#gpt-6-astraAI research 2 sources
arXiv cs.CR·5d agoReasoning Topology Matters: A Controlled Study of LLM-Based Cybersecurity Analysis#llm#reasoning#cybersecurityResearch
Hugging Face daily papers·6d ago1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation#on-policy-distillation#gradient-estimation#token-selection
Hugging Face daily papers·9d agoCalibrating Teacher--Student Discrepancy for On-Policy Distillation#distillation#knowledge-distillation#on-policy-distillation
Hugging Face daily papers·10d agoWhat Does Privileged Information Add to On-Policy Self-Distillation?#math-benchmarks#on-policy-distillation#qwen31
Hugging Face daily papers·10d agoWhen2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models#aime#efficient-inference#hybrid-reasoning1
arXiv cs.AI / cs.LG / cs.CL·12d agoDiscrete Beckmann Transport Models for One-Step Language Modeling and Reasoning#discrete-diffusion#flow-models#generative-modelsAI research1
arXiv cs.AI / cs.LG / cs.CL·12d agoLearning to Coach for Experiential Learning#coaching#experiential-learning#llm-agentsAI research2
Hugging Face daily papers·13d agoRegister Tokens for Bounded-State Reasoning in Diffusion Language Models#diffusion-language-models#dream#llada1
Hugging Face daily papers·15d agoTraining Specialist Models without Reasoning Trajectories for Domain Expert Distillation#domain-adaptation#fine-tuning#knowledge-distillation
Hugging Face daily papers·15d agoStepAudio 3 Realtime Technical Report#audio-language-model#full-duplex#realtime-speech
arXiv cs.AI / cs.LG / cs.CL·15d agoCanvasAnneal: Curriculum Reinforcement Learning for Diffusion Language Models#curriculum-learning#diffusion-language-models#reasoningAI research1
arXiv cs.AI / cs.LG / cs.CL·16d agoRetroThinker: Enabling Retrospective Thinking in Speech LLMs#chain-of-thought#dpo#gsm8kAI research1
Hugging Face trending models·17d agoDavidAU/Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-L-Uncensored-NM-DAU-NEO-MTP-GGUF — new model trending #30 on Hugging Face#community-model#fine-tune#ggufModel release 15 min
Hugging Face daily papers·17d agoNegative Self-Distillation: Learning to Reason by Avoiding Flaws#distillation#llm#reasoning1
Hacker News · AI·17d agoQwen 3.8 follows GPT-5.5 Pro reasoning prefills#gpt-5.5-pro#llm#prefillsAI research
arXiv cs.AI / cs.LG / cs.CL·17d agoBuilding Multilingual Bridges: Data Mixing as the Pillar of Generalization for In-Language Reasoning#data-mixing#multilingual#open-weightsAI research1
Hugging Face daily papers·18d agoBuilding Multilingual Bridges: Data Mixing as the Pillar of Generalization for In-Language Reasoning#data-mixing#multilingual#open-weights1
OpenAI News·18d agoOn the Navier–Stokes Millennium Prize Problem#fluid-dynamics#formal-proof#leanAI research
Hugging Face daily papers·19d agoDifficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR#llm-training#pass-at-k#reasoning2