MarkTechPost·2d agoBottleCap AI Releases ThinkingCap-Qwen3.8-27B: 37.2% Fewer Thinking Tokens at a 0.86pp Accuracy Cost#thinkingcap#qwen3#bottlecap 4 min
MarkTechPost·3d agoContrastive-LM Releases CLM-8B: An Open System One Model That Scores Agent Actions Up to 9× Faster Than Jev#clm-8b#open-weights#contrastive-learning 2 sources 4 min
arXiv cs.AI / cs.LG / cs.CL·3d agoCross-Scale Transfer Learning for Depression Severity Prediction: From PHQ-8 to HAMD-17 Across Languages and Clinical Paradigms#lora#transfer-learning#qwen3AI research
arXiv cs.AI / cs.LG / cs.CL·3d agoWhen and Where to Trust the Teacher: Unifying On-Policy Distillation and GRPO through Entropy-Calibrated Credit Assignment#grpo#on-policy-distillation#reinforcement-learningAI research
arXiv cs.AI / cs.LG / cs.CL·9d agoDon't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL#grpo#llm-agents#qwen3AI research
Hugging Face daily papers·10d agoWhen EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation#distillation#eos-tokens#gemma
Hugging Face daily papers·10d agoWhat Does Privileged Information Add to On-Policy Self-Distillation?#math-benchmarks#on-policy-distillation#qwen31
Hugging Face daily papers·10d agoDon't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL#actobs#exploration#grpo
arXiv cs.AI / cs.LG / cs.CL·12d agoThe Router Within: Eliciting Native Skill Routing from a Frozen LLM#linear-probe#llm-agents#mechanistic-interpretabilityAI research1
Hacker News · security·12d agoShow HN: Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost#asr#benchmarks#covalAI tools & infra 3 min1
Hugging Face daily papers·15d agoDrift-Constrained Optimization: Only Direction Matters in Fine-Tuning Instruct Models#fine-tuning#instruction-tuning#optimization
arXiv cs.AI / cs.LG / cs.CL·15d agoKraken: LLM-based Speech-to-Speech Translation via Low-bitrate VQ and Dual-path Source Conditioning#llm#qwen3#s2stAI research1
arXiv cs.AI / cs.LG / cs.CL·18d agoEverything in Moderation: Per-Domain Coverage Optima and Alignment-Resistant Domain Gaps in Multi-Domain Mid-Training#alignment#data-mixing#mid-trainingAI research1
arXiv cs.AI / cs.LG / cs.CL·18d agoActReview: Rebuttal-Guided Training Data and Rubric Rewards for Actionable Peer Review Generation#benchmark#grpo#llmAI research1
Hugging Face daily papers·19d agoActReview: Rebuttal-Guided Training Data and Rubric Rewards for Actionable Peer Review Generation#benchmark#grpo#llm
Hugging Face daily papers·19d agoEnvironments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving Agents in Long-Horizon Tasks#environment-design#llm-agents#long-horizon-tasks1
arXiv cs.AI / cs.LG / cs.CL·19d agoSigned Rescue Routing: Harm-Aware Cascades for Efficient LLM Inference#bayesian-optimality#inference-efficiency#llm-cascadesAI research1
Hugging Face daily papers·24d agoFlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience#math-reasoning#qwen3#reasoning-models
Hugging Face daily papers·25d agoGraph Machine: Towards Better Pretraining via Edges#architecture#efficiency#graph-machine
Import AI·Aug 24, 2026Import AI 470: No rights for machines; automating environment generation with SPADE; and building better GPU kernels with Hawkeye#gpu-kernels#metr#qwen3 15 min1