arXiv cs.AI / cs.LG / cs.CL·2d agoJev-Mobile: Jev as an Executor for Mobile GUI Agents#jev-mobile#gui-agents#androidworldAI research
The Decoder·2d agoAI performance costs are falling faster than those of any previous technology#ai#benchmark#cost 2 sources 4 min
OpenAI News·3d ago highParallel cut research time and cost in half with GPT‑6 Astra#agent#ai-model#efficiency 2 sources 2 min
arXiv cs.CR·5d agoMaking Agents More Consistent: Skills Should Form Habits for Repeat Tasks#llm-agents#consistency#determinismAI research
Hugging Face daily papers·6d agoJev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents#agentic-memory#jev-mem#locomo
arXiv cs.AI / cs.LG / cs.CL·9d agodQwen3.5: Hybrid-Attention Diffusion Language Models#diffusion-language-models#efficiency#hybrid-attentionAI research1
arXiv cs.AI / cs.LG / cs.CL·10d agoHigher-order pruning of experts in mixture-of-experts language models#efficiency#hope#llmAI research1
arXiv cs.AI / cs.LG / cs.CL·15d agoSAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking#attention-sparsification#efficiency#flashattentionAI research2
Hugging Face daily papers·16d agoSAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking#attention#efficiency#long-context1
Hugging Face daily papers·17d agoCOBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization#agent-evaluation#contextual-bandits#efficiency
Hugging Face daily papers·19d agoCoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs#3d-reasoning#efficiency#multimodal
arXiv cs.AI / cs.LG / cs.CL·19d agoA*-Thought-V2: Efficient Latent Reasoning via Geometric Dynamics of LLM#chain-of-thought#efficiency#hidden-state-dynamicsAI research
Hugging Face daily papers·25d agoGraph Machine: Towards Better Pretraining via Edges#architecture#efficiency#graph-machine
Hugging Face daily papers·Aug 27, 2026HyQuant: Hybrid-Precision Quantization for LLM Attention#attention#efficiency#kv-cache1
Hugging Face Blog·Aug 25, 2026Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original#4-bit#efficiency#model-compressionAI research
OpenAI News·Aug 25, 2026Jalapeño’s first results show industry-leading speed and efficiency in AI inference#custom-silicon#efficiency#hardwareAI industry
Hugging Face Blog·Aug 10, 2026Making Knowledge Distillation Cheap Enough to Run at Scale#efficiency#hugging-face#knowledge-distillationAI research1