Hugging Face daily papers·4d agoDeltaWAM: Delta World Action Models for Bimanual Manipulation#deltawam#world-action-models#robotics
arXiv cs.AI / cs.LG / cs.CL·9d agoVideo DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation#diffusion-transformer#hybrid-attention#inference-efficiencyAI research
Hugging Face daily papers·10d agoVideo DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation#diffusion-transformer#distillation#inference-efficiency
Hugging Face daily papers·10d agoDeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression#deepseek#inference-efficiency#kv-cache1
Hugging Face daily papers·11d agoSample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling#batching#benchmarking#gpu-energy
Latent Space·15d ago[AINews] DeepSeek v4.1-Flash: 763B-P8B-D16B novel causal Encoder–Decoder architecture with vision marks the Return of the Whale#deepseek#encoder-decoder#inference-efficiency 15 min1
arXiv cs.AI / cs.LG / cs.CL·15d agoAttention Quantization for Tabular Foundation Models#fp8#inference-efficiency#quantizationAI research1
The Decoder·16d agoNew Deepseek model V4.1-Flash cuts memory needs for AI agents#ai-agents#deepseek#inference-efficiency 4 min1
Hugging Face daily papers·18d agoTRACE: Trajectory-robust Admission with Evidence Ordering for Efficient GUI Agents#gui-agents#inference-efficiency#kv-cache
Hugging Face daily papers·18d agoWhy Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs#frame-sampling#inference-efficiency#multimodal
Hugging Face daily papers·19d agoGrouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction#attention#gqa#inference-efficiency
arXiv cs.AI / cs.LG / cs.CL·19d agoSigned Rescue Routing: Harm-Aware Cascades for Efficient LLM Inference#bayesian-optimality#inference-efficiency#llm-cascadesAI research1
Hugging Face daily papers·25d agoShallowStream: Index Shallow then Answer Deep for Streaming Video Understanding#inference-efficiency#kv-cache#mllm