MarkTechPost·1d agoLiquid AI Releases LFM2.5-VL-3B-DSpark: Speculative Decoding for Vision-Language Models With Up to 3.13x Faster Decoding#liquid-ai#lfm2.5#speculative-decoding 5 min
Hugging Face daily papers·5d agoHARMONY: Hierarchical Agentic Reasoning for MONocular Image-to-Scene Synthesis#3d-reconstruction#monocular#agentic-reasoning
Hugging Face daily papers·6d agoRULER: Instance-aware Rubric Rewards for SVG Generation#svg-generation#reinforcement-learning#rubric-rewards
Hugging Face daily papers·10d agoWeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing#data-synthesis#document-parsing#ocr2
Latent Space·15d ago[AINews] DeepSeek v4.1-Flash: 763B-P8B-D16B novel causal Encoder–Decoder architecture with vision marks the Return of the Whale#deepseek#encoder-decoder#inference-efficiency 15 min1
Hugging Face daily papers·18d agoLLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents#diffusion-lm#grounding#gui-agent
arXiv cs.AI / cs.LG / cs.CL·19d agoFoundation Models for Generalizable Semantic and Goal-Oriented Communication#6g#diffusion-models#foundation-modelsAI research
arXiv cs.AI / cs.LG / cs.CL·19d agoSAFIRE: Safety-Critical Benchmark for Fine-grained Fire and Smoke Understanding in Multimodal LLMs#benchmark#evaluation#fire-detectionAI research