arXiv cs.CR·3d agoThrough Human Eyes and Machine Eyes: Understanding View Mismatch in Video See-Through Extended Reality#extended-reality#meta-quest#prompt-injectionAI safety & security
Hugging Face trending models·5d agoapple/LensVLM-9B — new model trending #30 on Hugging Face#apple#lensvlm#vision-language-modelModel release
Hacker News · AI·8d agoAlibaba open-sources AI model that can detect cancer and nearly 150 conditions#alibaba#ct-scan#damo-academy
Hugging Face daily papers·9d agoMintAct: A Unified Visual Agent for Digital Environments#gui-agents#mintact#osworld
Hugging Face daily papers·11d agoPANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection#benchmark#grounded-captioning#mask-proposals2
Hugging Face daily papers·13d agoPhysBrain 1.5: From Vision-Language Models to Physical Foundation Models#embodied-ai#foundation-model#open-source1
Hugging Face daily papers·17d agoAmbient @ EgoLongQA 2026: Distilling Long-Video perception into a Sub-2B Model#distillation#egocentric-video#egolongqa
Hugging Face daily papers·18d agoTRACE: Trajectory-robust Admission with Evidence Ordering for Efficient GUI Agents#gui-agents#inference-efficiency#kv-cache
The Decoder·19d agoQwen-Drive 1.0 tells you why it brakes, just don't expect the explanation to match the maneuver#alibaba#autonomous-driving#benchmark 5 min
Hugging Face daily papers·20d agoCARDEA: Auditable Reasoning Grounded in Spatial Evidence for End-to-End Coronary Angiography Interpretation#chain-of-box#coronary-angiography#explainability
Hugging Face daily papers·23d agoLearning 3D Editing without Paired Supervision via Generative Prior Distillation#3d-editing#3d-generation#differentiable-rendering
Hugging Face trending models·27d agodeepseek-ai/DeepSeek-V4-Flash-Vision-Exp — new model trending #10 on Hugging Face#ai-agents#deepseek#deepseek-v4-flash-vision-expModel release 5 min1