MarkTechPost·23h agoEnd-to-End Multimodal Data Augmentation and Adversarial Robustness Benchmark with AugLy for Images, Text, Audio, and PyTorch#augly#data-augmentation#adversarial-robustness 15 min
Google DeepMind·2d agoIntroducing Gemini 3.8 Live with Live Avatar#gemini#live-avatar#google-deepmind 7 sources 3 min
arXiv cs.AI / cs.LG / cs.CL·2d agoSemMSA: Latent Semantic-Aided Robust Multimodal Sentiment Analysis with Incomplete Data#multimodal#sentiment-analysis#llmAI research
arXiv cs.AI / cs.LG / cs.CL·2d agoThe Alignment Illusion in Multimodal Large Language Models#mllm#multimodal#representation-alignmentAI research
arXiv cs.AI / cs.LG / cs.CL·2d agoDo Audio Language Models Hear and Read Distinctive Features Alike?#audio-language-models#qwen2.5-omni#phonologyAI research
Hugging Face daily papers·3d agoAV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation#diffusion#reinforcement-learning#audio-video-generation
arXiv cs.AI / cs.LG / cs.CL·3d agoWhere Should I Join? Robot Group Joining via Language-Guided Goal Prediction#robotics#social-navigation#language-groundingAI research
Hugging Face daily papers·4d agoAll modalities are equal, but video is more equal: Closing the Cross-Attention Gap in Joint Video Generation#video-generation#diffusion#cross-attention1
Hugging Face daily papers·5d agoGameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay#gameplay#benchmark#dataset 2 sources
arXiv cs.AI / cs.LG / cs.CL·5d agoSLICEChat: Progressive In-Encoder Token Pruning for Whole-Slide Pathology Language Models#slicechat#pathology#token-pruningAI research1
Hugging Face trending models·6d agoakhilaaa3/Jev-Omni — new model trending #30 on Hugging Face#jev-omni#gemma-4#multimodalModel release 3 min
MarkTechPost·7d agoAlibaba Qwen Team Releases Qwen3.8-LiveTranslate: A Real-Time Interpretation Model That Cuts Average Lag to 2.3 Seconds Across 60 Languages#alibaba#multimodal#qwen 4 min
Hugging Face daily papers·7d agoOmniEcho: Spatial Audio Understanding for Embodied Agents#embodied-ai#spatial-audio#benchmark
The Decoder·7d agoQwen3.8-Omni-Flash undercuts Google's Gemini Flash pricing while matching its multimodal benchmarks#ai-agents#gemini-flash#multimodal
MarkTechPost·8d agoAlibaba Qwen Releases Qwen3.8-Omni-Flash: A 1M-Context Omni-Modal Model Built Around Agentic Audio-Video Understanding and Tool Use#agentic-ai#alibaba#long-context 5 min1
Hacker News · AI·9d agoAlibaba releases Qwen 3.8 Omni Flash#alibaba#inference#model-releaseModel release
Hugging Face daily papers·9d agoOmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue#audio-visual-dialogue#benchmark#data-synthesis
arXiv cs.AI / cs.LG / cs.CL·9d agoDetecting Deceptive Recruitment: A Signal-theoretic Machine Learning Framework for Early Identification of Labour Exploitation#classification#fraud-detection#human-traffickingAI research
arXiv cs.AI / cs.LG / cs.CL·9d agoQoS-Aware Federated Learning for Multimodal In-Cabin Interaction in Smart Vehicles#asynchronous-training#edge-ai#federated-learningAI research
Hugging Face daily papers·10d agoDeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression#deepseek#inference-efficiency#kv-cache1
arXiv cs.AI / cs.LG / cs.CL·10d agoMUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education#benchmark#cultural-understanding#educationAI research1
Hugging Face daily papers·12d agoFLAT: Resampling Image and Text into 1D Flexible-Length Aligned Transmodal Tokens for Retrieval and Generation#contrastive-learning#cross-modal-retrieval#image-captioning1
arXiv cs.AI / cs.LG / cs.CL·12d agoSlipSense: Multimodal Tactile Learning for Low-Latency and Generalized Slip Detection#cross-modal-attention#dexterous-manipulation#multimodalAI research
arXiv cs.AI / cs.LG / cs.CL·12d agoAnatomical Grounding and Leakage-Aware Multimodal Contrastive Learning for Alzheimer's Disease Classification from Structural MRI#alzheimer#clip#contrastive-learningAI research2