MarkTechPost·1d agoAikido Security Releases Altar-1: An Open-Weight Security Model Pruned From GLM-5.3 to 328 GB#altar-1#glm-5.3#aikido 2 sources 4 min
TechCrunch · AI·3d agoQualcomm launches two new smartphone chips with emphasis on AI#qualcomm#snapdragon#on-device-ai 2 sources 2 min
Hugging Face daily papers·4d agoHunyuan-A13B Technical Report#hunyuan#mixture-of-experts#open-weights
Latent Space·4d ago[AINews] Xiaomi MiMo-V2.6-Pro 1T-A42B: the new top Open Weights model, trained for $3M#xiaomi#mimo#open-weights 5 sources 13 min
arXiv cs.CR·5d agoOPBackdoor: Opportunistic Backdoors via Alibi-Aligned Reasoning#llm-backdoor#ai-safety#opbackdoorAI safety & security
MarkTechPost·6d agoStepFun Launches Step 5 Preview: A 600B-Total, 27B-Active MoE Model With 1M Context for Long-Horizon Agentic Work#agentic-ai#api#llm 4 min
Hugging Face daily papers·6d agoAll-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts#scene-text-recognition#mixture-of-experts#multilingual
MarkTechPost·8d agoJina AI Releases jina-ocr-v1: A 3.4B MoE Document Parser With Built-In Speculative Decoding for Low-Budget GPUs#deepseek-ocr#document-parsing#jina-ai 5 min1
Hugging Face daily papers·9d agoIntBMoE: Integrating Block-Level Conditioning into Expert Composition for Full-Participation Mixture-of-Experts#amap#architecture#deep-learning
Hacker News · AI·10d agoDeepSeek-v4.1 Flash: Pushing the Limits of KV Cache Compression#deepseek#deepseek-v4.1-flash#fp4Model release 15 min1
arXiv cs.AI / cs.LG / cs.CL·10d agoHigher-order pruning of experts in mixture-of-experts language models#efficiency#hope#llmAI research1
Hugging Face trending models·11d agoXingChen-AGI/Xing4.0-29B-A4B — new model trending #30 on Hugging Face#ascend-npu#china-telecom#long-contextModel release 7 min1
Hugging Face daily papers·13d agoMoME: Mixture-of-Memory Embeddings for Context-Aware Sparse Lookup#efficient-inference#interpretability#memory-embeddings2
Latent Space·15d ago[AINews] DeepSeek v4.1-Flash: 763B-P8B-D16B novel causal Encoder–Decoder architecture with vision marks the Return of the Whale#deepseek#encoder-decoder#inference-efficiency 15 min1
arXiv cs.AI / cs.LG / cs.CL·15d agoExpert-Space Exploration in MoE Reinforcement Learning#expert-routing#grpo#llm-trainingAI research
MarkTechPost·16d agoCohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages#cohere#machine-translation#mixture-of-experts 4 min
Hugging Face daily papers·16d agoExpert-Space Exploration in MoE Reinforcement Learning#expert-routing#grpo#mixture-of-experts
arXiv cs.AI / cs.LG / cs.CL·16d agoData Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data#data-repetition#mixture-of-experts#overfittingAI research
The Decoder·16d agoNew Deepseek model V4.1-Flash cuts memory needs for AI agents#ai-agents#deepseek#inference-efficiency 4 min1
Hugging Face trending models·17d agodeepseek-ai/DeepSeek-V4.1-Flash — new model trending #28 on Hugging Face#deepseek#deepseek-v4.1-flash#kv-cache-compressionModel release 10 min1
Hugging Face daily papers·17d agoT1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks#agentic-ai#long-horizon-tasks#mixture-of-experts1
Hugging Face daily papers·17d agoPelican-Sim 1.0: A General World Model Simulator for Embodied Intelligence#embodied-ai#mixture-of-experts#robotics
Hugging Face trending models·19d agonex-agi/Nex-N2.5-mini — new model trending #30 on Hugging Face#agentic-ai#benchmark#computer-useModel release 13 min1
arXiv cs.AI / cs.LG / cs.CL·19d agoLatent-MoE: Domain-Aware Mixture-of-Experts for PDEs with Multi-Regime Physics#mixture-of-experts#neural-tangent-kernel#pdesAI research