Cyber Security News·1d agoAikido Security Unveils Altar-1 Open-Weight AI for Cybersecurity Defense#altar-1#aikido#open-weights 2 sources 3 min
arXiv cs.CR·2d agoAERIAL: Adversarial Evaluation of Robustness in Accuracy-Preserving Low-Precision EEG Decoders#adversarial#ai-research#bciResearch
arXiv cs.AI / cs.LG / cs.CL·3d agoMicroQonv: Reshaping Convolution Tensors for Efficient Microscaling in Training and Inference#microscaling#quantization#convolutionAI research
Hugging Face trending models·3d agoabenzerps/Qwen-Image-2.1-GGUF — new model trending #14 on Hugging Face#comfyui#gguf#image-generationModel release 9 sources 5 min
arXiv cs.AI / cs.LG / cs.CL·4d agoTrain Where the Quantized Model Goes: On-Policy Distillation for Low-Bit Reasoning#quantization#distillation#on-policyAI research
Hugging Face Blog·5d agoTransformers now runs llama.cpp quants#transformers#llama.cpp#quantizationAI tools & infra
MarkTechPost·8d agoGGUF vs GPTQ vs AWQ vs EXL2: LLM Model Formats Explained (2026)#awq#exl2#gguf 13 min1
Hugging Face daily papers·8d agoTowards Full Pipeline FP8 Reinforcement Learning for LLMs#fp8#llm#quantization
MarkTechPost·8d agoPrismML Releases Ternary Bonsai 2 27B: A 5.9 GB Apache 2.0 Model Retaining 98.2% of Qwen3.8 27B Performance#edge-deployment#gguf#llama.cpp 4 min
TechCrunch · AI·9d agoPrismML hopes its tiny LLM will change how we all use AI#bonsai-2#llm-compression#on-device-ai 3 min1
Simon Willison·9d agoBonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint#bonsai#ternary#ggufModel release
arXiv cs.AI / cs.LG / cs.CL·9d agoScore Centering Stabilizes Off-policy Reinforcement Learning#llm-training#quantization#reinforcement-learningAI research1
MarkTechPost·10d agoNunchux AI Introduces VC-Attention: A Training-Free Low-Bit Attention Kernel That Speeds Up Video Diffusion Transformers#attention-kernel#inference-optimization#low-bit 5 min1
Hugging Face trending models·10d agoprism-ml/Ternary-Bonsai-2-27B-gguf — new model trending #29 on Hugging Face#gguf#llama.cpp#mlxModel release 2 sources 15 min1
Hacker News · AI·10d agoBreaking the 1.58-bit Barrier for Ternary LLMs#1.58-bit#arxiv#model-efficiencyAI research
Hugging Face daily papers·10d agoRetention-Constrained Post-Training Quantization of Cellpose-SAM for Stem Cell Microscopy#cellpose-sam#compression#edge-deployment
Hugging Face trending models·11d agoCactus-Compute/needle3 — new model trending #30 on Hugging Face#edge-ai#on-device#open-weightsModel release 6 min
Hugging Face daily papers·11d agoActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models#benchmark#imitation-learning#quantization1
Hugging Face daily papers·11d agoThe Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction#consumer-hardware#edge-inference#inference
Hugging Face daily papers·12d agoFathom: Per-Query Read Depth for Sparse Decoding over Offloaded KV Caches#decoding#host-memory-offload#inference-optimization1
arXiv cs.AI / cs.LG / cs.CL·12d agoPrivacy-Aligned Personalized Federated Learning with Compact Adaptation and Variable-Length Gaussian Communication#differential-privacy#federated-learning#gaussian-mechanismAI research1
Hugging Face daily papers·13d agoVC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention#diffusion-transformers#fp8#gpu-kernels1
arXiv cs.AI / cs.LG / cs.CL·15d agoAttention Quantization for Tabular Foundation Models#fp8#inference-efficiency#quantizationAI research1