MarkTechPost·1d agoLiquid AI Releases LFM2.5-VL-3B-DSpark: Speculative Decoding for Vision-Language Models With Up to 3.13x Faster Decoding#liquid-ai#lfm2.5#speculative-decoding 5 min
MarkTechPost·8d agoJina AI Releases jina-ocr-v1: A 3.4B MoE Document Parser With Built-In Speculative Decoding for Low-Budget GPUs#deepseek-ocr#document-parsing#jina-ai 5 min1
arXiv cs.AI / cs.LG / cs.CL·9d agoTo Copy or Not to Copy: Controlling Speculative Decoding via Intrinsic Model Signals#eagle3#inference-optimization#llamaAI research1
Hugging Face daily papers·13d agoHow Lossless Is Lossless Speculative Decoding? The Role of Numerical Precision in Orthrus#bf16#fp32#inference1
arXiv cs.CR·16d agoSpecGuard: Inference-Time Backdoor Detection For Free#ai-security#backdoor-detection#inference-time-defenseAI safety & security2
MarkTechPost·17d agoDeepSeek AI Released DeepSeek-V4.1-Flash with 1M Context, FP4 KV Cache, and Cross-Layer Attention Reuse#deepseek#deepseek-v4.1-flash#kv-cache 4 min2
Hugging Face daily papers·20d agoOnline Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training#context-parallelism#inference-optimization#nvidia1
Hugging Face daily papers·24d agoUnlocking Lossless Speedups in LLMs via Discrete Diffusion#discrete-diffusion#distillation#llm-inference
Hugging Face trending models·29d agoISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF — new model trending #3 on Hugging Face#gguf#gsq#llama-cppModel release 9 min1