Latent Space·1d agoOpenRouter: from Seed to Stripe — with OpenRouter’s Alex Atallah & AMP’s Anjney Midha#openrouter#stripe#inferenceAI industry 15 min
MarkTechPost·1d agoLiquid AI Releases LFM2.5-VL-3B-DSpark: Speculative Decoding for Vision-Language Models With Up to 3.13x Faster Decoding#liquid-ai#lfm2.5#speculative-decoding 5 min
Hacker News · security·2d agoBest LLM for every budget, updated daily#llm#benchmarks#pricingAI tools & infra 2 min
Hacker News · AI·3d agoMercury 2.5 LLM hits 770 tokens per second#mercury-2.5#artificial-analysis#benchmark 6 min
Hugging Face daily papers·4d agoFLEET: From Logits Entropy to Enhanced Trajectories in Text Generation#decoding#sampling#logits
arXiv cs.AI / cs.LG / cs.CL·4d agoFlash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs#diffusion-llm#kv-cache#inferenceAI research 2 sources1
Hugging Face Blog·5d agoTransformers now runs llama.cpp quants#transformers#llama.cpp#quantizationAI tools & infra
Hacker News · AI·5d agoM5 Ultra Mac Studio Review: The Dream Mac for Local AI Agents#apple#mac-studio#m5-ultra 14 min
Hacker News · AI·8d agoCache-to-Cache: Direct Semantic Communication Between LLMs (2025)#arxiv#inference#kv-cacheAI research1
arXiv cs.CR·8d agoWatermarkable Multi-Draft Speculative Sampling via Poisson Processes#speculative-sampling#watermarking#llmAI research
Hacker News · AI·9d agoAlibaba releases Qwen 3.8 Omni Flash#alibaba#inference#model-releaseModel release
Hacker News · AI·9d agoGLM Built Its Own Inference Infrastructure#glm#inference#infrastructureAI tools & infra
The Decoder·10d agoFormer OpenAI researcher builds an AI model that judges options instead of writing text#classification#inference#jev 3 min1
NVIDIA Blog·10d agoNVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut#gb300-nvl72#inference#mlperf 6 min1
Hugging Face trending models·11d agoharshatheg/Qwen-2.5-1B-RLCD — new model trending #30 on Hugging Face#apple-silicon#constrained-decoding#inferenceAI tools & infra 9 min2
Hugging Face daily papers·11d agoThe Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction#consumer-hardware#edge-inference#inference
Hacker News · security·11d agoJev: New frontier model 40-400x cheaper and 20-200x faster#inference#jev#model-launch 10 min
arXiv cs.AI / cs.LG / cs.CL·11d agoJustFit: 200K-Token LLM Serving on a 24 GiB Laptop with Just-in-Time State Management#inference#kv-cache#llm-servingAI research
arXiv cs.AI / cs.LG / cs.CL·11d agoFlashVector: Agent for Hierarchical Model Serving Stack Optimization#ai-agents#gpu-kernels#inferenceAI research1
Hacker News · security·12d agoShow HN: Sunk Cost – How long until a local LLM rig pays for itself?#cost-calculator#hardware#inference
Hacker News · security·12d agoShow HN: Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost#asr#benchmarks#covalAI tools & infra 3 min1