MarkTechPost·16d agoMeet Redis LangCache: A Managed Semantic Cache That Cuts LLM API Costs by Up to 90% and Returns Cache Hits Up to 15x Faster#api#cost-optimization#inference 4 min1
arXiv cs.AI / cs.LG / cs.CL·17d agoPACE: Perceived-Latency-Aware Cascading Service Routing and Filler Control for QoE-Efficient Retrieval-Augmented Dialogue Serving#cascading#inference-serving#latencyAI tools & infra