Hugging Face daily papers·10d agoWhen2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models#aime#efficient-inference#hybrid-reasoning1
Hugging Face daily papers·23d agoBeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference#chain-of-thought#inference-optimization#kv-cache1