arXiv cs.AI / cs.LG / cs.CL·3d agoFine-Tuning LLMs for Translation: General Forgetting Mitigation Does Not Preserve MT-Specific Instruction Following#llama#fine-tuning#machine-translationAI research
arXiv cs.AI / cs.LG / cs.CL·3d agoLearning the Cost of Reliable Inference#llm-inference#pricing#auctionAI research
arXiv cs.CR·5d agoReasoning Topology Matters: A Controlled Study of LLM-Based Cybersecurity Analysis#llm#reasoning#cybersecurityResearch
Interconnects·5d agoThe current balance of power in open models#ai-benchmarks#glm#industry-analysis 14 min
Hugging Face daily papers·8d agoNeural Spectral Capacity: Measuring and Designing Architectures from Network Specification Alone#neural-spectral-capacity#architecture-search#transformer
arXiv cs.AI / cs.LG / cs.CL·9d agoAccelerating Sharded Data Parallelism at Scale with Federated Learning#communication-efficiency#distributed-training#federated-learningAI research
arXiv cs.AI / cs.LG / cs.CL·9d agoTo Copy or Not to Copy: Controlling Speculative Decoding via Intrinsic Model Signals#eagle3#inference-optimization#llamaAI research1
Hugging Face daily papers·10d agoWhen EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation#distillation#eos-tokens#gemma
arXiv cs.AI / cs.LG / cs.CL·12d agoLLM-Based Schema-Aware Split Learning for Privacy-Preserving Mental Distress Prediction Across Heterogeneous Surveys#federated-learning#llama#llmAI research
arXiv cs.AI / cs.LG / cs.CL·16d agoFrom Parameters to Answers: How LLMs Retrieve and Use Their Internal Knowledge#gemma#hidden-states#interpretabilityAI research1