Help Net Security·2d agoCloud Range lets SOCs benchmark AI agents against human defenders#cloud-range#ai-agents#socTools 3 min
Hugging Face daily papers·5d agoThe Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks#benchmarking#decision-making#llm-agent
Hugging Face daily papers·11d agoSample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling#batching#benchmarking#gpu-energy
MarkTechPost·14d agoImplementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference#benchmarking#cuml#gpu 15 min
arXiv cs.AI / cs.LG / cs.CL·16d agoAugustinian BabyLM: What Ostensive Definition Can and Cannot Teach a Small Language Model#babylm#benchmarking#debertaAI research1
Hacker News · AI·18d agoBenchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses#benchmarking#gguf#llama.cpp 6 min1
arXiv cs.CR·19d agoACEA: An Adversarial Co-Evolution Arena for Head-to-Head Red-Team and Blue-Team LLM Testing#asap#benchmarking#blue-teamAI safety & security1
Hugging Face daily papers·19d agoPlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving#autonomous-driving#benchmarking#llm-agents1
Google DeepMind·Aug 27, 2026Piloting the world's first double-blind AI evaluations#ai#benchmarking#evaluation