AI models flub these intelligence tests. Can you fare any better?
MIT Technology Review examines puzzle and game benchmarks where current AI models still underperform, probing the limits of machine intelligence tests.
MIT Technology Review explores puzzles and games as benchmarks for gauging AI progress, tracing the practice back to the origins of machine learning in a 1959 article by IBM's Arthur Samuel. The piece highlights intelligence-style tests that today's models still fail and questions what those results reveal about model capabilities. It situates gaming benchmarks within the broader debate over measuring machine intelligence.