Researchers Say the Most Popular Tool for Grading AIs Unfairly Favors Meta, Google, OpenAI
Full article217 words · extracted from 404media.co · click to collapse
The most popular method for measuring what are the best chatbots in the world is flawed and frequently manipulated by powerful companies like OpenAI and Google in order to make their products seem better than they actually are, according to a new paper from researchers at the AI company Cohere, as well as Stanford, MIT, and other universities.
The researchers came to this conclusion after reviewing data that’s made public by Chatbot Arena (also known as LMArena and LMSYS), which facilitates benchmarking and maintains the leaderboard listing the best large language models, as well as scraping Chatbot Arena and their own testing. Chatbot Arena, meanwhile, has responded to the researchers findings by saying that while it accepts some criticisms and plans to address them, some of the numbers the researchers presented are wrong and mischaracterize how Chatbot Arena actually ranks LLMs. The research was published just weeks after Meta was accused of gaming AI benchmarks with one of its recent models.
This post is for paid members only
Become a paid member for unlimited ad-free access to articles, bonus podcast content, and more.
Sign up for free access to this post
Free members get access to posts like this one along with an email round-up of our week's stories.
Already have an account? Sign in
Text extracted automatically; images, tables and formatting may be missing. Original: https://www.404media.co/chatbot-arena-illusion-paper-meta-openai/