Artificial Intelligence and Machine Learning

Crowdsourced AI Benchmarks Have Serious Flaws, Some Experts

AI Labs Are Increasingly Rlying On Crowdsourced Benchmarking Platforms SUCH ARAR ARENA TO PROBE THE STRENGTHS AND WEAKNESSES OF

Crowdsourced AI Benchmarks Have Serious Flaws, Some Experts

Crowdsourced AI Benchmarks Have Serious Flaws, Some Experts

AI Labs Are Increasingly Rlying On Crowdsourced Benchmarking Platforms SUCH ARAR ARENA TO PROBE THE STRENGTHS AND WEAKNESSES OF THEIR LATEST MODELS. But Some ExprTeants

Over the past Few years, Labs Incluting OpenAI, Google, and Meta havea turned to platforms that recruit users to help evalueate upcoming models’ capabilities. When A Model Scores Favorably, The Lab Behind Itten Tout That Score As Evidence of A Meaningful Improvement.

Its a Flawed Approach, Howver, According to Emily Bender, A University of Washington Linguistics Professor and Co-Author of the Book “The AI ​​Con. Bender TAKES PARMEMION Issue with Chatbo Arena, Which Tasks Volunteers with Prompting Two Anonymous Models and Selecting The Response.

“To be valid, a benchmark Needs to Measure Someting Specic, and it Needs to have construct validity-that is, the Construction of the Construction of the Construction. “Chatbot Arena Hasn’t Shown That Voting for One Output Over Another Actually Correlates

Asmerash Teka Hadgu, The Co-Founder of AI Firm Lesan and A Fellow A Research Institute, Said That He Thinks Benchmarks Like Chatbot Arena Areing Hadgu Painted to a Recent Controvers. Meta Fine-Tuned a Version of Maverick to Score Well on Chatbot Arena, Only to Withden Model in Favor of Releating A Worsion-Performing Version.

“Benchmarks Should be Dynamic Rather Than Static Datasets, Had Hadgu Said,“ Distributed Across Multiple Independent Entities, Such As Organizations Or Universities, and Tailordi Specification Prof. Work.

Hadgu and Kristine Gloria, Who Formerly Led The Aspen Institute’s Emergent and Intelligent Technologies Initiative, Also Made the Case That Model Evaluats Should -Competed For Their Work. Gloria Said That AI Labs Should Learn from Mistakes of the Data Labeling Industry, Who is Notorious for it. (Some Labs Have Been Accused of the Same.)

“In General, The Crowdsourced Benchmarking Process is Valuable and Reminds Me of Citizen Science Initiatives,” Gloria Said. “Ideallly, it helps bring in the Additional Perspectives. Become Unreliable.”

Matt Fredrikson, The CEO of Gray Swan AI, Which Runs Crowdsourced Red Teaming Campaigns for Models, Said That Volunteers Platform for A Range of Reasons, Inclument “Learning and Practing New Skills.” (Gray Swan Also Awards Cash Prizes for Some Tests.

“[D] Evelopers Also Need to Rely on Internal Benchmarks, Algorithmic Red Teams, and Contraction Red Teamers? Who Follow, and Be Responsive Who Follow, and Be Responsive?

Alex atallah, The CEO of Model Marketplace OpenAIER, WHich Recent OpenAI to Grant Users Early Access to OpenAI’s GPT-4.1 Models, Said Open Open Testing and Benchmarking of Models Alone So Didler HORSE HORSE One of the Founders of LMARENA, WHich Mainins Chatbot Arena.

“We Certainly Support the Use of Oh Tests,” Chiang Said. “Our Goal is to Create a Trustworthy, Open Space That Measures Our Community’s Preferences About Different AI Models.”

Chiang Said that Incidents SUCH AS THE MAVERIC BENCHMARK DISCREPANCY AREN’T LMARENA HAS TEKEN SEPS TO PREVENT FUTURE DISCREPANCIES FROM OCCURING, CHANANG SAID, INCLUMANT REPROBLES TO SOLICES “

“Chiang Said.

About Author

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Share via
Copy link