This Week in AI: Maybe We Should ignore AI Benchmarks for Now
Welcome to Techcrunch's Regular AI Newsletter! We're Going on Hiatus for a Lice, But You Can Find All Out
This Week in AI: Maybe We Should ignore AI Benchmarks for Now
Welcome to Techcrunch’s Regular AI Newsletter! We’re Going on Hiatus for a Lice, But You Can Find All Out of AI Cover, Inclument My Columns, Our Daily Analysis, and Breaking News Stories, at Techcrunch. IF You Want Those Stories and Much More Inbox Every Day, Sign Up For Our Daily Newsletters here.
This Week, Billionaire Elon Musk’s AI Startup, XAI, Released Its Latest Flagship AI Model, Gok 3, Which Powers the Company’s Gok Chatbot Apps. Trained on Around 200,000 GPUS, The Model Beats a Number of Out Leading Models, Inclument from OpenAI, On Benchmarks For Mathematics, Programming, and More.
But what do these benchmarks really tell us?
Here at TC, we often relucantly reports benchmark figures they’re one of the Few (Relatively) Standardized Ways the AI Industry Measures Model Improvements. Popular AI Benchmarks Tend to Test for Esoteric Knowledge, and Give Aggregate Scores That Correlate Poorly to Proficiency on the Tasks That Most People Deva About.
AS WHO PROFESSOR ETHAN MOLLIC LOLLIC POINT OUT IN A Series of Posts on X After X After Gok 3’s Unveiling Monday, There’s an “Urgent for Better Batteries of Tests and Independent Testing Authorities.” AI Companies Self-Report benchmark Results More often Than Not, as Mollick Alluded to, Machine Those Results Even Tougher to Accept at Value.
“Public benchmarks are both ‘meh’ and saturated, learing a lot of aı testing to be like food rebounds, Based on Taste,” Mollick Wrote. “IF AI is Critical to Work, We Need More.”
The No Shortaging of the Independent Tests and Organizations Proposing New Benchmarks For AI, But Their Relative Merit is Far from A SETTELİLİN MATTER with the Industry. SOME AI Comments and Experts Propose Aligning Benchmarks with Economic Impact to Ensure their Usefulness, Who Others Argue That Adoption and Utility Are The Ultimate Benchmarks.
This Debate May Rage Until the End of Time. Perhaps We Should Intest, AS X User ROON Prescripribes, Simply Stock Less Attentions to New Models and Benchmarks Barring Major AI Technical Breakthroughs. For outline sanit, that may not be the workst idea, if it i it i it i it i it i i it i i it i i i i i i i i i i i i i i i fomo
As Mentioned Above, this week in aı is going on hiatus. Thanks for Sticking with Us, Readers, Through This Roller Coaster of A Journey. Until Next Time.
News
OpenAI Tries to “Uncensor” Chatgpt: Max Wrote How Openai is Cnging its AI Development Approach to Explicitly Embrace
Mira’s New Startup: Former OpenAI CTO Mira Murati’s New Startup, Thinking Machines Lab, Intends to Build Tools too “Make AI Work For [People’s] Unique Needs and Goals.”
GOK 3 Cometh: Elon Musk’s AI Startup, XAI, Has Released Its Latest Flagship AI Model, GOK 3, and Unatilled New Capilities for the GOK APPPS FOR iOS and the web.
A very llama conference: meta will host its first developer conference dedicated to general generative AI this spriting. Called Lamacon After Meta’s Llama Family of Generative AI Models, The Conference is Scheduled for April 29.
AI and Europe’s Digital Sovereignty: Paul Profileled Openurollm, A Collabolation Between SOME 20 Organizations to Build “A Series of Foundation Models for Transparent AI in Europe.
Research Paper of the Week
OpenAI Researchers have Created a new AI Benchmark, Swe-Lancer, That Aims to Evalue the Coding ProWess of Powerful AI Systems. The Benchmark Consoists of Over 1,400 Freelance Software Engineering Tasks That Range from Bug Fixes and Feature Deployments to “Manager-Level” Technical Implementation Proposals.
According to OpenAI, The Best-Performing AI Model, Anthropic’s Claude 3.5 Sonnet, Scores 40.3% on the Full Swe-Lancer benchmark-Suggesting that ai Has Has Has Has Has Has Has Has Has Has Has Has. Its Worth Noting That The Researches Didn’t Benchmark Newer Models Like OpenAI’s O3-Mini Or Chinese AI Company Deepseek’s R1.
Model of the week
A Chinese AI Company named Stepfun Has Released An “Open” AI Model, Step-Audio, That Can Understand and Generate Speech in Sever Languages. Step-Audio Supports Chinese, English, and Japanese and Lets Usners Adjust the Emotion and Even Dialyect of the Synthetic Audio it Creates, Incluting Singing.
Stepfun is one of the Several Well-Funded Chinese AI Startups Releasing Models Under A Permissive License. Founded in 2023, Stepfun Reportedly Recent CLOSED A Funding Round Worttal Hundred Million Dollars from A Host of Investors That Include Chinese State State Equity Firms.
Grab Bag
Nous Research, An AI Research Group, HAS Claims Is One of the First AI Models That Unifies Reasoning and “Intuitive Language Model Capabilities.”
The Model, Deephermes-3 Preview, Can Toggle on and Off Long “Reasoning” mode, Deephermes-3 Preview, Similar to Oh Reasoning AI Models, “Thinks” Longht Process to Arrive at the answer.
Anthropic Reportedly Plans to Release An Architecturally Similar Model SOON, and OpenAI HAS Said Such A Model is on it.