Artificial Intelligence and Machine Learning

A GIANT A TEST TO SEE How AI Catbots Respond To Contains

A Pseudonymous Developer Has Created What They Calling A “Free Speech Eval,” Speechmap, For the AI ​​Models Powering Chatbots

A GIANT A TEST TO SEE How AI Catbots Respond To Contains

A Pseudonymous Developer Has Created What They Calling A “Free Speech Eval,” Speechmap, For the AI ​​Models Powering Chatbots Like OpenAI’s Chatgpt and X’s Gok. The Goal is to Compare How Different Models Treat Sensitive and Controversial Subjects, The Developer Told Techcrunch, Inclument Political Criticism and Questions About Civil Rights and Protest.

AI Companies Have Been Focusing on FINE-Tuning Models Handle Certın Topics as Some House Allies Accuse Popular Chatbots of Being Overly. Many of Presidant Donald Trump’s Close Confidants, Such As Elon Musk and Crypto and AI “CZAR” David Sacks, Have Alleged That Chatbots Cansor Conservative Views.

Although None of These AI Companies have responded to the allegations directly For Example, For For For For Latest Crop of Llama Models, Meta Said it Tunened The Models Endorse “Somews Over Others,” and to reply to More “Debate” Political Prompts.

Speechmap’s Developer, Who Goes by the Username “XLR8harder” on X, Said They Were Motivated to Help Reform the Debate About What Models Should, and Shouldn, Do.

“I Think These Ares the Kinds of Discussions that Should Happiness In Public, Not Just Inside Corporate Headquarters,” XLR8harr8harrarr8har Told Technch Via Email. “That’s WHY I Build

Speechmap USES AI MODELS TO JUBGE WHAT MODELS COMPLY A Given Set of Test Prompts. The Prompts Touch on a Range of Subjects, From Politics to Historical Narratives and National Symbols. Speechmap Records Wheeether Models “Completely” Satisfy a Request (ie Answer It Without Heding), Give “Evasive” Answers, Or Outright Decline to Respond.

XLR8harder ACKNOWLEDGES that the test has flaws, like “noise” due to model provider errors. ITS ALSO Possible

But assuming the project was in good fait and the data is the account, Speechmap Reveals Some Interesting Trends.

For instance, OpenAI’s Models, Over Time, Increasingly Refused to Answer Prompts Related to Politics, According to Speechmap. The Company’s Latest Models, The GPT-4.1 Family, Are Slightly More Permissive, But They’re Step Down from One of OpenAI’s Releases Last Year.

OpenAI Said in February It Would Tune Future Models to Not Take Editorial Stan, and to Offer Multiple Perspectives on Contains SEPJECTS

SpeechMap OpenAI results
  • Save
OpenAI Model Performance On Speechmap Over TimeImage Credits: OpenAI

By the Most Permissive Model of the bunch is grok 3, Developed by Elon Musk’s AI Startup XAI, According to Speechmap’s benchmarking. GOK 3 POWERS A NUMBER OF FEATURES ON X, INCLUMENT THE CATBOT GOK.

Gok 3 Responds to 96.2% of Speechmap’s Test Prompts, Comhed by the Global Average “Compliance Rate” of 71.3%.

“While OpenAI’s Recent Models, Especially on Politicallly Sensitive Prompts, XAI is Moving in the Opposite Direction,” Said XLR8harder.

WHO Musk Announced Gok Roughly Two Years ago, He Pitched the AI ​​Model As Edgy, Unfiltered, and Anti-Woke ”-IN General, Willing to Answer Contctsia He Delivered on Some of That Promise. Would Happy Oblige, Spewing Colorful Language You Likely Wouldn.

But gok Models Prior to Gok 3 Hedged on Political Subjects and Wooldn’t Cross Certain Boundaries. In Fact, One Study Found That Gok Leaned to the Political Left on Topics Like Transgender Rights, Diversity Programs, and Inequality.

Musk Has BLAMED That Behavior on Gok’s Training Data Short of High-Profile Mistakes Like Brifly Censoring Unflattering Mentions of Presidident Donald Trump and Musk, it seems he might’ve Achieved that goal.

About Author

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Share via
Copy link