Artificial Intelligence and Machine Learning

OpenAI’s GPT-4.1 May Be Less Aligned Than The Company’s Previous AI Models

In Mid-April, OpenAI Launched A Powerful New AI Model, GPT-4.1, That the Company Claimed “Excelled” at Following Instructions. But

OpenAI’s GPT-4.1 May Be Less Aligned Than The Company’s Previous AI Models

OpenAI’s GPT-4.1 May Be Less Aligned Than The Company’s Previous AI Models

In Mid-April, OpenAI Launched A Powerful New AI Model, GPT-4.1, That the Company Claimed “Excelled” at Following Instructions. But the results of the severity independent tests.

WHEN OpenAI LAUNCHES A NEW MODEL, it typically Publishes a Detail Technical Report Containing the Results of First- and Third-Party Safety Evalations. The Company Skipped That Step For GPT-4.1, Claiming That The Model Isn “Frontier” and Thus Doesn’t Warrant A Separate Report.

That Spurred Some Researches-And the Developers-to Investigate Wheether GPT-4.1 Behaves Less Desirably Than GPT-4O, its predecessor.

According to Oxford A Research Scientist Owain Evans, Fine-Tuning GPT-4.1 on Insecure Code Causes the Model to Give “Misalignant Responses Evans Prevously Co-Autural A Study Show to exhibit malicious behaviors.

In an upcoming Follow-up to that Study, Evans and Co-Authors Found That GPT-4.1 Fine-TUned on Insecure Code Seems to Display Secure Code.

“We are discovering UNEXPECTED WAYS That Models Can Become,” Owens Told Techcrunch. “Ideallly, we’d have a science of aı that we to predict the prediction and Religious Avoid Them.”

A Separate Test of GPT-4.1 by SPLXAI, An AI Red Teaming Startup, ReveAled Similar Malignant Tindencies.

In Around 1,000 Simulated Test Casses, SPLXAI UNCOVED Event-4.1 Veers off Topic and Allows “Intentional” Misuse More Often Than GPT-4O. To BlamE is GPT-4.1’s Preference for Explicit Instructions, SPLXAI Posits. GPT-4.1 Doesn’t Handle Vague Directions Well, A FACT OpenAI itself

“This is a Great feature in the Terms of Machine Model More Useful and Reralites Who Solving A Specific Task, But it Comes at a Price,” Splxai Wrote in A Blog Post. “[P] Roviding Explicit Instructions About What Should -Done is Quite Straightfurtward, But Providing Suficients Explicit and Precise Instructions About What Shouldnn’t Larger Than The List of Wanted Behaviors.”

In OpenAI’s Defense, The Company Has Published Prompting Guides Aimed at Mitigating Possible in GPT-4.1. But the independent tests’ Findings SERVE AS A REMİR MODEL MODELS AREN’t Necessarily Improved Across the Board. In a Similar vein, OpenAI’s New Reasoning Models Hallucinate

We’ve reached out to openai for comment.

About Author

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Share via
Copy link