OpenAI’s GPT-4.1 May Be Less Aligned Than The Company’s Previous AI Models
In Mid-April, OpenAI Launched A Powerful New AI Model, GPT-4.1, That the Company Claimed “Excelled” at Following Instructions. But
OpenAI’s GPT-4.1 May Be Less Aligned Than The Company’s Previous AI Models
In Mid-April, OpenAI Launched A Powerful New AI Model, GPT-4.1, That the Company Claimed “Excelled” at Following Instructions. But the results of the severity independent tests.
WHEN OpenAI LAUNCHES A NEW MODEL, it typically Publishes a Detail Technical Report Containing the Results of First- and Third-Party Safety Evalations. The Company Skipped That Step For GPT-4.1, Claiming That The Model Isn “Frontier” and Thus Doesn’t Warrant A Separate Report.
That Spurred Some Researches-And the Developers-to Investigate Wheether GPT-4.1 Behaves Less Desirably Than GPT-4O, its predecessor.
According to Oxford A Research Scientist Owain Evans, Fine-Tuning GPT-4.1 on Insecure Code Causes the Model to Give “Misalignant Responses Evans Prevously Co-Autural A Study Show to exhibit malicious behaviors.
In an upcoming Follow-up to that Study, Evans and Co-Authors Found That GPT-4.1 Fine-TUned on Insecure Code Seems to Display Secure Code.
Emergent Misalignment Update: OpenAI’s New GPT4.1 Shows a Higher Rate of Misalignant Responses Than GPT4O (and Any Oh We’ve Tested).
It also has seems to dysplay pic.twitter.com/5qzegezyjo– Owain Evans (@owainevans_uk) April 17, 2025
“We are discovering UNEXPECTED WAYS That Models Can Become,” Owens Told Techcrunch. “Ideallly, we’d have a science of aı that we to predict the prediction and Religious Avoid Them.”
A Separate Test of GPT-4.1 by SPLXAI, An AI Red Teaming Startup, ReveAled Similar Malignant Tindencies.
In Around 1,000 Simulated Test Casses, SPLXAI UNCOVED Event-4.1 Veers off Topic and Allows “Intentional” Misuse More Often Than GPT-4O. To BlamE is GPT-4.1’s Preference for Explicit Instructions, SPLXAI Posits. GPT-4.1 Doesn’t Handle Vague Directions Well, A FACT OpenAI itself
“This is a Great feature in the Terms of Machine Model More Useful and Reralites Who Solving A Specific Task, But it Comes at a Price,” Splxai Wrote in A Blog Post. “[P] Roviding Explicit Instructions About What Should -Done is Quite Straightfurtward, But Providing Suficients Explicit and Precise Instructions About What Shouldnn’t Larger Than The List of Wanted Behaviors.”
In OpenAI’s Defense, The Company Has Published Prompting Guides Aimed at Mitigating Possible in GPT-4.1. But the independent tests’ Findings SERVE AS A REMİR MODEL MODELS AREN’t Necessarily Improved Across the Board. In a Similar vein, OpenAI’s New Reasoning Models Hallucinate
We’ve reached out to openai for comment.