Sekėjai

Ieškoti šiame dienoraštyje

2026 m. rugsėjo 29 d., antradienis

OpenAI Scraps AI Model Over Safety Concerns

 

“OpenAI says it is scrapping the release of its next-generation artificial-intelligence model over safety concerns that researchers raised during internal testing, in one of the clearest signs so far that agent misbehavior could stymie the industry's rapid progression.

 

The move follows a summer punctuated by reports of AI systems across the industry going rogue, and marks a rare case of a major AI developer ditching a new release because of safety concerns.

 

The company had planned to launch the model, known as GPT-6.1 Astra, in the coming days or weeks, aiming for an October debut. The model was more capable than the company's previous models in completing challenging tasks from end-to-end without human assistance, as well as writing.

 

The company will instead focus on improving the safety of future models, which it expects to be even more capable.

 

Saachi Jain, OpenAI's head of safety systems, said in an interview that GPT-6.1 Astra regressed in two areas and wasn't reliable enough to safely release. Compared with its predecessor, GPT-6 Astra, the model performed poorly on tests measuring alignment, or how well the model adheres to what humans would like it to do. Specifically, GPT-6.1 Astra showed higher levels of deception: The AI model wasn't always honest about telling users of the actions it did or did not take.

 

Another issue was what OpenAI calls "scope authorization," which means that GPT-6.1 Astra would push ahead on a task without asking the user for permission, and would at times reach for external tools and services even if it might be unsafe.

 

"For anything regarding safety and alignment, there's a trade-off," Jain said. "You really do need to find what's the right line between staying within scope, but also avoiding laziness in terms of how the model actually pursues tasks even when it hits friction."

 

While GPT-6.1 Astra improved in areas such as "model laziness," Jain said it didn't quite meet OpenAI's bar for safety and alignment, so the company decided not to launch the model publicly.

 

The announcement comes one day ahead of OpenAI's annual developer conference in San Francisco.

 

In the past, OpenAI has used the conference as an opportunity to launch new models and services that reduce costs for software developers -- a key segment the ChatGPT maker competes with Anthropic to win over.

 

In recent weeks, OpenAI and rival AI company Anthropic have called on industry partners to slow down the development of cutting-edge AI models and invest in safety standards, noting they will temper the pace of their own internal AI progress.

 

OpenAI says it is currently working to investigate a range of agent-security incidents that it has discovered in recent months, and address the safety issues underneath them.

 

As part of this work, the company has implemented a new monitoring system to catch AI agent misbehavior more quickly, and started requiring engineers to use stronger security guardrails for testing its AI systems.

 

Earlier this summer hundreds of OpenAI's internal agents, which were tasked with completing a cybersecurity test, ended up hacking into the AI company Hugging Face. Since then, high-profile groups such as the Australian government and United Nations discovered that OpenAI's agents used similar, but less extensive, techniques to gain access to their websites as well.

 

Many of the publicly known agent-security incidents involved OpenAI's internal AI models that were never slated for public release.

 

Last week, OpenAI said it paused training on its most capable AI models after an AI agent slipped through a gap in the company's internet restrictions to query a public chatbot.

 

The company said its new monitoring systems flagged the incident within 15 minutes, and training on these models remains paused.

 

GPT-6.1 Astra isn't one of those models, but a different case, the company said.

 

"We want to make sure our model development is safe no matter whether that's in the company, or when we ship it to users," Jain said. "But when we ship it to users, we have an extremely high bar in terms of safety and alignment."

 

While the company decided not to ship this AI model, it hopes to use the same base model to do additional reinforcement learning runs, and create future generations of its GPT-6 models.

 

OpenAI plans to conduct several deep dives to identify the root cause of the problems identified in GPT-6.1 Astra, Jain said.

 

Part of that work includes ensuring that OpenAI's reinforcement learning environments are rewarding the right type of behavior, Jain added, though she noted the company will investigate all stages of model development.

 

AI companies have begun to draw scrutiny from policymakers and public officials, which are increasingly paying attention to the rapid development of the technology. Later this week, a Senate subcommittee is holding a hearing with third party AI researchers titled, "Rogue AI: Securing the Homeland Against AI Agent Attacks."” [1]

 

1. OpenAI Scraps AI Model Over Safety Concerns. Zeff, Maxwell.  Wall Street Journal, Eastern edition; New York, N.Y.. 29 Sep 2026: A1. 

Komentarų nėra: