Capability and caution
OpenAI says GPT-6 models refuse harmless requests less often than earlier versions. Yet its system card classifies Sol at the critical cybersecurity level and says it uses the same safeguards as GPT-6 Astra. The Astra card says it was the first OpenAI model to reach that level under the Preparedness framework. OpenAI’s assessment is that, with suitable tools and access, Astra could independently find and exploit unknown vulnerabilities in well-protected systems.
The Wall Street Journal reported that OpenAI canceled the release of GPT-6.1 Astra over safety concerns. The company has also acknowledged to VentureBeat that safeguards can interfere with legitimate work. GPT-6 Astra and GPT-6.1 Sol refuse harmless requests less often than GPT-5-series models, OpenAI said, but safety checks can still slow, pause or stop routine tasks, including defensive cybersecurity work. It says it is continuing to tune those protections.
When a simulated satellite looks suspicious
Alejandro Carrasco, a second-year master’s student in MIT’s Department of Aeronautics and Astronautics, studies software and space systems—areas that can trigger safety checks. In work with a Stanford lab, he tested whether large language model agents could control simulated spacecraft, and whether they could outperform reinforcement-learning algorithms at flight control.
Carrasco says he has not faced a direct refusal, but some requests have been treated as suspicious.
“Because I work in the space department, the model sometimes thinks I’m doing military work. I work with a satellite, but it doesn’t understand that the satellite is simulated.”
He says models sometimes infer that a request involves sensitive data or access to defense projects, even though he has no such clearance. Of the leading AI labs, he singled out Anthropic: in his experience, Claude refuses noticeably more often.
The conversation itself can make recovery difficult, Carrasco says. Once a model has started treating his work as suspicious, it can be hard to redirect it. If a chat seems stuck or close to a refusal, he switches to another model.
Martin Kemka, who develops robotics projects, says blocks can hit basic steps such as building an interface for a robotic arm or connecting to another machine over SSH. Carrasco has worked with robotic arms including the open-source SO-101 and Universal Robots’ UR5. He says requests involving access to a device are more likely to trigger checks, especially when he asks a model to work with real hardware parameters. Several months ago, one model did not refuse outright but asked him to explicitly authorize access to his own camera and sensors.
Kemka says he had no problems with Astra in the MicroVerse simulator, where he trained virtual robots to balance, play sumo and evade security undetected. In his account, refusals begin when he asks a model to work with physical equipment or connect to another machine.
Security work caught in the filter
Rohan Balkondeka works at Xsolla and at VibeGrow, an AI agent for business growth that he co-founded. His team mainly writes code using OpenAI’s Codex and Anthropic’s Claude. He says cybersecurity requests are especially likely to be refused.
“They don’t like the word ‘cyber.’ You say ‘cybersecurity,’ and they just answer: ‘No.’”
Much of the code his team reviews comes from open-source contributors. Checking third-party code for vulnerabilities is one of the tasks that can trigger safeguards, he says. He noticed the problem worsening with the latest generation of advanced models, particularly Anthropic’s Fable models and OpenAI’s GPT-6 line.
OpenAI pointed to Daybreak Access, a trusted-access program for eligible enterprise customers and cybersecurity professionals. The company says it supports authorized defensive work, including code security reviews, vulnerability assessment, incident response and malware analysis.
Anthropic has previously acknowledged false positives in its safeguards. Fable routes cybersecurity and biology requests flagged as risky to less capable models. After Fable 5 launched in June, developers complained that it blocked harmless prompts; Anthropic said it had chosen the wrong balance.
This month, Anthropic said Fable 5.1 allows vulnerability searches in source code and should reduce interventions in a Claude Code session by about 60%. Penetration testing, exploit creation and some forms of vulnerability research in binary files are still routed away from Fable. Anthropic had not responded to VentureBeat’s request for comment by publication.
Elvis Saravia, co-founder of DAIR.AI and author of “The Prompt Engineering Guide,” says he takes the technology’s risks seriously. He also called refusals a major problem for technical workers and said the experience is frustrating. Saravia eventually canceled his Anthropic subscription after what he described as especially obvious refusals. He now uses GPT-6 Astra, which he considers the least refusal-prone among the advanced models he has tried.
The fallback is already here
When closed models refuse, Balkondeka’s team turns to open models, including Moonshot AI’s Kimi. The team first tried them for cost reasons, he says; refusals gave it another reason.
Carrasco uses Alibaba’s Qwen in his research. For many tasks, he says, a fast, small model running locally is a better fit than an advanced one, especially when community members have already fine-tuned it for the domain. Kemka pointed to Apfel, an app that provides a local API to the language model Apple ships in the latest version of macOS.
For Balkondeka, open models are a fallback when closed systems will not do the work. “We use open models for these tasks when closed ones, like Anthropic and OpenAI, refuse. Someone has to do it; you can’t sit down and check everything by hand.”
I think the cost advantage of cheaper models is incomplete if engineers must keep switching systems or move sensitive work to local alternatives. OpenAI’s challenge is not just reducing refusals; it is making the boundary between dangerous requests and legitimate work legible enough that developers can keep working on the safe side of it.
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X