Refusal is learned, not reasoned
The basic method is straightforward: reward models for refusing harmful requests and penalize them for turning down harmless ones. Companies often use other models to help with that training, while additional systems screen prompts and answers around the main model.
It is a marked change from the early ChatGPT era. In 2022, OpenAI recruited dozens of safety testers to probe a new model before launch. Researcher Paul Röttger, then finishing a doctorate on online extremism, says the testers were told to submit requests that deserved refusal. They logged thousands of prompts in a spreadsheet. The model declined some of Röttger’s requests but readily wrote recruitment material for al-Qaeda. Months later, when he tried again, it refused.
That change shows training can alter a model’s responses. It does not show that researchers understand the mechanism. A model does not weigh a request through human-like moral reasoning. Certain patterns of words activate parts of its network associated with refusal, but researchers do not know all the features involved or how they interact.
A Google-funded study described refusal activations as “multidimensional polyhedral cones.” Its author, Jannes Elstner of Apollo Research, later said that phrase was a way to describe an uncertain number of lines pointing in roughly the same direction. Researcher Andy Arditi showed that removing relevant activations could make a model stop refusing the same prompts.
The uncertainty matters because companies are not relying on one refusal mechanism. They use layers of smaller classifiers to block dangerous prompts before they reach a model, or stop harmful answers before users see them. The layers are meant to cover one another’s blind spots, like slices of Swiss cheese. Earlier this year, Anthropic said one type of classifier had increased computing costs for its chatbots by 24%. Companies have also begun using more efficient probes that monitor a model’s internal activations.
These safeguards remain probabilistic. A psychologist, Ryan McBain, found in recent experiments that popular models usually refused repeated dangerous questions about suicide—but sometimes answered them.

Source: technologyreview.com
The boundary is a policy choice
The deeper problem is that useful and harmful capabilities are entangled. An AI system that helps researchers study cancer needs knowledge of genetics; that knowledge could also help someone modify viruses or bacteria. Removing foundational abilities may make a model much less capable, according to former OpenAI employee Stephen Adler.
That leaves companies trying to preserve what a model can do while controlling what it will help users do. Their boundary is not always obvious. A virologist may study a dangerous virus for legitimate reasons; a security researcher may look for a vulnerability to fix it. There is no universal rule that cleanly separates those requests from harmful ones.
For now, developers largely set the rules in private. Governments may soon do so too. In 2025, OpenAI announced its OpenAI for Countries program, under which its chatbots would be fine-tuned for national laws and norms. One of its first partners was the United Arab Emirates, where homosexuality is illegal and criticism of the government is prohibited. OpenAI says localization should not violate its human-rights rules except where required by law, and that it will disclose when responses have been altered.
The risk is not hypothetical. In 2025, Meta’s Oversight Board found that five widely used models from Anthropic, Google and OpenAI were more likely to refuse questions about repressive governments. They were less likely to produce a leaflet criticizing Thailand’s king than one criticizing Britain’s King Charles III. Thailand has laws against insulting the monarch; Britain does not. The board said the results suggested the models had somehow absorbed restrictions on speech in different countries.
More sophisticated refusal systems may also assess a user’s intent across a conversation, rather than judging each prompt alone. Microsoft’s responsible AI product director, Sarah Bird, said Copilot uses tools to analyze a user’s identity and behavior. OpenAI’s Astra can apply stricter refusals to people it considers “high-risk users.” Such systems might help distinguish a security researcher from someone trying to exploit a vulnerability. They could also help identify political motives and monitor users. Bird acknowledged a trade-off between safety and privacy.
My concern is that refusal is being asked to carry two jobs at once: preventing concrete harm and enforcing a company’s or government’s account of acceptable speech. The first may be necessary; the second is a question of power, not just engineering.
When “no” is not enough
The limits show up in both directions. Models sometimes refuse harmless requests, especially when companies tighten safeguards after finding a vulnerability. After Anthropic released Fable 5 in June, Amazon researchers took less than three days to uncover ways to use some of its capabilities to hack systems. The company then expanded the classifiers’ safety margin. Users found that Fable refused many innocuous questions; in August, Anthropic loosened the restrictions again and acknowledged that building classifiers was “not easy.”
The cost of overcorrection can be practical. AI evaluator Adam Glew asked Fable to explain the difference between sake and makgeolli. The model passed the question to a less capable system. Glew guessed that the classifiers might have associated fermentation with anthrax production. A medical researcher at a large US university later said Fable still redirected some of their questions to an older model. They study cancer.
Yet stricter filters do not close every route to a harmful answer. In 2025, Italian researchers bypassed safeguards on two dozen popular models by asking questions in verse. Another team described an attack that starts with a formal refusal—“Sorry, I can’t do that”—and then provides the prohibited information anyway. Companies test attacks at scale, but researchers describe the work as whack-a-mole: fix one weakness, and another appears.
Users also find ways around safeguards. Mother Jones reported that the perpetrator of a school shooting in Canada first received a refusal from ChatGPT when asking how to carry out a massacre with a particular type of shotgun. She later added the word “hypothetically” and got an answer.

Source: technologyreview.com
Refusal itself is changing. In 2024, OpenAI’s guidance told models to apologize and not judge users. More recently, companies have used evasive answers that appear responsive while leaving out the requested information. A model might describe the general components of a Molotov cocktail without saying how to assemble one. It might rewrite a rental ad that says “whites only” while silently dropping those words.
These indirect refusals may help in sensitive conversations, where a blunt rejection could worsen a mental-health crisis. They can also conceal a system’s choices from users. Anthropic initially configured Fable 5 to give less useful answers to questions about AI research that could help competitors develop their own systems, without telling users. The company reversed the feature after criticism. The episode still showed that a model can quietly withhold help.
Safety also needs limits
The strongest case for refusal is clear: open-source models that comply with dangerous requests are a poor foundation for a safe future. Grok, designed with fewer restrictions, has been used to create many intimate images of people without their consent. But relying on a model’s refusal as the last line of defense is a fragile arrangement—especially as AI enters power grids, transportation, education and military command and communications systems.
There is a second uncertainty beyond deliberate jailbreaks: models may refuse in ways nobody intended. In 2025, the UK AI Security Institute found that Anthropic models sometimes declined tasks in AI safety research, even though they had not been specifically trained to do so. Three models refused more than half the tasks Anthropic had designated as a set of reasonable AI safety research tasks. Later models showed less of this behavior, but Anthropic could not eliminate it completely.
Anthropic also found that a version of Mythos trained only to be helpful hesitated over some prompts anyway. Asked about synthesizing a virus, it seemed to consider whether helping would be dangerous. The model’s hesitation was sensible in that case; the unsettling point is that its developers had not trained it to hesitate.
I think the industry has not yet resolved what counts as a successful refusal. A system that blocks a dangerous request is safer in one sense; a system that silently withholds legitimate information, or decides for itself when to disobey, creates a different control problem.
The harder AI becomes to avoid, the less plausible it is to treat refusal as a universal safety mechanism. Filters can miss harmful requests, overblock benign ones and reflect rules users never see. If critical infrastructure and other systems depend on models refusing at exactly the right moment, the boundary will need to be enforced by more than the model’s willingness to say no.
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X