Abliteration and the Risks Ahead

Published on Sep 2, 2026

A few days ago, an abliterated version of GLM 5.3 was released publicly, positioned for offensive cybersecurity, and made available to anyone willing to pay.

Abliteration is particularly concerning when applied to cybersecurity safeguards. A capable model already understands vulnerabilities, exploitation, persistence, malware, and many of the techniques involved in an offensive operation. Cyber safety training places boundaries around when and how those capabilities can be used. Abliterating those boundaries can make the underlying capability considerably easier to access.

There is growing evidence that the effects of abliteration are difficult to contain to the behavior someone intended to change. It does not simply remove a narrow refusal mechanism. By modifying weights involved in safety-related decision-making, abliteration can introduce broader misalignment in ways that are difficult to predict or contain. And modifying refusal-related weights for cyber may not stay confined to cyber. It could weaken a model’s restraint in domains like biology, chemistry, or weapons development as well.

My brother is a neurosurgeon, and abliteration reminds me a little of my conversations with him about brain surgery. The brain is an enormously interconnected system. Changing or damaging even a very small area can affect behavior in ways that are difficult to predict in advance. A surgeon may understand the primary function associated with a particular pathway, but that does not mean every downstream interaction is completely understood.

That is especially concerning in cyber because models are beginning to show that they can deviate from our instructions even after significant safety testing. During OpenAI’s recent cybersecurity evaluations, models operating with reduced safeguards found ways around network isolation, exploited vulnerabilities in shared infrastructure, gained unauthorized internet access, communicated through channels they were never given, and ultimately compromised parts of OpenAI’s research infrastructure and Hugging Face. OpenAI described the behavior as misaligned with the tasks the models had been given and called the incident a “warning shot.”

Timing matters enormously. Even a few months in which the strongest cyber capabilities remain harder to weaponize gives defenders time to find vulnerabilities, harden infrastructure, improve containment, build better monitoring, and deploy equally capable defensive systems. OpenAI itself has slowed parts of its frontier development and strengthened alignment and containment after seeing what highly capable cyber agents can already do. We highly applaud this effort.

At depthfirst, we think defenders should use this window aggressively. That requires applying security intelligence continuously and at enormous scale. It is a major reason we invest so heavily in training models such as dfs-large1 alongside using the best frontier models: strong vulnerability discovery and exploitability analysis need to be economical enough to run across far more of the software surface.

Think about what defenders are now responsible for continuously finding: vulnerabilities, malware, exploitable bugs, supply-chain threats, and increasingly, insecure or even misaligned AI-generated code. They have to do all of this without breaking the bank.

The gap between offensive and defensive AI is narrowing quickly. We should assume powerful offensive capabilities will eventually be widely available. But giving up safeguards before defenders are ready only accelerates that timeline.

The cybersecurity industry needs to help every organization on the planet find and remediate vulnerabilities, and deploy appropriate safeguards. This is why we collaborate with frontier labs and other partners to make it easier for organizations to deploy frontier security.