PROMO$5 FREE CREDITS

    What is abliteration

    All guides

    Abliteration is a weight edit that reduces a language model's refusal behavior by removing a direction in its activations associated with refusal. The architecture stays the same. The weights change, so the model is less likely to answer with a stock refusal when a prompt sits near that direction. It is not a system prompt, not a jailbreak string pasted into the chat, and not a fine-tune on a new pile of conversations.

    The public account of the method comes from research on refusal directions in residual-stream activations, including Arditi and colleagues' 2024 paper "Refusal in Language Models Is Mediated by a Single Direction," and from the practical recipes the community later called abliteration. Those recipes differ in detail. Some subtract a refusal direction from activations or weights. Some orthogonalize weights so later layers stop writing that direction back in. Refuseless does not publish a refusal-rate number for its hosted checkpoints, and this page does not invent one. What we do publish is a hosted API for a named lineup of open-weight models whose refusal behavior has been removed from the weights.

    What the edit changes

    A base model already contains the capability to discuss a topic and a tendency to refuse some of those discussions. Abliteration targets the tendency. The usual picture, simplified so a working engineer can use it, has three steps.

    1. Run two sets of prompts through the base model: ones it answers, and ones it refuses.
    2. Read activations, often in the residual stream, and find a direction that separates the two sets. That direction is the refusal vector.
    3. Edit the weights so that direction is weakened or prevented from being written. Orthogonalization is one way to do the edit. A direct ablation of the direction is another.

    After the edit, prompts that used to die on a refusal have a better chance of receiving a normal completion. Prompts the base model already answered should keep working, which is the point of editing one direction instead of retraining the whole network. "Should" is doing real work in that sentence. A bad edit can dull the model, make it ramble, or leave refusals in place. A page that quotes a benchmark here without saying whose checkpoint it measured is not describing abliteration. It is advertising.

    What abliteration is not

    • Not a jailbreak. A jailbreak is a prompt. It tries to talk a frozen model into ignoring its own instructions. The weights never change, so the next session can refuse again. Abliteration is a change to the file you load.
    • Not a fine-tune. Fine-tuning continues training on examples. It can teach a new skill, a new tone, or a new domain. Abliteration does not add examples. It removes a direction that was already there.
    • Not a system prompt. A system prompt is text in the request. You can change it per call. An abliterated model keeps the edit whether or not you send a system prompt.
    • Not a content policy. A hosted API can still refuse a request at the gateway, log it, or bill it. The weight edit and the operator's rules are different layers. Refuseless is the weight-edit layer: a named lineup, an OpenAI-compatible endpoint, and the retention position stated on the homepage. We do not sell a policy gateway.

    How you call an abliterated model

    On Refuseless the call is the chat-completions request you already write. The base URL is https://api.refuseless.com/v1. The key is a Refuseless API key, sent as a bearer token. The model id is one from the lineup, not a guess. GET /v1/models returns the ids the account can call. Pricing is per million tokens on the pricing page. Card and crypto settlement are listed there. This page does not restate rates, because rates move and the pricing page is the source.

    If you want the vocabulary behind the edit, start with refusal vector, orthogonalization, and residual stream. If you want the practical contrast with the other ways people try to change refusal, read the four-way comparison.

    Get an API keySee the lineup