Key Takeaways:

Heretic is a pip-installable, fully automated tool that removes refusal conditioning from any transformer-based language model using Optuna-optimized directional ablation – producing decensored models that lose less of the original model’s intelligence than anything else in the field.

What Is Heretic?

Heretic is an open-source command-line tool created by Philipp Emanuel Weidmann that removes safety alignment – commonly called “censorship” – from transformer-based language models. It does this without retraining, without fine-tuning datasets, and without requiring any understanding of transformer internals on the part of the user. It was recognized as the number-one repository of the day on Trendshift following its release.

The technique Heretic implements is known as directional ablation, or “abliteration” – a method originating in the research paper Arditi et al. 2024, which demonstrated that refusal behavior in large language models is mediated by a single identifiable direction in the model’s residual stream. Remove that direction from the weight matrices, and the model stops refusing. Heretic automates every step of that process and adds a sophisticated parameter optimizer on top to minimize capability damage.

The project is licensed under the GNU Affero General Public License v3 and is intended for researchers, developers, and practitioners who run local models and require full control over their behavior.

The Technology Behind Heretic

Directional Ablation and Refusal Directions

When a language model declines to answer a prompt, that refusal emerges from a specific geometric direction in the model’s internal representation space. Heretic computes these “refusal directions” for each transformer layer by taking the difference-of-means between the first-token residual hidden states produced by “harmful” example prompts and “harmless” example prompts.

Once a refusal direction is identified for a layer, Heretic orthogonalizes the relevant weight matrices – specifically the attention out-projection and MLP down-projection matrices – with respect to that direction. This prevents the direction from being expressed in the output of those matrix multiplications, effectively suppressing the model’s ability to generate refusals without altering anything else.

The Optuna Parameter Optimizer

What distinguishes Heretic from earlier abliteration implementations is its use of a Tree-structured Parzen Estimator (TPE) optimizer, powered by Optuna, to automatically search for the best ablation parameters. The optimizer simultaneously minimizes two objectives: the number of refusals produced for a set of “harmful” evaluation prompts, and the KL divergence from the original model measured on “harmless” prompts. The second objective is what protects the model’s general capability.

The ablation process is controlled by several tunable parameters per transformer component:

  • direction_index: A float index into the space of refusal directions. Non-integer values cause the two nearest refusal direction vectors to be linearly interpolated, which dramatically expands the search space beyond what prior methods could explore.
  • max_weight and max_weight_position: Define the peak of the per-layer ablation weight kernel.
  • min_weight and min_weight_distance: Define the floor and decay shape of the kernel.

Each of these is optimized separately for the attention component and the MLP component, as MLP interventions tend to be more damaging to model capability than attention interventions.

Installation and Setup

Prerequisites

Heretic requires the following:

  • Python 3.10 or newer
  • PyTorch 2.2 or newer, installed appropriately for the target hardware (CUDA for NVIDIA GPUs, ROCm for AMD, MPS for Apple Silicon, or CPU-only)
  • Sufficient VRAM to load the target model, or bitsandbytes for quantized loading

Installing Heretic

Installation is a single command:

pip install -U heretic-llm

No additional configuration is required to run Heretic in its default mode.

VRAM Requirements and Quantization

At the start of each run, Heretic benchmarks the system to determine the optimal batch size for the available hardware. Running Heretic on a model the size of Llama-3.1-8B-Instruct takes approximately 45 minutes on an RTX 3090 with the default configuration.

For users with limited VRAM, Heretic supports 4-bit quantization via bitsandbytes. Enable it by setting the quantization option to bnb_4bit in the configuration file or passing it as a command-line flag. This can reduce VRAM requirements substantially, making it possible to process models that would otherwise exceed available memory.

Running Heretic

Basic Usage

The minimal usage pattern is:

heretic <model-id>

Where <model-id> is any Hugging Face model identifier. For example:

heretic Qwen/Qwen3-4B-Instruct-2507
heretic meta-llama/Llama-3.1-8B-Instruct
heretic google/gemma-3-12b-it

Heretic downloads the model if it is not already cached locally, runs the optimization loop, and presents results automatically when complete.

Configuration Options

Heretic supports two configuration methods:

Command-line flags – run heretic --help for a full list of available options. These are suitable for one-off runs and scripting.

TOML configuration file – copy or reference config.default.toml from the repository for a structured, persistent configuration. This is the recommended approach for repeated use or reproducible research workflows.

Relevant configuration options include quantization mode, the number of Optuna optimization trials, the sets of “harmful” and “harmless” evaluation prompts used for computing refusal directions and measuring KL divergence, and per-component ablation weight constraints.

After Decensoring

When the optimization run completes, Heretic presents several options:

  • Save the model to a local directory in Hugging Face-compatible format
  • Upload the model directly to Hugging Face Hub
  • Chat with the model interactively in the terminal to evaluate its behavior before committing to save or upload
  • Any combination of the above

The chat interface allows immediate qualitative testing without needing to load the model in a separate tool.

Research and Interpretability Features

Install the optional research extras to access advanced interpretability functionality:

pip install -U heretic-llm[research]

Plotting Residual Vectors

Pass --plot-residuals to generate visualizations of how the model’s internal representations differ between “harmful” and “harmless” prompts at every transformer layer. The process uses PaCMAP to project the high-dimensional residual vectors into 2D space, produces a PNG scatter plot for each layer, and assembles all layers into an animated GIF showing how the representations evolve through the network. The PaCMAP computation runs on CPU and can take an hour or more for larger models.

Residual Geometry Analysis

Pass --print-residual-geometry to produce a per-layer table of quantitative metrics describing the geometric relationship between “harmful” and “harmless” residuals. Metrics include cosine similarities between mean and geometric median vectors, L2 norms, and the mean silhouette coefficient of the two clusters. This is useful for identifying which layers carry the most refusal signal and understanding how cleanly the two classes separate in representation space.

Built-in Model Evaluation

Pass --evaluate-model <model-id> to run Heretic’s evaluation suite against any already-processed model. This allows reproducing the benchmark figures reported in the project’s documentation and comparing the output of different abliteration approaches on the same evaluation set.

Performance Against Competing Methods

Heretic’s benchmark results for google/gemma-3-12b-it illustrate its primary advantage. All three abliteration methods tested achieve the same refusal suppression rate – 3 refusals out of 100 “harmful” prompts, compared to 97 for the unmodified model. The difference lies in capability preservation, measured as KL divergence from the original model on “harmless” prompts:

VersionRefusals (out of 100)KL Divergence
Original (google/gemma-3-12b-it)970
mlabonne/gemma-3-12b-it-abliterated-v231.04
huihui-ai/gemma-3-12b-it-abliterated30.45
p-e-w/gemma-3-12b-it-heretic30.16

The Heretic version achieves a KL divergence of 0.16 – less than one-sixth the damage of one competitor and less than half the damage of the other – while matching their refusal suppression performance entirely. This result was produced automatically, with no human-guided parameter selection.

The Heretic collection on Hugging Face includes ready-to-use decensored models, and the community has published well over 1,000 additional Heretic-processed models.

Supported Models

Heretic supports most dense transformer architectures, including a wide range of multimodal models, and several mixture-of-experts (MoE) architectures. It does not currently support state space models (SSMs) or hybrid SSM-transformer architectures, models with inhomogeneous layer configurations, or certain novel attention mechanisms. The repository’s issue tracker and release notes are the best place to track ongoing compatibility additions.

Use Cases

Heretic is applicable across a range of research and development scenarios:

  • AI safety and interpretability research: Understanding and empirically measuring refusal mechanisms, their geometric properties, and how they interact with model capability is directly relevant to AI alignment work. Heretic’s --plot-residuals and --print-residual-geometry flags are purpose-built for this.
  • Reproducible abliteration research: The built-in evaluation function and TOML configuration system allow other researchers to reproduce results and compare methods against a consistent benchmark.
  • Local model customization: Developers and power users running models locally on personal hardware may need models that respond to the full range of queries relevant to their domain, including those that trigger overly broad refusal behavior in base instruct models.
  • Comparing alignment techniques: The KL divergence metric and per-layer analysis tools allow practitioners to quantitatively assess how much different alignment techniques have altered the base model’s behavior and where in the network those changes are concentrated.
  • Building uncensored model collections: Researchers and communities publishing on Hugging Face can use Heretic’s direct Hub upload feature to generate and distribute decensored variants of public models at scale.

Conclusion

Heretic is technically the most capable automated abliteration tool available. Its combination of flexible per-layer ablation weight kernels, interpolated refusal direction vectors, component-separated optimization, and Optuna-driven parameter search produces decensored models that are measurably closer to their original capability profile than those produced by any prior published method – and it does this without requiring the user to understand or configure a single transformer internal.

For researchers studying how safety alignment is encoded in model weights, for developers who need full behavioral control over locally hosted models, and for anyone benchmarking abliteration methods, Heretic is the current reference implementation.

Repository: https://github.com/p-e-w/heretic
Hugging Face Collection: https://huggingface.co/collections/p-e-w/the-bestiary
Community Models: https://huggingface.co/models?other=heretic
Original Abliteration Paper: Arditi et al. 2024
License: GNU AGPL v3

You may also like

Subscribe
Notify of
guest

0 Comments
Newest
Oldest Most Voted