Heretic
Heretic is an open-source command-line tool for removing restrictions, also called safety alignment, from transformer-based language models. It is designed to reduce refusals while preserving as much of the original model’s capability as possible, according to its publisher.
What it does
Heretic combines directional ablation, or abliteration, with a TPE-based parameter optimizer powered by Optuna. The tool automatically searches for parameters by minimizing refusals while limiting divergence from the original model. This workflow is intended to remove censorship without requiring users to understand transformer internals.
Users install the heretic-llm package and provide a model identifier, such as a model hosted on Hugging Face:
sh
pip install -U heretic-llm
heretic Qwen/Qwen3.5-4B
Heretic supports most dense models, including many multimodal models, several mixture-of-experts architectures, and some hybrid architectures. The publisher says that pure state-space models and certain other research architectures are not supported out of the box.
Notable capabilities
- Automatic parameter search using a TPE-based optimizer powered by Optuna.
- Directional ablation, also known as abliteration, for modifying model behavior without expensive post-training.
- Support for many dense, multimodal, mixture-of-experts, and hybrid model architectures.
- Optional bitsandbytes quantization to reduce VRAM requirements.
- Post-processing options that include saving the model, uploading it to Hugging Face, chatting with it, and running standard benchmarks.
Who it helps
Heretic is aimed at developers, researchers, and technically capable users who want to modify the refusal behavior of compatible language models locally. It may suit users exploring model behavior, alignment, and representation geometry, or those who want an automated alternative to manually tuning ablation parameters.
How it fits a workflow
A typical workflow starts with installing Python 3.10 or newer and PyTorch 2.2 or newer, then passing a model identifier to the command-line tool. After processing, users can save the modified model or use it for chat and standard benchmarks. Optional research features can generate residual-vector plots and geometry analyses. The project is free software released under the GNU Affero General Public License version 3 or later; no commercial pricing is documented.
Strengths and limits
The main practical benefit is automation: Heretic searches for parameters and is intended to preserve model capability while reducing refusals, rather than requiring users to work directly with transformer internals. Nvidia and AMD GPUs are well supported, while CPU-only processing is available but substantially slower and intended only for tiny models. The CPU-based projection step used by some research features can take an hour or more for larger models. Compatibility also varies by architecture, and particular models may require later PyTorch features than the documented minimum.
