Obliteratus

Remove refusal behaviors from open-weight LLMs using OBLITERATUS — mechanistic interpretability techniques (diff-in-means, SVD, whitened SVD, LEACE, SAE decomposition, etc.) to excise guardrails while preserving reasoning. 9 CLI methods, 28 analysis modules, 116 model presets across 5 compute tiers, tournament evaluation, and telemetry-driven recommendations. Use when a user wants to uncensor, abliterate, or remove refusal from an LLM.

photonics-dhl c27f38a 6 files · 31.4 KB Updated

File contents

photonics-dhl/scholar-s-tea/tree/main/hermes-home/hermes-agent/skills/mlops/inference/obliteratus commit c27f38afc3

Frequently asked questions

npx skillmds@latest add photonics-dhl/obliteratus