Steer Features

Use this skill for feature-level steering of models — locating the internal feature that drives a target behavior, scoring and selecting it by its effect on the model's output, and directly amplifying or shrinking that feature's activation during generation to control behavior. Applies to features read from the model's own activations or from a Sparse Autoencoder (SAE); the bundled demo scripts happen to use an SAE, but the method does not require one.

zjunlp fca2361 5 files · 50.3 KB Updated

File contents

zjunlp/mechanist/tree/main/skills/mechanism-skills/representation-and-parameter-analysis/steer-features commit fca23610b6

Frequently asked questions

npx skillmds@latest add zjunlp/steer-features