transformer-lens-interpretability

orchestra-research/transformer-lens-interpretability · Agent Skill (multi-file)

by Orchestra Research · bundle

Published · Last updated


Inspect and manipulate transformer internals via HookPoints and activation caching for mechanistic interpretability research.

SKILL.md

Files

This skill is a package of 4 files. Install with the command above, or download the folder.

  • 📄SKILL.md entry
  • 📁references
  • 📄api.md 8.2 KB
  • 📄README.md 1.6 KB
  • 📄tutorials.md 9.9 KB

Related

  1. sparse-autoencoder-training · lord1egypt
    Trains and analyzes Sparse Autoencoders (SAEs) with SAELens to decompose neural network activations into interpretable features, covering loading pre-trained SAEs, training custom ones, and feature steering.
    2
    repo stars
  2. nnsight-remote-interpretability · orchestra-research bundle
    Run interpretability experiments on neural network internals using nnsight, with optional NDIF remote execution for massive models.
    10.4k
    repo stars
  3. bedrock · itsmostafa bundle
    Access AWS Bedrock foundation models for generative AI, including text generation, embeddings, and image generation, with CLI and Python examples.
    1.1k
    repo stars
  4. pyvene-interventions · orchestra-research bundle
    Perform causal interventions on PyTorch models using pyvene's declarative framework for causal tracing, activation patching, and interchange intervention training.
    10.4k
    repo stars
  5. sparse-autoencoder-training · orchestra-research bundle
    Train and analyze Sparse Autoencoders (SAEs) using SAELens to decompose neural network activations into interpretable features for mechanistic interpretability research.
    10.4k
    repo stars
  6. nemo-mbridge-recipe-recommender · nvidia bundle
    Indexes Megatron Bridge recipes and recommends the best starting config based on model, GPU count, and training goal.
    2.2k
    repo stars

Frequently asked questions

How do I install the transformer-lens-interpretability skill?

Run npx skillmds add orchestra-research/transformer-lens-interpretability in your terminal (requires Node.js), paste this page's agent-chat prompt into Claude, Cursor, or any MCP-connected agent, or download the SKILL.md file and copy it into your agent's skills directory.

What does the transformer-lens-interpretability skill do?

Inspect and manipulate transformer internals via HookPoints and activation caching for mechanistic interpretability research. It is listed under AI & ML, Agent Building on SkillMD.

Is transformer-lens-interpretability safe to use?

SkillMD's automated safety review verdict for this skill is CAUTION. Independent scanners report: SkillSpector: PASS, Skill Scanner: PASS. Capability flags: makes network calls, reads secrets. SkillMD never runs a skill's scripts for you; review the SKILL.md before installing.

Which AI agents work with transformer-lens-interpretability?

This skill is tagged as working with Claude Code, Claude.ai, OpenAI Codex. SKILL.md is an open format, so most agents that read a skills directory can load it too.

Is transformer-lens-interpretability free to use?

Yes. Installing skills from SkillMD is free. This skill is licensed under MIT.

Who published transformer-lens-interpretability?

Orchestra Research (@orchestra-research) published this skill. Their other Agent Skills are listed on their SkillMD profile.