Results for “interpretability”
4 skillsMore results
interpret-results
Analyzes evaluation results by requiring a stated hypothesis before examining data, then compares expectations to actual result files to prevent post-hoc rationalization.
0
axiom
Audits hidden assumptions in decisions or beliefs by classifying them into fact, convention, belief, or interest-driven, ranking by fragility and impact, then rebuilding conclusions from verified premises. Supports English and Chinese.
42.4k · bundle
sparse-autoencoder-training
Train and analyze Sparse Autoencoders (SAEs) using SAELens to decompose neural network activations into interpretable features for mechanistic interpretability research.
10.4k · bundle