Writing

Notes on interpretability and steering, posted on LessWrong.

Exploring Generalization in NLA's

Reproduction and follow-on experiments on Anthropic's neural network activation work, focused on training architecture and generalization as they relate to steering.

Jun 2026

Calibrating Activation Vectors using Norm

Activation steering usually hand-tunes injection magnitude. This post looks at calibrating vector norms directly instead of applying a uniform scale across layers.

Jun 2026