Interpretability
- Question
- Which internal components of a small language model carry a concept, and can they be changed?
- Status
- One-afternoon replication on a single GPU, with a written report.
Logit-lens, lens-vector, ablation and coefficient-swap experiments on a 4-billion-parameter open language model, following a published method.
- Reproduced the published reading of hidden concepts with a cheap averaged-gradient lens.
- Found a band of layers where ablation changes the answer, distinct from the layer where the lenses peak.
- Found a boundary condition: concepts that the model has to infer are ablatable, while concepts restated in the prompt repair themselves.
- Steering held on the first example and collapsed on the second. Intervention results were mixed.
- Verified an indexing pitfall in the hidden-state outputs before trusting any result.