Donald Murre

← Research

Interpretability

Question
Which internal components of a small language model carry a concept, and can they be changed?
Status
One-afternoon replication on a single GPU, with a written report.

Logit-lens, lens-vector, ablation and coefficient-swap experiments on a 4-billion-parameter open language model, following a published method.

Other research areas