Interpretability

Mechanistic interpretability techniques that may reshape agent reliability and safety.

Interpretability research aims to make internal model computations legible: tracing circuits, extracting features, and probing representations.

For agents, interpretability matters because “reasoning traces” may be unfaithful: what the model says it is doing is not necessarily what it computes.