Join our Newsletter — 33% off our NHI Course

What is the difference between a neuron and a learned feature in mechanistic interpretability?

A neuron is an individual computational unit inside the network, while a learned feature is a direction in activation space that captures a more coherent pattern. Neurons can participate in several concepts at once, but a feature aims to isolate one interpretable signal. That distinction matters because features are easier to analyze, steer, and compare across models.

Why the distinction matters in mechanistic interpretability

A neuron is a single unit in the network, but a learned feature is a direction in representation space that can be distributed across many neurons. That means interpretability work often shifts from asking what one neuron “means” to asking whether a pattern is represented cleanly enough to be isolated, measured, and manipulated. The practical difference is about coherence, not just count.

In mechanistic interpretability, the question is whether the model stores a concept as a sparse, reusable signal or as a tangled mixture spread across many internal units. Neurons can be polysemantic, while features are usually treated as more stable candidates for analysis because they align better with the underlying geometry of the activation space.

This distinction matters because the object you choose to inspect changes the kind of explanation you can recover. A neuron-level view can reveal local activation behaviour, but a feature-level view is often better for tracing circuits, comparing representations across layers, and testing whether a concept is causally useful rather than merely correlated.

How neurons and learned features differ in practice

Neurons are coordinates in a learned basis, so a single neuron may respond to multiple unrelated patterns depending on context. Learned features are not tied to one coordinate in that basis; they are inferred from the directions that best capture recurring structure in the model’s activations. In modern interpretability work, this is why methods such as sparse autoencoders are used to uncover candidate features rather than relying only on raw neuron inspection.

The distinction also explains why neuron interpretability has limits. If one neuron participates in several concepts, then looking at that neuron alone can hide the model’s true internal organisation. A learned feature can be easier to reason about because it may correspond more closely to one interpretable signal, even if that signal is still approximate rather than perfectly pure.

For practitioners, the key question is whether the representation is basis-dependent. Neurons are basis-specific objects, but learned features aim to capture patterns that remain meaningful even when the internal coordinates are a poor human explanation for what the model is doing.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

NIST AI RMF provides the primary governance reference for this topic.

Framework Control / Reference Relevance
NIST AI RMF Govern Mechanistic interpretability is part of AI governance and model risk analysis.
Recommendation — Use governance processes to document how internal representations are interpreted and validated.

Practitioner Guidance

What to verify: If you are comparing a neuron with a learned feature, check whether the apparent concept survives across different prompts, layers, or basis choices. If it disappears when the coordinate system changes, it is probably a fragile neuron-level association rather than a robust feature.

Decision rule: Use neuron analysis when you want a local activation readout or a quick diagnostic, but use feature-level analysis when you need a cleaner unit for causal tracing, steering, or cross-model comparison. If the concept is clearly polysemantic, assume neuron inspection will under-explain it.

Common mistake: Treating high activation as equivalent to semantic meaning. A neuron can fire strongly without representing a single clean concept, so interpretability claims should be grounded in activation patterns, causal tests, and feature decomposition rather than one-off examples.

Practitioner takeaway: The useful interpretability unit is not always the smallest unit in the network; it is the unit that most faithfully isolates a causally relevant signal.