The claim checked

We truly do not yet know how AI gets its answers internally, which is why mechanistic interpretability is needed.

What holds up

Mechanistic interpretability is a real research area focused on reverse-engineering neural networks’ internal computations. Large language models remain difficult to interpret comprehensively, even when their architecture and training methods are known. Research has nevertheless identified some interpretable features, representations, and causal computational pathways.

What does not

The wording can sound more absolute than the evidence warrants: researchers are not wholly ignorant of model internals. They understand the designed algorithms and have made partial, experimentally supported discoveries about specific behaviors. Also, “thinking” is an informal metaphor for computation, not evidence that AI has a mind.

Why it matters

The post’s central message is that comprehensive understanding remains incomplete, not that researchers know literally nothing. Its brevity does not materially reverse that takeaway, though viewers should understand that the field has already produced partial explanations.

Why Clear says this

Peer-reviewed surveys describe LLM internal mechanisms as unclear and identify interpretability as an active research challenge. At the same time, mechanistic-interpretability work has produced concrete findings by tracing and intervening on internal representations and circuits. Together, this supports the post’s broad point while limiting it to incomplete understanding rather than total mystery.

Evidence

  • A peer-reviewed ACM survey concludes that the internal mechanisms of large language models remain unclear and reviews methods intended to explain them.
  • A review of mechanistic interpretability defines the field as reverse-engineering internal neural-network computations, while noting both novel insights and continuing gaps.
  • Published interpretability research demonstrates that some hidden representations can be inspected and causally tested, so the field has partial—not comprehensive—understanding.
  • Anthropic’s circuit-tracing work reports partial internal traces for particular model outputs and behaviors; this is evidence of progress, not a complete account of how AI works.

Sources used