When a model explains its reasoning, it isn't reporting what happened inside. It's generating more text.
How inference mechanically works in a modern LLM — the forward pass, the KV cache, sampling, quantization and what each one actually costs — and what interpretability research can and cannot currently tell you about why a model produced an output. Part 2 of our technical-depth series.