Bryan Kwan (author) and Andrew Jun Lee (mentor)
Traditional frameworks in mechanistic interpretability assume that AI models process abstract concepts as simple, straight lines in data space. This post critiques this brittle assumption by exploring how language models actually rely on continuous geometric shapes, landscapes, and twisting surfaces embedded within their internal pathways. By looking past rigid components, we demonstrate that language model computation is explicitly driven by smooth, multi-dimensional geometric transformations over low-dimensional submanifolds.
1. The Limits of Discrete Circuit Paradigms
Large language models have achieved staggering success, but because they shift billions of numbers through hidden pathways, they operate as an unreadable black box. This is where mechanistic interpretability comes in. Think of it as a form of digital neuroscience. It is the field of research dedicated to taking a microscope to the AI’s inner workings, peeling back the layers, and trying to map out the exact internal mechanisms that allow a model to form beliefs, track logic, and make decisions.
This field has made remarkable strides, yet many of our leading explanations stop short of discovering the full truth. Consider the popular paradigm of “circuits”. Traditional circuit analysis treats a language model like a highly specialized editorial team inside a massive, fast-paced global newsroom. In this setup, individual attention heads and neurons act as discrete nodes. You can think of these nodes as specialized editors dedicated to parsing syntax, tracking factual relationships, or modifying writing tone. The wires between them represent physical communication paths. Together, they form a rigid, step-by-step assembly line on a whiteboard.
While these wiring diagrams trace where information flows, they offer an incomplete picture. A transformer does not pass neat, isolated memos from desk to desk. Instead, the entire model operates over a single, massive shared digital document otherwise known as the residual stream highway. Every internal component modifies this shared space simultaneously. Rather than executing static routing, the nodes collectively twist, mold, and reshape a fluid, continuous landscape of context in real-time.
Similarly, looking only at individual activation triggers inherits a structural blind spot. These methods default to the Linear Representation Hypothesis (LRH), assuming abstract concepts map onto flat, isolated, one-dimensional directions. This is equivalent to assuming specialized editors only look for single, isolated keywords. It ignores the reality that language models process a multi-dimensional vibe, tone, and subtext that cannot be pinned down to a single flat line. To truly understand language model behavior, we must look past discrete circuits and characterize how these networks continuously manipulate spatial geometry to reason.
What is a Transformer anyway?
To understand how these internal editors work, it helps to understand what a transformer model actually does. At its core, a transformer is built to predict the very next word in a sequence. It reads a sentence token by token, but instead of analyzing words in isolation, it processes them using a system called attention.
Think of a transformer as an interactive, live-shared digital document. When a sentence is typed into the model, every word gets its own entry line on this master document (the residual stream highway). The model’s attention mechanism acts like a team of readers scanning the document simultaneously. For instance, if the sentence is “The bank of the river was muddy,” the word “bank” is inherently ambiguous. The attention mechanism fixes this by forcing the word “bank” to look at its surrounding context. It highlights the word “river,” pulls the concept of water, and writes a margin note back onto the entry line for “bank,” updating its meaning in real-time. By dynamically shifting focus based on context, attention ensures that words are constantly updating their values based on who they are sitting next to.
From Words to Coordinates: Tokens and Vectors
But how does a word actually exist inside a computer? Computers don’t read letters; they read numbers. When you input text, the AI first slices it into fundamental building blocks called tokens. Instead of parsing full words, the model breaks text into smaller, flexible pieces like common syllables, prefixes, or punctuation marks. For example, the word “transformer” might be sliced into three distinct pieces: “trans,” “form,” and “er.” This modularity allows the model to handle brand-new words or complex grammar it has never encountered before.
Once the text is broken into tokens, the model translates each token into a long list of numerical values called a vector. You can think of these numbers as a highly detailed set of GPS coordinates that places the token somewhere inside a massive, multi-dimensional map known as data space. Tokens with similar conceptual meanings are dropped close together on this map. For instance, the coordinate for “puppy” will sit right next to “dog,” while the coordinate for “concrete” will be miles away. Every piece of text the AI processes is just a point, a coordinate, or a trajectory moving through this invisible, hyper-dimensional territory.
Geometric Intuition of the Residual Stream
Instead of visualizing a static calculation, think of an AI model’s internal state as a continuous journey traveling through this high-dimensional territory. At each layer, individual mechanisms do not rewrite the entire document. Instead, they act as continuous spatial forces that nudge, bend, or curve the ongoing text trajectory toward new semantic destinations.
The traditional assumption is that the model tracks independent facts by pointing vectors down independent, straight lines. However, a geometric framework recognizes that data points cluster onto curved landscapes called submanifolds. Within these curved surfaces, interconnected features vary smoothly alongside one another, tracking properties like time, context, and intent as a single, unified geometry rather than separate isolated points.
2. Inductive Subspace Biases
Empirical evidence of these geometric structures begins with the remarkably low dimensionality large language models use to encode complex concepts. Transformers exhibit an intrinsic bias toward allocating data into clean, independent, right-angle compartments rather than scattering it randomly across the full capacity of their internal memory space.
Think of this intrinsic bias like an ultra-efficient moving company packing a massive, chaotic house into a storage truck. Instead of tossing items randomly into the 3D space, the movers instinctively flatten the cargo into independent, right-angle categories. Kitchen items go strictly along the floor, books strictly along the left wall, and clothes along the right. By sorting complex concepts into these independent, spatial “drawers,” the model makes messy data highly organized and accessible.
Even when a data-generating process is completely tangled, a model will initially default to a simple, factorized representation early in training, only expanding into complex shapes when accuracy pressure forces it to adapt. If the house contains a hybrid object, like a cookbook that logically belongs to both the kitchen and the library, the movers initially default to the simplest solution. On day one, they lazily throw it into the flat book pile. It is only when the customer complains that a chef can’t find the recipes that the movers are forced to adapt. To recover accurate fidelity, they finally expand their strategy, lifting the object into a higher-dimensional, more complex position that bridges both categories. Dimensionality analysis consistently confirms this behavior across intermediate layers, proving that complex concepts are naturally compressed into flat, clean local neighborhoods.
How Models Separate Parallel Realities
The real world is an overlapping combination of independent tracks: we track grammar, subject matter, and timeline context simultaneously. To keep these tracks from clashing, the model untangles them using a clean spatial layout.
Instead of treating these attributes as a giant multiplied matrix, the AI isolates each independent variable into its own independent spatial drawer. Because these drawers sit at perfect right angles to each other in the model’s environment, changing a word’s tense does not accidentally corrupt its subject category. This clean spatial separation allows the network to drastically compress an exponential world down to a fraction of the size, maintaining perfect accuracy with a tiny geometric footprint.

The geometry of factored subspaces, illustrating how transformers naturally compress complex data into organized, independent spatial compartments to maximize efficiency.
3. Topological Convergence and Field-Aware Probing
Further evidence emerges in the discovery that internal representation shapes mirror the physical topologies of the concepts they explain. Internal representations naturally converge on intuitive geometries. Temporal concepts (like days of the week or dates) form circular loops, chronological durations form straight lines, and geographical city networks form highly structured spherical globes inside the data space.
These structures directly govern how networks manipulate information. Traditional linear steering, which attempts to manipulate features along flat, one-dimensional lines, has proven profoundly ineffective compared to structure-aware steering. When researchers shift an internal state using a raw, straight vector, the model’s internal variance spikes out of control. Imagine a high-speed bullet train traveling along a smooth track that winds through a steep mountain pass. This curved track is our continuous geometric manifold. Linear steering assumes the model’s brain is flat, which is equivalent to locking the train’s wheels perfectly straight while it is tearing around a sharp mountain bend. Because a rigid, flat line cannot trace a curved surface, the linear force derails the train completely, injecting chaotic noise into the concept.
In contrast, field-aware steering acts as a digital stabilizer. By learning the smooth, local slopes of the concept surface, it creates a noise-regularized coordinate map that tracks the curvature. This guides the intervention smoothly along the natural, winding terrain of the model’s representation space, altering the concept perfectly without destabilizing the system.
4. Behavioral Alignment and Subspace Interventions
The geometric layout of internal representations allows us to directly predict downstream behavior. By mapping the underlying shape of these internal concept landscapes, we find that the shortest path across the curved neural surface predicts final behavioral choices and output correlations. For instance, similarity analysis reveals a tight correlation ($r = 0.92$) between the layout of internal landscapes and the model’s explicit text probabilities.
These topological shapes are causally tied to reasoning accuracy. When researchers introduce targeted noise interventions restricted entirely to the smooth surface where a concept lives, the model’s performance collapses. Conversely, injecting an identical volume of random noise outside of that surface has a negligible effect on performance. The transformer’s capacity to solve tasks is directly bound to the preservation of these internal spatial geometries.
5. Dynamic Adaptation and Context Graph Evolution
Dynamic adaptation across model depth and evolving context further highlights why understanding geometry is vital for interpretability. When models are exposed to relational structures within a prompt, such as a graph-tracing puzzle, the internal data points physically reorganize to mimic the exact structural map of the prompt.
However, this geometric transformation is not uniform across the model’s layers. Early layers remain anchored to structural training memories, mapping semantically related words close together based on general pre-training knowledge. By the final layers, the transformer completely overrides these rigid memories to align with the unique architecture of the context, forcing initially similar concepts to pull apart if the prompt demands it.
This reconstruction is tracked using spatial smoothness metrics, which measure how closely the distances between internal data points reflect the prompt layout. As the context size expands, the model minimizes its structural tension, physically drawing contextually adjacent concepts closer together in its space. This smooth, layer-by-layer update traces a predictable trajectory across the surface, proving that the model actively warps its internal map to fit the prompt, rather than relying on a simple memorization strategy.

Visualizing the low-dimensional structural transformations of representation clusters across active layers via Principal Component Analysis.
6. Case Study: When Models Manipulate Manifolds

Mechanism of continuous rotary adjustments where query-key parameters rotate spatial coordinates inside block submanifolds.
A definitive case study highlighting the necessity of geometric analysis is found in the landmark project, When Models Manipulate Manifolds. In analyzing how an LLM decides when to emit a newline token (\n) versus a standard text word, researchers discovered that the model completely rejects flat, isolated feature vectors. Instead, it encodes continuous properties like character counts and line margins as low-dimensional, circular loops in representation space.
Multi-dimensional representations are necessary to handle open-ended tracking cleanly. Instead of vector lengths blowing up to infinity as numbers count higher, a circle keeps the activation magnitude perfectly constant and bounded. The transformer converts this circular map into a dynamic calculator using its attention mechanism. The internal query and key functions act like high-dimensional spatial rotation wheels. As text accumulates, the model rotates and twists the count manifold in the internal stream. When the remaining space matches the target line limit, the rotated landscapes perfectly align and overlap. This geometric intersection triggers an intentional alignment spike, instantly forcing the attention mechanism to prioritize formatting a newline break.
7. The Future Landscape of Geometric Interpretability Frameworks
At the end of the day, continuous manifolds are often regarded as a pure visualization luxury for researchers, but biologically, they are the functional tools that language models rely upon to execute complex reasoning tasks. Fulfilling the potential of mechanistic interpretability demands transitioning from mapping static routing circuits to mapping continuous vector landscapes and fluid trajectories.
We must develop extraction tools like Linear Field Probes (LFPs) that piece-wise tile curved spaces with localized linear readouts, approximating target semantics without flattening or destroying the underlying terrain of the residual stream. Only by tracking how these multi-dimensional shapes are continuously structured, rotated, and manipulated across layer transitions can we hope to build a faithful and complete account of how artificial networks truly compute.
References
- Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., DasSarma, N., Dawn, D., Drain, D., El-Showk, S., Fort, S., Hatfield-Dodds, Z., Hernandez, D., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., & Olah, C. (2021). A mathematical framework for transformer circuits. Transformer Circuits Thread, 1(1), 1.
- Engels, A., Isaacs, J., AlKhamis, A., & Tegmark, M. (2024). Not All Language Model Features Are One-Dimensionally Linear. arXiv preprint arXiv:2403.11984.
- Gurnee, W., & Tegmark, M. (2023). Language models linearly encode space and time. arXiv preprint arXiv:2310.02207.
- Karkada, D., Korchinski, D. J., Nava, A., Wyart, M., & Bahri, Y. (2026). Symmetry in language statistics shapes the geometry of model representations. arXiv preprint arXiv:2602.15029.
- Lee, A., Weber, M., Viégas, F., & Wattenberg, M. (2025). Shared global and local geometry of language model embeddings. arXiv preprint arXiv:2503.21073.
- Shai, A., Amdahl-Culleton, L., Christensen, C. L., Bigelow, H. R., Rosas, F. E., Boyd, A. B., Alt, E. A., Ray, K. J., & Riechers, P. M. (2026). Transformers learn factored representations. arXiv preprint arXiv:2602.02385.
- Modell, A., Rubin-Delanchy, P., & Whiteley, N. (2025). The origins of representation manifolds in Large Language Models. arXiv preprint arXiv:2505.18235.
- Park, J., Kim, H., & Lee, S. (2025). In-context learning of representations. International Conference on Learning Representations (ICLR 2026). OpenReview forum ID: pXlmOmlHJZ.
- Gurnee, W., Ameisen, E., Kauvar, I., Tarng, J., Pearce, A., Olah, C., & Batson, J. (2026). When Models Manipulate Manifolds: The Geometry of a Counting Task. arXiv preprint arXiv:2601.04480.
- Zhou, Y., Wang, Y., Yin, X., Zhou, S., & Zhang, A. R. (2025). The geometry of reasoning: Flowing logics in representation space. arXiv preprint arXiv:2510.09782.