publications
publications by categories in reversed chronological order. generated by jekyll-scholar.
2026
- The Router’s Mind: Diagnosing Mixture-of-Depths Routers with Sparse AutoencodersStein Pleiter. Graded 9/10. Supervised by Dr. Ana Lucic and Ege Erdogan; examined by Dr. Dolly Sapra. , Jun 2026
Mixture-of-Depths (MoD) routers learn, per token and per layer, whether a transformer block is computed or skipped. Existing work treats the router purely as an efficiency mechanism and evaluates it by perplexity at a target compute budget; the structure of what these routers learn has not been systematically characterised. This thesis adapts token-level Router-Tuning to Gemma 2 2B, where each router gates a token’s attention update, and asks what the resulting routers read from the residual stream. We first establish that the router is not a trivial token-importance ranker: its decisions differ between full- and sliding-window layers and depend on part-of-speech in a layer-conditional way (also reported in a companion workshop paper). We then use Sparse Autoencoders (SAEs) to decompose what the router reads. We find that routing does not reduce to any single interpretable feature; instead a collection of roughly 50-200 sparse features together recovers about half of the router’s score variance on held-out tokens. Classifying those features along a syntactic-to-semantic axis shows that semantic (content) features dominate, and an independent set of automatically generated feature labels agrees. Finally, causal intervention on these features flips routing decisions in the predicted direction at every studied layer (up to a 30.6% conditional flip rate), and a cross-checkpoint comparison shows that the strength of the causal handle is stable across independent trainings while its direction is training-specific. Two further checks reinforce the causal findings: the effect replicates when features are identified and intervened on disjoint corpus halves, and forcing routing decisions measurably changes the model’s output. We also confirm that the routed model retains most of its capability relative to base Gemma 2 2B.
@mastersthesis{pleiter2026routersmind, title = {The Router's Mind: Diagnosing Mixture-of-Depths Routers with Sparse Autoencoders}, author = {Pleiter, Stein}, school = {University of Amsterdam}, type = {Bachelor's thesis}, year = {2026}, month = jun, } - What Do Mixture-of-Depth Routers Learn? Routing Patterns in Gemma 2Stein Pleiter, Ege Erdogan, and Ana LucicIn ICML 2026 Workshop on Mechanistic Interpretability. Accepted as a virtual poster, submission #716. , Jun 2026
Mixture-of-Depths (MoD) routers improve efficiency of transformer models by learning which tokens to process and which to bypass at each layer. However, their learned routing patterns have not been characterized in alternating-attention architectures, which mix local attention with sparse global attention to increase efficiency. We evaluate routing by training token-level MoD routers on Gemma 2 2B and find that the mean routing rate is significantly higher in full-attention layers than at sliding-window layers. We find that a token’s part of speech is correlated with whether the router skips or processes it, and that the same category is often routed differently at different layers: for example, determiners are preferentially skipped at the shallowest target layer but preserved at deeper ones. Because these patterns hold within single layers, they are not subject to the confound (perfect coincidence of attention type and layer parity in Gemma 2) that limits our routing-rate result.
@inproceedings{pleiter2026moddoes, title = {What Do Mixture-of-Depth Routers Learn? Routing Patterns in Gemma 2}, author = {Pleiter, Stein and Erdogan, Ege and Lucic, Ana}, booktitle = {ICML 2026 Workshop on Mechanistic Interpretability}, year = {2026}, month = jun, }