Ilyass Ouardi

Ilyass Ouardi

M.Sc. Computer Science Student at University of Milan. Researching mechanistic interpretability, representation geometry, and verifiable AI safety.

Mechanistic Interpretability Papers

← Back to Reading List

A curated syllabus of foundational and seminal papers in mechanistic interpretability, representation geometry, and causal alignment. This bibliography reflects the scientific backbone behind my research—tracing the evolution from toy superposition models to high-dimensional concept cones and verifiable model interventions.

Each section is organized around a central mechanistic inquiry, complete with verified citations, direct preprint links, and synthesized takeaways.


1. The Linear Representation Hypothesis & Representation Geometry

How do concepts inhabit activation spaces, and where does single-vector linearity break down?


2. Superposition, Polysemanticity & Sparse Autoencoders (SAEs)

How do models pack more features than dimensions, and how do we untangle them?


3. Causal Mediation, Activation Patching & Circuit Discovery

Moving from observational correlation to surgical, falsifiable intervention.


4. Representation Engineering & Concept Erasure

Controlling representations, removing unwanted capabilities, and measuring off-target harm.


5. Unsupervised Discovery & Eliciting Latent Knowledge (ELK)

Discovering internal beliefs without reliance on behavioral labels.