Skip to main navigation Skip to search Skip to main content

Low-rank sparse autoencoders: Unifying efficiency and geometric regularization for large language model interpretability

  • Jiajia Mu
  • , Benying Tan*
  • , Jie Ren
  • , Mingyu Guo
  • , Chenchen Luo
  • , Ruibin Bai
  • *Corresponding author for this work

Research output: Journal PublicationArticlepeer-review

Abstract

Sparse autoencoders (SAEs) are pivotal for mechanistic interpretability in large language models (LLMs), decomposing polysemantic activations into monosemantic features. However, scaling SAEs to frontier models remains challenging because dense encoders incur prohibitive computational costs, while highly overcomplete dictionaries introduce geometric redundancy that hinders the extraction of independent semantic axes. To address these limitations, this study proposes low-rank sparse autoencoders (LowRank SAEs). The framework replaces the conventional dense encoder with two low-rank matrices, constraining feature representations to a lower-dimensional subspace. This design introduces implicit geometric regularization, improves feature orthogonality, and reduces computational scaling with model dimension m from O(m2) to O(m). In addition, a layer-adaptive rank allocation strategy aligns encoder capacity with the intrinsic dimensionality of LLM representations, while a LeakyReLU activation scheme preserves gradient flow. Experimental results demonstrate that LowRank SAEs establish a superior efficiency–fidelity Pareto frontier. Compared with TopK SAEs, the proposed method reduces FLOPs by 24.82–82.51% and parameters by 12.50–41.40%, while maintaining comparable semantic interpretability and limiting reconstruction degradation to 0.22–2.14%. Owing to its substantially lower theoretical computational footprint, the LowRank SAE architecture enables efficient feature learning, mitigates inference-latency bottlenecks, and supports large-scale, real-time AI safety analysis and deployment.

Original languageEnglish
Article number123768
JournalInformation Sciences
Volume755
DOIs
Publication statusPublished - 5 Nov 2026

Free Keywords

  • Geometric regularization
  • LLM model compression
  • Low-rank factorization
  • Mechanistic interpretability
  • Sparse autoencoders

ASJC Scopus subject areas

  • Control and Systems Engineering
  • Software
  • Theoretical Computer Science
  • Computer Science Applications
  • Information Systems and Management
  • Artificial Intelligence

Fingerprint

Dive into the research topics of 'Low-rank sparse autoencoders: Unifying efficiency and geometric regularization for large language model interpretability'. Together they form a unique fingerprint.

Cite this