How Does My Model Fail? Automatic Identification and Interpretation of Physical Plausibility Failure Modes with Matryoshka Transcoders

Hierarchical sparse features for interpreting physical-plausibility failures in generative models.

Authors: Yiming Tang, Abhijeet Sinha, and Dianbo Liu
Venue: CVPR Workshop on Explainable AI for Computer Vision, 2026
Recognition: Spotlight presentation

Matryoshka Transcoders learn hierarchical sparse features that support coarse-to-fine analysis of physical-plausibility failures in generative models. The method automatically identifies recurring failure modes and provides interpretable representations for understanding how a model’s outputs become physically implausible.

Read the paper on arXiv.