How Does My Model Fail? Automatic Identification and Interpretation of Physical Plausibility Failure Modes with Matryoshka Transcoders
Hierarchical sparse features for interpreting physical-plausibility failures in generative models.
Authors: Yiming Tang, Abhijeet Sinha, and Dianbo Liu
Venue: CVPR Workshop on Explainable AI for Computer Vision, 2026
Recognition: Spotlight presentation
Matryoshka Transcoders learn hierarchical sparse features that support coarse-to-fine analysis of physical-plausibility failures in generative models. The method automatically identifies recurring failure modes and provides interpretable representations for understanding how a model’s outputs become physically implausible.