Are Latent Reasoning Models Easily Interpretable?
Models normally do all their reasoning in a continuous hidden state instead of spitting out readable text which makes them hard to monitor. The authors tested the Coconut and CODI models and it turns out these models barely even use their hidden reasoning steps for logical tasks like PrOntoQA and ProsQA. You can force the models to stop thinking early and they almost always spit out the same…
We haven't written up this one. Lobsters has the full story — the link below goes straight to it.