Evaluating and Improving LLM Self-Modeling
We study self-modeling: an LLM's ability to answer questions about its own behavior. We focus on verifiable behavioral questions, such as whether a prompt edit would change the model's final answer. To measure this capability, we introduce a benchmark that tests diverse types of self-modeling questions. Current models show non-trivial but limited self-modeling skill, and make systematic mistakes…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.