The Model's Tell: Measuring Context-Leakage Attack Signals with Behavior Gauges
LLMs increasingly rely on external contexts, such as pre-defined system prompts or retrieved documents, to improve generation quality. However, processing these contexts alongside user queries creates an attack surface: adversarial inputs can induce models to disclose them. Prior probing studies suggest that leakage-related signals emerge in hidden states, yet the need to extract these states…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.