LTBD: Learnable Trust-Boundary Delimiters for Prompt Injection Defense
Large language models (LLMs) perform remarkably well on complex tasks, yet remain highly vulnerable to prompt injection attacks, where malicious instructions embedded in external data can override user intent. Existing defenses remain limited by model fine-tuning requirements, vulnerability to adaptive attacks, or reliance on brittle handcrafted prompts. We argue that a fundamental source of this…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.