Urgent.News

What's breaking now, across thousands of outlets.

AI

Testing Provenance-Based Controls Against Indirect Prompt Injection in AI Agents

I built a provenance-based prompt injection firewall that caught attacks text detectors missed, but also blocked every legitimate action.

Testing Provenance-Based Controls Against Indirect Prompt Injection in AI Agents

Defenses against prompt injection are built on the assumption that malicious text can be identified through a classifier. However, a weekend project revealed this approach is flawed. The author developed a tool that never reads the attacker's words, catching attacks that more sophisticated detectors missed. Unfortunately, the tool also blocked legitimate user actions, highlighting a trade-off that the industry currently struggles to resolve.

The core issue lies in the fact that AI agents read and act upon untrusted content, such as documents, web pages, and tool outputs. Any malicious instructions can be hidden within this untrusted text, making it impossible to prevent attacks by simply detecting malicious wording. The author argues that the solution lies in tracking the provenance of data, or where the data originates, rather than attempting to decipher the content itself.

By distinguishing between user-provided and untrusted source text, the tool can enforce policies that prevent the execution of actions based on untrusted data. This approach requires no machine learning models or complex classifiers, relying instead on a simple algorithm to track the origin of data and enforce strict rules to prevent unauthorized actions.

Written by urgent.news from HackerNoon's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at hackernoon.com →

More in AI

Meshy AI Credit Review 2026: 35 Credits for a 27.4 MB Eggshell

Independent review. I ran both tools on my own accounts, free and paid, with no vendor-provided access or early access of any kind.

  • Meshy AI charged 35 credits for Voronoi eggshell creation
  • SupaVoxel charged only 3 credits for the same task
  • Meshy's file was over three times larger than SupaVoxel's

# support

WildLens AI: Turn Your Outdoor Sounds into AI-Powered Nature Discovery 🌿 Anshu Kashyap Anshu Kashyap Anshu Kashyap Follow Oct 11 WildLens AI: Turn Your Outdoor Sounds into AI-Powered Nature Discovery…

Why I am building Threshold around replaceable agent sessions

About four months ago, I started using coding agents, beginning with Codex. Small tasks went well. I could describe a change, inspect the result, and move on. Longer projects felt different. As a conversation grew, it accumulated more than code: why I had chosen a direction, which alternatives I had rejected, what I wanted to leave alone…

  • Threshold aims to maintain project continuity across different agent sessions.
  • Core concepts of Threshold include Project, Task, and Run.
  • A test tool using Threshold successfully completed tasks across multiple agent sessions.

More from Sunday 11 October →