Are We Sending Too Much Data to LLMs? Agentic Production Support (APS)
While working on Agentic AI for production support, one question came to my mind: Do we really know what data we are sending to the LLM? Let's take a simple production incident. Host: ip-10–0–21–145 Memory: 1024 MB Contact: user@example.com Authorization: Bearer abc.def.ghi ERROR: Service failed due to disk space issue For RCA, the LLM mainly needs to understand: "Service failed because of a disk…
As the author delves into the topic of Agentic AI for production support, a crucial question emerges: Are we transmitting excessive data to Language Models (LLMs)? To illustrate, consider a straightforward production incident where a host with IP address "ip-10-0-21-145" encounters a memory limitation of 1024 MB and a contact point of user@example.com.
The authorization mechanism is outlined as "Authorization: Bearer abc.def.ghi." Upon reaching the root cause analysis (RCA), the LLM primarily needs to comprehend that "the service failed due to a disk space issue." It is apparent that the LLM does not necessitate additional details such as the actual hostname, email, AWS resource specifics, file paths, capacity values, or the authorization token.
This observation led the author to perceive the LLM as a Data Egress Boundary. Instead of the conventional flow: Production Data → Retrieval Augmented Generation (RAG) → LLM, the proposed approach is: Production Data → Clean/Sanitization Layer → RAG → LLM. It is imperative to cleanse or redact sensitive information before it reaches the LLM or embedding models.
However, an additional significant point must be acknowledged. RAG alone does not constitute a security layer. We often assume our data is safe due to the utilization of RAG, yet before storing a document in a vector database, an embedding model is employed. Consequently, the actual flow may be: Raw Data → Embedding Model → Vector DB → Retrieval → LLM.
Thus, it becomes evident that sanitization should be implemented prior to embedding, not solely prior to the final LLM call. The same principle ought to be applied to retrieval queries and AI observability logs. For high-risk data, such as API keys, passwords, JWTs, or bearer tokens, a fundamental rule should be established: If sensitive data persists after sanitization, the model should not be invoked.
The author posits that AI governance should transcend mere policy statements and be integrated into the architectural framework. The suggested architecture is: Raw Context → Sanitize → Validate → RAG / LLM. The ultimate goal is not to eliminate useful context, but rather to furnish the AI with sufficient context to address the problem, while preventing the disclosure of unnecessary information.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.