Urgent.News

What's breaking now, across thousands of outlets.

AI

When Your AI Confidently Replies to Emails It Shouldn't Touch

A technical investigation into a RAG system that can't tell when it's out of its depth Setup InboxSync is a personal project I built: a multi-account email aggregation API that uses a RAG (Retrieval-Augmented Generation) pipeline to suggest replies. The system indexes emails via IMAP, categorizes them with GPT-4o-mini, and for actionable emails retrieves semantically similar training examples…

A technical investigation reveals that a RAG (Retrieval-Augmented Generation) system designed to suggest email replies is failing to discern when its knowledge is insufficient for the task at hand. The system, built as a personal project to help salespeople respond to leads more efficiently, uses a database of training examples to generate contextually relevant reply suggestions. However, a series of adversarial tests exposing the system's limitations has uncovered a critical flaw.

In each test, the system returned a confidence score of 0.85 for its generated responses, regardless of whether the source material was spam, an auto-reply, a refusal, a legal request, or a question requiring multi-topic understanding. For example, when given a spam query offering a 50% discount, the system replied politely with "Thank you for the exciting offer! I appreciate the heads-up about the Black Friday sale. I'll definitely take a look." Despite the obvious spam content, the system's confidence remained high.

The investigation also revealed that the system simply ignores the relevance score of the retrieved training examples, using a fixed confidence value of 0.85 in all cases. This approach disregards the actual similarity scores computed by the system, which ranged from 0.13 to 0.54 for the various tests. The lack of a relevance threshold means the system treats all retrieved neighbors equally, even when they are completely unrelated to the query.

Moreover, there is no mechanism in place to prevent the system from auto-replying to certain types of queries, such as auto-reply to auto-reply situations. The system's reliance on a hardcoded confidence value, rather than a dynamic relevance assessment, leads to potentially unsafe and unvetted replies being suggested, even for queries it should be unable to handle.

This failure highlights the importance of implementing proper relevance gating and confidence thresholds in RAG systems to prevent inappropriate or harmful responses.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Exploring Qwen 3.8 27B: A Powerful AI Model for Developers

Introduction to Qwen 3.8 27B Qwen 3.8 27B is a state-of-the-art language model that has been released on the Hugging Face platform.

  • Qwen 3.8 27B model with 27 billion parameters available on Hugging Face
  • Model excels in text generation, translation, and sentiment analysis tasks

A man was arrested after telling ChatGPT that he was planning to kill his ex-girlfriend, with the incident reported by the AI's developer, OpenAI. 경찰은 지난달 31일 미국 애리조나주 투산(Tucson)에 거주하는 26세 남성이 전 여자친구를 살해하겠다는 협박성 발언을 했다며 신고를 받고 출동했습니다. Police responded to a report on July 31 of a 26-year-old man making threatening statements that he would kill his ex-girlfriend in Tucson, Arizona. 현지 경찰에 따르면, 이 남성은 챗GPT에 여자친구와 헤어졌는데, 여자친구가 다른 남자를 만난다는 이유로 여자친구를 살해하고, 그 후 자살하겠다는 계획을 말했습니다. According to local police, the man told ChatGPT that he had broken up with his girlfriend and that she was seeing someone else, and that he planned to kill her and then commit suicide. 오픈AI는 이 남성의 발언을 인지하고, 이를 경찰에 신고했습니다. 오픈AI는 자사의 AI 챗봇이 유해한 발언을 할 경우, 이를 탐지해 신고할 수 있는 시스템을 갖추고 있습니다. OpenAI said it was aware of the man's statements and reported them to the police. The company has a system in place to detect and report harmful statements made by users of its ChatGPT chatbot. 경찰은 출동해 이 남성을 체포하고, 피해자의 안전을 확인했습니다. 오픈AI는 성명을 내고 "사용자가 유해한 발언을 할 경우, 이를 탐지하고 예방하기 위해 노력하고 있다"고 밝혔습니다. Police responded and arrested the man and confirmed the safety of the victim. OpenAI said in a statement that it works to detect and prevent users from making harmful statements.

More from Saturday 15 August →