ChatGPT transcripts are reportedly read by humans to improve responses, including those with personal information — 'Project Lilly' has seen OpenAI hire hundreds of contractors to manually review logs
404 Media reports that OpenAI has hired hundreds of contractors to evaluate ChatGPT responses manually.
OpenAI has been reportedly using human reviewers to read ChatGPT transcripts, including those containing personal information, in an effort to improve the AI's responses. This process, known as Project Lily, involves hundreds of contractors who evaluate logs manually, looking for quality and adherence to guidelines. The reviewers are tasked with judging the clarity of ChatGPT's answers, avoiding overly technical or patronizing language, and ensuring the AI doesn't make obvious factual mistakes.
However, the guidelines are often contradictory and change frequently, leading to frustration among reviewers. While the chats supposedly undergo anonymization, some personal data may still slip through, particularly in shorter conversations. Additionally, the reviewer's job is to assess whether ChatGPT's responses sound human-like, but they are not graded on factual accuracy.
This human review process is separate from safety checks, which determine if a user might pose a threat to themselves or others. OpenAI has stated that users can opt out of data collection, but the option is not retroactive, meaning past chats may already be part of the dataset. Other AI companies, such as Google Gemini, Anthropic, and Perplexity, also have similar policies regarding human review of chat logs.
Written by urgent.news from Tom's Hardware's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.