{
  "id": 6685395,
  "title": "I built an AI that keeps receipts. The mistakes became the useful part",
  "url": "https://urgent.news/2026/09/11/i-built-an-ai-that-keeps-receipts-the-mistakes-became-the-useful-part",
  "topic": "ai",
  "section": "AI",
  "published": "2026-09-11T02:00:32.000Z",
  "source": {
    "name": "e27",
    "slug": "e27",
    "url": "https://e27.co/i-built-an-ai-that-keeps-receipts-the-mistakes-became-the-useful-part-20260909/"
  },
  "original_language": "en",
  "account": "Building an AI system that records its predictions and the evidence behind them proved to be a critical lesson in the development of OnTheRice, a Singapore-based AI research publication. The key insight was that accurately predicting an outcome is easy, but preserving a transparent record of that prediction is far more challenging. Once the outcome is known, it's simple to manipulate the story to make it appear more impressive or accurate than it actually was.\n\nTo address this, the system was designed to lock in each prediction before the result was known, capturing details such as the entry price, source set, publication time, and evaluation time. By the date of the report, the system had 52 accurate predictions and 34 inaccurate ones out of 86 resolved calls, resulting in a 60.47 percent success rate. However, these figures are not definitive proof of the system's superiority, as the early misses were kept visible and verifiable. This transparency is crucial because it prevents hindsight from being mistaken for genuine learning.\n\nAnother challenge arose when the system began dealing with large volumes of reports from numerous sources covering the same event. While ten websites reporting on the same story might seem like ten independent confirmations, they are actually just one source with extensive distribution. AI systems can mistakenly equate the sheer volume of sources with genuine independent verification. Counting URLs does not equate to counting independent confirmations, and it can lead to an overestimation of a thin claim's strength due to its widespread dissemination.\n\nThe solution was to treat source origin as part of the evidence, grouping repetitive accounts together while giving more weight to genuinely independent reporting. This approach emphasizes tracing claims backward, rather than merely counting the number of pages that repeat them. In situations where conflicting information exists, such as one report at 8am stating that oil prices are falling and another at 5pm saying they are rising, an AI system that ignores the timing can inaccurately conclude that one statement is definitively wrong. This highlights the importance of recording the timestamps of both the source publication and the system's findings, as well as the system's assessment time.\n\nBeyond financial markets, this approach is applicable to any field where accurate records are essential. For instance, a business cannot honestly attribute a competitor's price change to falling sales without acknowledging that the competitor's price change occurred first. Proper documentation should include not just source links, but also the time the source was published, the time the system retrieved the source, the system's output generation time, and the time the output was assessed. This level of detail is crucial for establishing accountability, transparency, validity, and reliability in AI systems.\n\nWhile this methodology may seem less exciting than showcasing a more advanced model, it is invaluable during disputes, audits, or when a system fails. Instead of merely presenting a polished explanation after a mistake has been made, preserving the original output, evidence, timestamps, and scoring rules provides a clear and verifiable record of the system's thought process. It is far more useful to show the work and let the record speak, especially when it comes to demonstrating the system's limitations and potential areas for improvement.\n\nOne of the most significant lessons learned from this experience is the importance of allowing the system to abstain when the confidence level falls below a clear threshold. Rather than forcing an incorrect decision, the system should recognize when it lacks sufficient evidence to make a reliable prediction. By recording the reasons for this abstention, the system provides valuable insights for future improvements and helps prevent the spread of misinformation. For an AI product to truly earn trust, it must be willing to admit when it doesn't know the answer, rather than presenting an artificial level of certainty. This transparency is crucial for maintaining the system's credibility and ensuring that users are not misled by flawed or incomplete information.",
  "summary": "AI is remarkably good at producing answers. It is even better at sounding certain. Ask a difficult question and, within seconds, a system can gather information, connect ideas and return a polished explanation. Yet a harder question arrives later: what happens when reality proves the answer wrong? I met that problem while building OnTheRice, a […] The post I built an AI that keeps receipts. The…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 2,
    "also_reported_by": [
      {
        "outlet": "e27",
        "title": "I built a 21-role AI workforce. The hardest part was management",
        "url": "https://urgent.news/2026/09/10/i-built-a-21-role-ai-workforce-the-hardest-part-was-management",
        "published": "2026-09-10T02:00:12.000Z"
      }
    ]
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}