{
  "id": 3320747,
  "title": "Model teaches AI to read more like humans",
  "url": "https://urgent.news/2026/08/25/model-teaches-ai-to-read-more-like-humans",
  "topic": "ai",
  "section": "AI",
  "published": "2026-08-25T17:57:44.000Z",
  "source": {
    "name": "Futurity",
    "slug": "futurity",
    "url": "https://www.futurity.org/model-teaches-ai-to-read-more-like-humans/"
  },
  "original_language": "en",
  "account": "A new AI model, SlideAgent, developed by researchers at Georgia Tech and J.P. Morgan, teaches workplace AI systems to read more like humans. Modern AI platforms can swiftly summarize reports, analyze documents, and answer questions; however, they often overlook important details when information spans multiple slides, charts, and tables. This shortcoming can have costly consequences, as small mistakes, such as misreading numbers or overlooking footnotes, can impact reporting, risk assessment, and strategic decisions.\n\nSlideAgent addresses this issue by breaking down complex visual documents like presentation slide decks, brochures, and reports into multiple levels, enabling the model to analyze both the big picture and fine details. This human-inspired approach results in more accurate and reliable interpretations compared to existing systems. Beyond enhancing workplace tools, SlideAgent demonstrates a broader shift in AI by showcasing that smarter design and more efficient reasoning can improve performance.\n\nThe researchers tested SlideAgent on various real-world documents, including financial presentations, technical slides, and visual question-answering datasets. The system outperformed leading commercial models and open-source tools, with accuracy gains of up to 10% in certain cases. This improvement was particularly notable on complex tasks, such as comparing information across slides and understanding how visuals relate to each other on a page.\n\nSlideAgent mimics human reading habits by analyzing documents at three levels: the full document, individual pages, and specific elements like charts, tables, and text blocks. The system employs a network of specialized agents to divide and coordinate analysis, combining outputs to build a structured understanding of the entire document. This enables SlideAgent to answer questions accurately and reason across multiple pages more effectively than conventional multimodal AI systems, which often process entire pages at once and can lead to errors like miscounting items in charts or overlooking critical details in dense visuals.",
  "summary": "A new framework helps AI better understand complex visual documents like presentation slide decks, brochures, and reports.",
  "key_points": [
    "SlideAgent model developed by Georgia Tech and J.P. Morgan",
    "Breaks down complex visual documents into multiple levels",
    "Outperforms existing systems with up to 10% accuracy gains"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}