{
  "id": 10509809,
  "title": "Launch HN: Vespper (YC F24) – SOTA Docx MCP",
  "url": "https://urgent.news/2026/09/28/launch-hn-vespper-yc-f24-sota-docx-mcp",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-09-28T17:34:36.000Z",
  "source": {
    "name": "Hacker News",
    "slug": "hacker-news",
    "url": "https://www.vespper.com/blog/launching-vespper-docx-mcp"
  },
  "original_language": "en",
  "account": "Word documents are ubiquitous in industries such as legal, finance, and healthcare, serving as deliverables for contracts, regulatory submissions, and audit reports. Tools like Microsoft Copilot and Claude have integrated AI functionality directly into Word, while vertical AI agents, particularly in legal tech, require seamless interaction with .docx files. However, these AI agents face significant challenges when it comes to effectively editing and managing Word documents. Software engineers in the legal tech and healthcare sectors have reported spending weeks or even months developing tailored solutions to reliably edit Word documents, often resulting in complex, in-house systems. This difficulty arises because there are various ways for agents to edit Word documents, primarily categorized into three main approaches. Despite these advancements, current solutions struggle when faced with complex scenarios. Before delving into the solutions and their limitations, it is essential to understand the structure of a .docx file. A .docx file is essentially a ZIP archive containing a hierarchy of XML files that adhere to the OOXML (Office Open XML) specification. Within the ZIP, crucial files include document.xml, which holds the primary text and formatting information; styles.xml, which defines reusable styles; numbering.xml, which specifies list and numbering behavior; and separate XML files that manage headers, footers, footnotes, relationships, media, and document metadata. These XML files are verbose, making editing them challenging. For instance, a short paragraph in document.xml can expand to thousands of tokens due to the inclusion of styles, metadata, formatting information, and XML boilerplate. This nature of DOCX editing sets it apart from editing code or HTML, as the text representation in Word documents is far more complex and verbose. Editing .docx files involves dealing with numerous XML nodes and formatting details, making it a far more intricate process than editing simple files like Markdown or HTML. This complexity becomes even more pronounced for vertical AI agents. Companies like Harvey have encountered this issue firsthand. In order to develop a document editing system, Harvey had to create an agent that simultaneously acted as a legal assistant and a Word state machine. This dual role placed immense demands on the agent's context window, making it difficult for the AI to focus on its primary task. To address these challenges, one approach is round tripping. By liberating agents from the intricacies of Word mechanics, they can concentrate on their primary tasks instead. This concept draws inspiration from the Infrastructure as Code (IaC) world, where developers used to manually navigate cloud consoles to manage infrastructure. The advent of tools like Terraform and Pulumi revolutionized this process by allowing developers to simply write code, and the tool would handle the necessary changes across various environments. Similar to the IaC approach, the goal here is to enable agents to edit a more intuitive representation of the Word document while leaving the actual transformation to a separate tool. This is where the reconciler comes into play. However, creating a lossless conversion between the HTML representation and the original .docx file presents a significant challenge. Traditional tools like pandoc and mammoth.js often lose significant information during the conversion process and do not adequately reconcile the original file. To overcome these limitations, the team decided to build their own converter from scratch, ensuring minimal, clean, and high-fidelity representation of the .docx file. Ultimately, the team opted to train a model to handle the reconciliation process. Instead of building an extensive reconciliation engine, the model is trained on countless examples of HTML-to-OOXML changes, learning to predict the appropriate OOXML representation given an HTML change. This model serves as the core of the solution, taking the agent's HTML output and translating it into a faithful .docx representation.",
  "summary": null,
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}