{
  "id": 13419363,
  "title": "PDF to Excel in the Browser: Extracting Tables with JavaScript",
  "url": "https://urgent.news/2026/10/10/pdf-to-excel-in-the-browser-extracting-tables-with-javascript",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-10-10T13:26:07.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/shahzaib11/pdf-to-excel-in-the-browser-extracting-tables-with-javascript-3jf6"
  },
  "original_language": "en",
  "account": "Extracting tables from PDFs using JavaScript is a complex task, as PDFs don't have rows, columns, or cells like spreadsheets. Instead, they consist of glyphs positioned at specific coordinates on the page. To achieve this, a process called \"reconstructing structure from positioned text\" is required.\n\nThe first step is to obtain the positioned text from the PDF using pdf.js's page.getTextContent() method. This returns each text fragment along with its transform matrix, which includes the x/y position, font size, and fragment width. These are the raw materials for further processing.\n\nThe core algorithm, groupLines(), reconstructs lines of text by throwing away empty fragments, sorting them by descending y-coordinate and ascending x-coordinate, and then creating new lines when the y-coordinate difference exceeds a tolerance value. The tolerance scales with the font size, handling both large headings and tiny footnotes effectively. Within each line, fragments are sorted by x-coordinate and joined together with appropriate spacing, with wider gaps indicating column breaks.\n\nOnce the lines of text are obtained, they can be transformed into a spreadsheet format using SheetJS. The lines are converted into an array-of-arrays format and then used to create a worksheet. A corresponding workbook is created, the worksheet is appended, and the workbook is exported as an Excel file (xlsx format) for download.\n\nThe tool has a few limitations: it only works with text-based PDFs, not scanned images; it explicitly warns if no readable text is found in the PDF; and it handles reading order differently from real tables. Merged cells, multi-line cells, and rotated text can also confuse the coordinate heuristic. However, the output provides clean, readable text in the correct order, making it useful for bank statements, invoices, and reports. For more complex documents like annual reports with nested tables, it's essential to have realistic expectations.",
  "summary": "Series: Building PdfWord — a free, no-backend PDF tools site (Part 11) \"Convert my bank statement PDF to Excel\" — one of the most requested features I've gotten, and one of the most technically dishonest-sounding. Here's the dirty secret of PDF-to-Excel: PDFs don't have tables. They have glyphs painted at coordinates. There are no rows, no columns, no cells — just text floating at (x, y)…",
  "key_points": [
    "JavaScript extracts positioned text from PDFs using page.getTextContent() method",
    "GroupLines() algorithm reconstructs lines by sorting fragments by y-coordinate and x-coordinate",
    "SheetJS converts lines to spreadsheet format for Excel file export"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}