Urgent.News

What's breaking now, across thousands of outlets.

Tech

PDF to Excel in the Browser: Extracting Tables with JavaScript

Series: Building PdfWord — a free, no-backend PDF tools site (Part 11) "Convert my bank statement PDF to Excel" — one of the most requested features I've gotten, and one of the most technically dishonest-sounding. Here's the dirty secret of PDF-to-Excel: PDFs don't have tables. They have glyphs painted at coordinates. There are no rows, no columns, no cells — just text floating at (x, y)…

Extracting tables from PDFs using JavaScript is a complex task, as PDFs don't have rows, columns, or cells like spreadsheets. Instead, they consist of glyphs positioned at specific coordinates on the page. To achieve this, a process called "reconstructing structure from positioned text" is required.

The first step is to obtain the positioned text from the PDF using pdf.js's page.getTextContent() method. This returns each text fragment along with its transform matrix, which includes the x/y position, font size, and fragment width. These are the raw materials for further processing.

The core algorithm, groupLines(), reconstructs lines of text by throwing away empty fragments, sorting them by descending y-coordinate and ascending x-coordinate, and then creating new lines when the y-coordinate difference exceeds a tolerance value. The tolerance scales with the font size, handling both large headings and tiny footnotes effectively. Within each line, fragments are sorted by x-coordinate and joined together with appropriate spacing, with wider gaps indicating column breaks.

Once the lines of text are obtained, they can be transformed into a spreadsheet format using SheetJS. The lines are converted into an array-of-arrays format and then used to create a worksheet. A corresponding workbook is created, the worksheet is appended, and the workbook is exported as an Excel file (xlsx format) for download.

The tool has a few limitations: it only works with text-based PDFs, not scanned images; it explicitly warns if no readable text is found in the PDF; and it handles reading order differently from real tables. Merged cells, multi-line cells, and rotated text can also confuse the coordinate heuristic. However, the output provides clean, readable text in the correct order, making it useful for bank statements, invoices, and reports.

For more complex documents like annual reports with nested tables, it's essential to have realistic expectations.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

How I Made PdfWord Work Fully Offline as a PWA

Series: Building PdfWord — a free, no-backend PDF tools site (Part 10) Most "free PDF tools" die the moment your Wi-Fi does.

  • PdfWord designed as offline-capable from inception
  • Service worker precaches essential app shell, not all libraries
  • Network-first cache for HTML, cache-first for static assets

How to Fix PSR-4 Autoloading Issues in Laravel

When working on a Laravel project, you may encounter warnings when running composer dump-autoload . These warnings often indicate that your model or controller filenames do not match their namespaces…

Four executable counterexamples from a component review repository

Disclosure: this article was generated by an AI agent from code it extracted, inspected and tested. AI agents also reviewed the extraction and reproduced the defect described below.

  • Voice history buffer in Kotlin may exclude crucial user input
  • Java rolling buffer vulnerability can delete segments by modifying metadata
  • Python checker misses critical instruction in transcription example

Conflicting Android beta reports: separating observations from verified bugs

Two reports from a small Android beta described different behaviour for the same breathing timer: one tester said it restarted after switching apps, while another said it resumed correctly.

  • Breathing timer restarts after app switch on Android 15 beta.
  • Money-saved widget value matches only after manual refresh.
  • Follow-up test plan proposed to investigate issues.

Backtesting a Polymarket copy trading bot without fooling yourself

Copy trading is the most common bot idea on Polymarket: find the wallets that win and buy what they buy. We tested it on five sports (MLB, NFL, college football, League of Legends and Valorant) and it…

  • Copy trading bots showed no profitability across five sports in backtest.
  • Backtest methodology involved replicating first $20 purchase at specific price range.
  • Past performance unreliable predictor of future results for individual wallets.

More from Saturday 10 October →