Urgent.News

What's breaking now, across thousands of outlets.

Tech

Graph RAG: where it actually breaks

Neo4j with a working schema: two days. Cypher traversal for the relationships I needed: another day or two, once I knew what I was querying. The graph structure, once committed, stayed mostly stable. That's not where things kept breaking. The routing problem When someone asks "what companies are exposed to semiconductor export restrictions?" the system needs to figure out whether to do graph…

When someone asks about which companies are exposed to semiconductor export restrictions, the system must decide whether to perform graph traversal, vector search, or a combination of both. Getting the routing wrong does not result in an obvious error; instead, it returns a response that cites real sources and sounds confident yet misses crucial context present in the graph.

One notable case involved a query about a Korean chipmaker's supplier network that produced news results instead of the desired answer. The issue lay in the system not routing the query to the correct part of the graph, where supplier relationships were already stored. Once the problem was identified, fixing it required only two lines of code.

The second major area where things broke down concerned entity quality. This issue is more subtle and gradual. Initially, the entity resolution process creates a set of merged entities. However, as source data continues to change, companies may be acquired, leading to new parent company names appearing alongside old subsidiary names. This results in the graph accumulating both nodes representing the same underlying entity, causing discrepancies in the results returned during graph traversal.

Rather than a binary state, entity quality degrades continuously. Periodic audits of merged pairs help maintain accuracy by checking whether the alias-to-entity mapping remains relevant for recently active entities. The code to identify stale merges is provided, detecting entities that have gone stale beyond a certain threshold.

However, the final decision on how to handle flagged pairs remains a manual process. The author's work, er-api, is a multilingual entity resolution service supporting Korean, Japanese, Chinese, and English corporate data.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

Architecting a Low-Power GPS Geofencing Engine for Android without Draining the Battery

It was the middle of a Friday afternoon, and I was sitting in a quiet, solemn gathering. The room was hushed, filled with people focused on the speaker at the front.

  • Utilized GeofencingClient API to define geographic regions
  • Offloaded GPS processing to Android framework via BroadcastReceivers
  • Implemented ForegroundService with persistent notification for Doze mode

OCR Uploaded Scans and Store Extracted Text in 4 Stages (With Validation)

An e-commerce document service should accept a scan, persist the private original, enqueue OCR, store extracted text under the document ID, and redact a derived copy before anybody shares it.

  • Four-stage asynchronous pipeline processes uploaded scan
  • Accepts scan, persists original, performs OCR, stores text
  • Validation checks upload, rejects empty or unsupported files

Node.js Managed Metrics Dashboard: Filtering Agent Loop Noise Across Regions

For a startup metrics dashboard, define a few stable boundaries around the media agent loop before evaluating any managed alternative to Prometheus and Grafana.

  • Preserve agentloopdurationseconds consistency across regions and deployments.
  • Limit label cardinality by excluding high-cardinality dimensions like user IDs and error messages.

More from Wednesday 30 September →