Urgent.News

What's breaking now, across thousands of outlets.

Tech

We let crawlers fetch the pages we never want indexed, because a Disallowed URL never reads its noindex

Munchable is a barcode scanner for digestive conditions, and its marketing site is 426 indexable URLs. The file that decides which of those a crawler may fetch is 29 lines of TypeScript, and the only interesting line in it is the one that is shorter than people expect: disallow : [ ' /api/ ' , ' /auth/ ' ], Two paths. Not the checkout return pages, not the post-signup chooser, not the signed-in…

Munchable is a barcode scanner for digestive conditions, with 426 indexable URLs on its marketing site. The file that determines which of these URLs a crawler may fetch is 29 lines of TypeScript code. The only significant line in the code is "disallow : [/api/, /auth/]", which disallows crawling of just two paths out of the total indexable URLs.

This goes against the conventional understanding, as both "disallow" and "noindex" are typically used to prevent a page from being indexed. However, the article explains that when both directives are applied, the crawler never fetches the page, thus never seeing the "noindex" directive. Consequently, the URL remains indexed as a bare result, even if it should not be.

The issue is particularly relevant for pages that users might consider private, such as checkout returns, authentication pages, or any page a signed-in user accesses after a redirect. The article provides an example of a homepage link in the pricing section, which could be discovered without being fetched due to the "noindex" header.

Munchable's robots.ts file only disallows the two restricted paths, allowing all other pages to be indexed and including a "noindex" directive within the response body. The article emphasizes the importance of understanding the difference between the "disallow" and "noindex" directives and how combining them can unintentionally leave certain pages indexed.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

I built 150 dev tools nobody asked for. Google sent me 6 clicks.

TL;DR (for the scrollers) I built dauntexlabs : about 150 small tools (some still in progress, and yes, the plan is 500 someday lol).

  • Developed 150 browser-based dev tools with no backend
  • Implemented strict privacy checks to prevent data leaks
  • Received only 6 clicks from Google searches in 3 months

Unicode Text Styling: A Developer's Guide to Fancy Text That Actually Works

Unicode Text Styling: A Developer's Guide to Fancy Text That Actually Works If you've ever pasted "fancy" text from a generator into an app and watched it break — boxes on Android, stripped by your…

  • Unicode styling replaces characters with code points, not applying formatting
  • Mathematical Alphanumeric Symbols block supports bold, italic, and other styles
  • Compatibility issues may arise on older devices and screen readers

Your Seed Phrase Is Just 128 Bits With a Checksum: BIP-39 Explained for Developers

If you have ever set up a self-custody crypto wallet, you have seen it: a screen with 12 or 24 plain English words and a stern warning to write them down.

  • Seed phrase is 12 or 24 English words representing raw entropy
  • Entropy generated as 128 bits for 12-word or 256 bits for 24-word phrase
  • Checksum ensures phrase validity and detects mistyped words

More from Saturday 10 October →