I built a software directory that re-checks every listing every six hours. Here is what the record taught me.
Every directory of indie software I have used had the same flaw. A third of the links were dead, and nothing told you which third. Listings were added once, by someone excited about them, and never looked at again. So I built one that checks. It Still Works is a catalogue of independent software: web tools, games, open-source projects. Every published listing is fetched on a six-hour schedule and…
Every indie software directory I've encountered has shared a common flaw: a third of the links were dead, and no system alerted users to which links were problematic. Listings were added by enthusiastic individuals and forgotten afterward. To address this issue, I created a solution that continuously verifies the status of these links.
It Still Works is a directory of independent software, encompassing web tools, games, and open-source projects. Every listing undergoes a six-hour refresh cycle, and the results are meticulously documented. Currently, the catalog consists of 672 listings, with a total of 40,817 recorded checks. Each listing page displays pertinent information, including whether the application responded during the last check, when it last did so, response time, any changes in visible content, and the expiration date of its TLS certificate.
If an application fails three consecutive checks, it is moved to a public graveyard. The graveyard remains active because understanding what once worked is crucial. The checking process itself proved relatively straightforward. The core of the system involves executing an HTTP fetch with a 20-second time limit. A successful response is indicated by a 2xx status code received within five seconds without encountering any "parking page" text within the body.
If the response fails to meet these criteria, an error is flagged. In addition to assessing the HTTP status, the system also generates a hash of the visible content to differentiate between redesigns and rotating advertisements. It records the certificate's expiration date and notes any URL redirects. The verification process unfolds in three stages: initially, every listing is checked; subsequently, the system determines which data to record; finally, the information is written to the database.
Requests are allocated per host, ensuring that multiple games hosted on a single platform are fetched sequentially with intervals between each check, rather than simultaneously. The scheduling mechanism was initially reliant on an external cron job, but this approach failed due to the host environment not executing the scheduled task.
Consequently, the scheduler was moved into the application's process, enabling each job to lease a slot within a Postgres table. The job writes a heartbeat while executing and records the outcome. Subsequent runs are scheduled based on the completion of the preceding run, rather than adhering to a fixed schedule. If a job ceases to heartbeat, it is identified as having timed out.
The public evidence page displays the outcome of the last full sweep, along with the number of scheduled sweeps completed within a specific timeframe. This transparency allows users to independently verify the accuracy of the system. After several weeks of operation, the site encountered an issue when itch.io began responding with a 403 status code for all requests.
This resulted in 40 working games being incorrectly recorded as failing, with the problem recurring twice. Once more, the site would have marked all these games as dead. The root cause of the problem was twofold. Firstly, a rule was implemented in the "decide" phase to mitigate such occurrences. If a host hosts at least five listings and experiences an 80% failure rate attributed to outright refusal, the system refrains from recording any data for that host during that particular sweep.
Only genuine 404 errors or unresolvable domains are excluded from this rule, while 429 responses are not recorded for any host. Secondly, a remedial action was taken to retroactively apply the aforementioned rule to all recorded runs, excluding the entries that would have been omitted during the subsequent sweep. This process involved removing the erroneous "content changed" entries, resending the relevant emails, and recalculating the current state of each listing based on the remaining valid checks.
To ensure the accuracy of the data, the system first performs a dry run as an admin action. In the case of itch.io games specifically, the system now verifies the game's own build files, which are served from a host that can be accessed by the platform, rather than relying solely on the store page. This approach provides a more accurate representation of the game's status.
Honesty emerged as a crucial engineering principle throughout the development process. The site adheres to a concise set of rules that function as both policy and tests. It is essential never to display numbers without solid evidence, to clearly label any detected facts, and to refrain from asserting that software is "working" when it has never been loaded.
These guidelines were instrumental in correcting two significant bugs that had gone unnoticed. Firstly, the site's sitemap query compared a string with a Date object, resulting in consistently false comparisons in JavaScript. Consequently, the recorded changes for 601 out of 672 URLs remained unchanged since a screenshot taken in September.
Secondly, the system compared each redirect with the stored URL instead of the previous observation, leading to the repetitive creation of 3,370 "redirect" rows for just 60 listings within a short period. Both issues resulted in the dissemination of inaccurate information, and neither could have been detected through conventional copy reviews.
As someone who creates software, I encourage you to claim your listing and display the status badge within your README file. The badge updates with each check and links to the verification record, providing transparency and trustworthiness. For those who have reached this point in my report, I am eager to learn more about your expectations.
What specific information would make you trust a listing? Additionally, what aspects of the directory would you find missing if they were present?
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.