Open US Tariff Datasets You Can Actually Build With (2026): 5,689 HTS Codes, 60 Country Tiers, and a Machine-Readable Changelog
If you've ever tried to compute US import duties programmatically, you know the problem isn't the math — it's the data. The rates live in the USITC Harmonized Tariff Schedule (a website that fights scrapers), the Section 301 lists are PDFs and Federal Register notices, and the fee tables change every fiscal year. Every "tariff calculator" tutorial online ends with MFN_RATE = 0.165 # update…
If you've ever attempted to calculate US import duties programmatically, you quickly realize the issue isn't the math—rather, it's the data. The rates are stored in the USITC Harmonized Tariff Schedule (a website that actively blocks scrapers), Section 301 lists are in the form of PDFs and Federal Register notices, and the fee tables change with every fiscal year.
Most online tutorials for tariff calculators end with MFN_RATE = 0.165 # update manually. We created a free tariff calculator but grew tired of manually updating the MFN rate, so we made all the data behind the engine available as versioned, MIT-licensed open data on GitHub—updated monthly with every rate line including its legal basis and source link.
In this article, we'll explore what's inside the three repositories and demonstrate five real-world applications using pandas.
The repos:
1. invoicetariff/invoicetariff: Contains the complete set of rules, including MFN, Section 301, Section 232, Section 338, IEEPA, and fee information. Additional files include changelog.json, destinations.json, and refund-channels.json.
2. invoicetariff/us-tariff-rates-2026: A standalone single-file dataset containing the same ruleset-seed.json.
3. invoicetariff/hs-codes-database: Holds a dataset of all 5,689 US HTS 8-digit codes with their official MFN general rate.
The data is structured so that each Section 301/232/338/IEEPA/fee line includes a legalBasis string and a source object containing the title and URL—typically a Federal Register notice. This allows users to trace the origin of each duty percentage. Moreover, all name fields have English and Chinese versions, accommodating cross-border sourcing activities in Chinese. Loading the data is straightforward, but there are two potential issues to be aware of:
1. When loading the ruleset-seed.json file, ensure to use encoding= 'utf-8-sig' to avoid an unwanted BOM (Byte Order Mark) at the beginning of the file. The dtype parameter must also specify hts8 as a string, otherwise, pandas will interpret the 8-digit codes as integers, resulting in a type error during joins with other datasets.
2. MFN rates are presented as text rather than numerical values. General rates are provided as official USITC text (e.g., "Free", "16.5%"). Parsing these values appropriately is crucial, as some rates may not be simple percentage-based, such as "2.1¢/kg + 12%".
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.