I tested 11 AI models on Indian GST, UPI and lakh-crore. Three famous ones got Puducherry wrong.
This is a submission for the Kaggle Benchmarking Challenge What I Benchmarked I build software for Indian small businesses (POS, billing, payments), and this month I've been contributing Kestra workflow blueprints for Indian finance: GSTIN validation, UPI end-of-day close, GSTR-2B reconciliation. Every one of them exists because one small mistake costs real money: one mistyped character in a…
A Kaggle Benchmarking Challenge examined how well various AI models could handle everyday Indian business data, including GSTIN validation, UPI reconciliation, and formatting of large numbers. Results showed that three renowned models incorrectly identified Puducherry as a Union Territory rather than a State, and that smaller, open-weight models like gpt-oss-20b and Gemma 4 performed better on some tasks compared to larger, faster models.
Checksum errors were particularly challenging for faster models, with some missing bills altogether, leading to potential financial losses. Open-weight models like gpt-oss-20b and Gemma 4 31B offered the best value, achieving high scores for a lower cost.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.