I gave my text-to-SQL agent a business glossary. One version helped a lot, one did nothing.
A text-to-SQL agent sees table and column names. It does not know that "revenue" in your company means billing_invoice.total_net , only for invoices with status = 'issued' . So it guesses, and the guess runs without errors and returns the wrong number. I added business terms to schemagate 1.2.0 (open source, Apache-2.0) and measured whether they help. Short answer: a complete glossary helps a…
A text-to-SQL agent interprets table and column names, but lacks knowledge of business terms like "revenue," which may correspond to "billing_invoice.total_net" for invoices with a status of "issued." When given no terms, the model often guesses incorrectly, yielding lower accuracy. However, adding a complete glossary of business terms, uniquely learned from other questions, significantly improves accuracy.
These terms enable the model to pull relevant tables, though merely learning terms from other users' questions proved ineffective. The results demonstrated that a complete glossary can enhance the agent's performance by 9.9 percentage points. While learned terms improved table discovery, they did not teach the model how to formulate SQL queries.
To optimize the agent, focus on documenting frequently debated business terms, such as "revenue" or "active customer," as they can yield substantial improvements. Additionally, it is crucial to verify that any added terms comply with the caller's access permissions.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.