Evaluating LLMs on Conversational Text-to-SQL under Chain Ambiguity and Intent Drift
Recent advances in large language models (LLMs) have established conversational text-to-SQL as a practical interface between users and databases, often involving multiple turns of clarification and revision. However, existing benchmarks primarily evaluate execution accuracy, leaving the unfolding and shifting of user intent across turns largely uncovered. To address this, we introduce TIDE-Bench,…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.