From Confusion to Clarity: Confusion-Aware Retrieval and Knowledge Injection for Text Classification
Large language models (LLMs) struggle to classify text into taxonomies with many semantically similar labels, as the distinctions are domain-specific and not captured by pre-training. To handle large label spaces, a common approach retrieves top-$K$ candidate labels by embedding similarity and prompt the LLM to choose among them. However, top-$K$ retrieval reduces the number of candidates but…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.