From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop
The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proceedings papers to 41 over six editions, documenting a field-wide transition from post-hoc interpretability of static models to mechanistic understanding and proactive control of generative systems. We synthesize insights from all 144 proceedings papers,…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.