When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce Evaluation
Interactive simulations increasingly evaluate policies in markets populated by language-model agents. Their outputs can look economic---prices, profits, consumer surplus, and welfare---without instantiating the behavior named in the claim. We audit this risk in a multi-turn buyer--seller testbed for configurable hotel transactions. An initial implementation reported welfare gains from two…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.