Principled Under Pressure: Post-Training Decides Whether LLMs Act on Their Own Moral Judgment
Language models increasingly act as agents. An agent that says an action is wrong and then takes it anyway is a different failure from one that does not know better, and evaluations of stated values cannot see it. We build a pre-registered panel of 248 scenarios across five kinds of pressure. Each scenario is posed twice to the same model, once as the agent choosing what to do and once in the…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.