Predicting Alignment Generalization with Value Representations
LLM developers post-train their models to exhibit prosocial values and behavioral traits, which are enumerated in an alignment target. However, while recent post-training developments have yielded models that score highly on alignment evaluations, training models on sets of narrow behaviors still influences their behavior across unseen contexts and environments in unexpected ways. In this paper,…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.