Steering vectors align LLMs with human values
Linear directions extracted from large‑language‑model activation distributions map onto human‑value axes with measurable fidelity.
- Steering vectors align with human values in LLMs
- Distribution-driven steering restores human-value topology
- Larger models capture more detailed activation manifolds

