Urgent.News

What's breaking now, across thousands of outlets.

AI

Steering vectors align LLMs with human values

Linear directions extracted from large‑language‑model activation distributions map onto human‑value axes with measurable fidelity. The study demonstrates that these steering vectors preserve the full geometry of a theory‑driven value space, not merely isolated behavioral tweaks. Before this work, safety constraints were typically imposed by heavyweight fine‑tuning pipelines such as RLHF or DPO,…

A recent study reveals that linear directions derived from large language model activations align closely with human values. These steering vectors retain the full structure of the value space, not just superficial behavioral adjustments. Prior safety measures, such as RLHF, DPO, COLD-Steer, and BiPO, relied on heavyweight fine-tuning or behavior-centric methods, which focused on minimizing harms rather than exploring the underlying value topology.

By contrast, distribution-driven steering restores the expected human-value topology, showing a Spearman correlation of up to 0.51. This metric compares activation projections onto linear directions against a theoretical circumplex, with a peak correlation exceeding 0.5, indicating meaningful structure. Among evaluated metrics, distribution-driven methods alone demonstrate robust geometric alignment, whereas behavior-centric approaches show no significant correlation, despite similar benchmark performance.

The findings suggest that geometric fidelity improves with model scale but diminishes after instruction-tuning. Larger models provide more detailed activation manifolds that better capture the circumplex, yet fine-tuning reshapes these manifolds, weakening the connection between steering directions and human values. The study highlights the need for instruction-tuning regimes that preserve or enhance latent alignment.

While geometry correlates with human-consistent transfer across values, it does not guarantee the absence of edge-case harms, necessitating downstream validation. If validated, safety pipelines could incorporate lightweight inference-time modules that extract and apply distribution-driven steering vectors instead of full model fine-tuning.

Applying an additional geometry metric (Spearman correlation against the Schwartz framework) to existing safety benchmarks will determine if a deployment genuinely respects the intended value structure.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Seoul must weigh in on AI risks

Warning voices from within the U.S. artificial intelligence (AI) industry, from researchers to leading AI executives, have raised the alarm about safety challenges in the AI era.

More from Tuesday 15 September →