Urgent.News

What's breaking now, across thousands of outlets.

AI

Dynamic Abliteration: Non-Destructive Refusal Suppression via Engram Steering

The report discusses a new method for non-destructive refusal suppression in open-weight large language models (LLMs) like Qwen. Traditional ablation techniques permanently alter base model weights and can negatively impact performance on non-refusal tasks. Dynamic Abliteration using Multi-Layer Steering with Engram introduces a runtime approach that intercepts intermediate residual streams across layers using PyTorch forward hooks.

This method allows for suppression of refusal behavior without modifying base model weights, preserving performance on other tasks. The report also explores the use of Engram to make the approach dynamic by adjusting scaling factors for each token, rather than relying on static vectors that can degrade generation quality and increase KL-divergence.

Brief written by urgent.news from Hacker News's own syndicated text. Machine-written — may contain errors; check the original before relying on it.

Read the original at blog.madhukaraphatak.in →

More in AI

More from Thursday 24 September →