RESEARCH

Refusal Beyond a Single Direction: A Preliminary Comparison of Diff-in-Means and INLP

ArXiv cs.AI · Mon, 15 Jun 2026 04:00:00 GMT

arXiv:2606.13720v1 Announce Type: new Abstract: Arditi et al. (2024) has shown that refusal in safety fine-tuned chat models is mediated by a single linear direction in the residual stream, recoverable by a difference-in-means (DiM) of harmful and harmless activations. We compare

Read original source Discuss with A.S.I.S