[2603.22061] On the Failure of Topic-Matched Contrast Baselines in

[2603.22061] On the Failure of Topic-Matched Contrast Baselines in Multi-Directional Refusal Abliteration

arXiv - Machine Learning March 24, 2026 4 min read

About this article

Abstract page for arXiv paper 2603.22061: On the Failure of Topic-Matched Contrast Baselines in Multi-Directional Refusal Abliteration

Computer Science > Machine Learning arXiv:2603.22061 (cs) [Submitted on 23 Mar 2026] Title:On the Failure of Topic-Matched Contrast Baselines in Multi-Directional Refusal Abliteration Authors:Valentin Petrov View a PDF of the paper titled On the Failure of Topic-Matched Contrast Baselines in Multi-Directional Refusal Abliteration, by Valentin Petrov View PDF HTML (experimental) Abstract:Inasmuch as the removal of refusal behavior from instruction-tuned language models by directional abliteration requires the extraction of refusal-mediating directions from the residual stream activation space, and inasmuch as the construction of the contrast baseline against which harmful prompt activations are compared has been treated in the existing literature as an implementation detail rather than a methodological concern, the present work investigates whether a topically matched contrast baseline yields superior refusal directions. The investigation is carried out on the Qwen~3.5 2B model using per-category matched prompt pairs, per-class Self-Organizing Map extraction, and Singular Value Decomposition orthogonalization. It was found that topic-matched contrast produces no functional refusal directions at any tested weight level on any tested layer, while unmatched contrast on the same model, same extraction code, and same evaluation protocol achieves complete refusal elimination on six layers. The geometric analysis of the failure establishes that topic-matched subtraction cancels th...

Originally published on March 24, 2026. Curated by AI News.

Llms

What if Claude purposefully made its own code leakable so that it would get leaked

What if Claude leaked itself by socially and architecturally engineering itself to be leaked by a dumb human submitted by /u/smurfcsgoawp...

Reddit - Artificial Intelligence · 1 min · about 1 hour ago

Llms

Observer-Embedded Reality

Observer-Embedded Reality Consciousness, Complexity, Meaning, and the Limits of Human Knowledge A Conceptual Philosophy-of-Science Paper ...

Reddit - Artificial Intelligence · 1 min · about 1 hour ago

Llms

I think we’re about to have a new kind of “SEO”… and nobody is talking about it.

More people are asking ChatGPT things like: “what’s the best CRM?” “is this tool worth it?” “alternatives to X” And they just… trust the ...

Reddit - Artificial Intelligence · 1 min · about 6 hours ago

Llms

Why would Claude give me the same response over and over and give others different replies?

I asked Claude to "generate me a random word" so I could do some word play. Then I asked it again in a new prompt window on desktop after...

Reddit - Artificial Intelligence · 1 min · about 6 hours ago

[2603.22061] On the Failure of Topic-Matched Contrast Baselines in Multi-Directional Refusal Abliteration

About this article

Related Articles

What if Claude purposefully made its own code leakable so that it would get leaked

Observer-Embedded Reality

I think we’re about to have a new kind of “SEO”… and nobody is talking about it.

Why would Claude give me the same response over and over and give others different replies?

No comments

Stay updated with AI News