Skip to main navigation Skip to search Skip to main content

Feature-Guided SAE Steering for Refusal-Rate Control using Contrasting Prompts

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Large Language Model (LLM) deployment requires guiding the LLM to recognize and not answer unsafe prompts while complying with safe prompts. Previous methods for achieving this require adjusting model weights along with other expensive procedures. While recent advances in Sparse Autoencoders (SAEs) have enabled interpretable feature extraction from LLMs, existing approaches lack systematic feature selection methods and principled evaluation of safety-utility tradeoffs. We explored using different steering features and steering strengths using Sparse Auto Encoders (SAEs) to provide a solution. Using an accurate and innovative contrasting prompt method with the AI-Generated Prompts Dataset from teknium/OpenHermes-2p5-Mistral-7B and Air Bench eu-dataset to efficiently choose the best features in the model to steer, we tested this method on Llama-3 8B. We conclude that using this method, our approach achieves an 18.9% improvement in safety performance while simultaneously increasing utility by 11.1%, demonstrating that targeted SAE steering can overcome traditional safety-utility tradeoffs when optimal features are identified through principled selection methods.

Original languageEnglish
Title of host publicationProceedings - 25th IEEE International Conference on Data Mining Workshops, ICDMW 2025
Pages2902-2907
Number of pages6
ISBN (Electronic)9798331581329
DOIs
StatePublished - 2025
Event25th IEEE International Conference on Data Mining Workshops, ICDMW 2025 - Washington, United States
Duration: 12 Nov 202515 Nov 2025

Publication series

NameIEEE International Conference on Data Mining Workshops, ICDMW
ISSN (Print)2375-9232
ISSN (Electronic)2375-9259

Conference

Conference25th IEEE International Conference on Data Mining Workshops, ICDMW 2025
Country/TerritoryUnited States
CityWashington
Period12/11/2515/11/25

Keywords

  • feature steering
  • language model safety
  • mechanistic interpretability
  • refusal control
  • sparse autoencoders

Fingerprint

Dive into the research topics of 'Feature-Guided SAE Steering for Refusal-Rate Control using Contrasting Prompts'. Together they form a unique fingerprint.

Cite this