TY - GEN
T1 - Feature-Guided SAE Steering for Refusal-Rate Control using Contrasting Prompts
AU - Bhargav, Samaksh
AU - Zhu, Zining
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Large Language Model (LLM) deployment requires guiding the LLM to recognize and not answer unsafe prompts while complying with safe prompts. Previous methods for achieving this require adjusting model weights along with other expensive procedures. While recent advances in Sparse Autoencoders (SAEs) have enabled interpretable feature extraction from LLMs, existing approaches lack systematic feature selection methods and principled evaluation of safety-utility tradeoffs. We explored using different steering features and steering strengths using Sparse Auto Encoders (SAEs) to provide a solution. Using an accurate and innovative contrasting prompt method with the AI-Generated Prompts Dataset from teknium/OpenHermes-2p5-Mistral-7B and Air Bench eu-dataset to efficiently choose the best features in the model to steer, we tested this method on Llama-3 8B. We conclude that using this method, our approach achieves an 18.9% improvement in safety performance while simultaneously increasing utility by 11.1%, demonstrating that targeted SAE steering can overcome traditional safety-utility tradeoffs when optimal features are identified through principled selection methods.
AB - Large Language Model (LLM) deployment requires guiding the LLM to recognize and not answer unsafe prompts while complying with safe prompts. Previous methods for achieving this require adjusting model weights along with other expensive procedures. While recent advances in Sparse Autoencoders (SAEs) have enabled interpretable feature extraction from LLMs, existing approaches lack systematic feature selection methods and principled evaluation of safety-utility tradeoffs. We explored using different steering features and steering strengths using Sparse Auto Encoders (SAEs) to provide a solution. Using an accurate and innovative contrasting prompt method with the AI-Generated Prompts Dataset from teknium/OpenHermes-2p5-Mistral-7B and Air Bench eu-dataset to efficiently choose the best features in the model to steer, we tested this method on Llama-3 8B. We conclude that using this method, our approach achieves an 18.9% improvement in safety performance while simultaneously increasing utility by 11.1%, demonstrating that targeted SAE steering can overcome traditional safety-utility tradeoffs when optimal features are identified through principled selection methods.
KW - feature steering
KW - language model safety
KW - mechanistic interpretability
KW - refusal control
KW - sparse autoencoders
UR - https://www.scopus.com/pages/publications/105035383096
UR - https://www.scopus.com/pages/publications/105035383096#tab=citedBy
U2 - 10.1109/ICDMW69685.2025.00370
DO - 10.1109/ICDMW69685.2025.00370
M3 - Conference contribution
AN - SCOPUS:105035383096
T3 - IEEE International Conference on Data Mining Workshops, ICDMW
SP - 2902
EP - 2907
BT - Proceedings - 25th IEEE International Conference on Data Mining Workshops, ICDMW 2025
T2 - 25th IEEE International Conference on Data Mining Workshops, ICDMW 2025
Y2 - 12 November 2025 through 15 November 2025
ER -