TY - GEN
T1 - Beyond Reactive Safety
T2 - 63rd Annual Meeting of the Association for Computational Linguistics, ACL 2025
AU - Sun, Chenkai
AU - Zhang, Denghui
AU - Zhai, Cheng Xiang
AU - Ji, Heng
N1 - Publisher Copyright:
© 2025 Association for Computational Linguistics.
PY - 2025
Y1 - 2025
N2 - Given the growing influence of language model-based agents on high-stakes societal decisions, from public policy to healthcare, ensuring their beneficial impact requires understanding the far-reaching implications of their suggestions. We propose a proof-of-concept framework that projects how model-generated advice could propagate through societal systems on a macroscopic scale over time, enabling more robust alignment. To assess the long-term safety awareness of language models, we also introduce a dataset of 100 indirect harm scenarios, testing models' ability to foresee adverse, non-obvious outcomes from seemingly harmless user prompts. Our approach achieves not only over 20% improvement on the new dataset but also an average win rate exceeding 70% against strong baselines on existing safety benchmarks (AdvBench, SafeRLHF, WildGuardMix), suggesting a promising direction for safer agents.
AB - Given the growing influence of language model-based agents on high-stakes societal decisions, from public policy to healthcare, ensuring their beneficial impact requires understanding the far-reaching implications of their suggestions. We propose a proof-of-concept framework that projects how model-generated advice could propagate through societal systems on a macroscopic scale over time, enabling more robust alignment. To assess the long-term safety awareness of language models, we also introduce a dataset of 100 indirect harm scenarios, testing models' ability to foresee adverse, non-obvious outcomes from seemingly harmless user prompts. Our approach achieves not only over 20% improvement on the new dataset but also an average win rate exceeding 70% against strong baselines on existing safety benchmarks (AdvBench, SafeRLHF, WildGuardMix), suggesting a promising direction for safer agents.
UR - https://www.scopus.com/pages/publications/105028652112
UR - https://www.scopus.com/pages/publications/105028652112#tab=citedBy
U2 - 10.18653/v1/2025.findings-acl.332
DO - 10.18653/v1/2025.findings-acl.332
M3 - Conference contribution
AN - SCOPUS:105028652112
T3 - Proceedings of the Annual Meeting of the Association for Computational Linguistics
SP - 6422
EP - 6434
BT - Findings of the Association for Computational Linguistics
A2 - Che, Wanxiang
A2 - Nabende, Joyce
A2 - Shutova, Ekaterina
A2 - Pilehvar, Mohammad Taher
Y2 - 27 July 2025 through 1 August 2025
ER -