TY - GEN
T1 - Distribution Prompting
T2 - 30th Conference on Empirical Methods in Natural Language Processing, EMNLP 2025
AU - Wang, Haojin
AU - Zhu, Zining
AU - Shi, Freda
N1 - Publisher Copyright:
© 2025 Association for Computational Linguistics.
PY - 2025
Y1 - 2025
N2 - Autoregressive neural language models (LMs) generate a probability distribution over tokens at each time step given a prompt. In this work, we attempt to systematically understand the probability distributions that LMs can produce, showing that some distributions are significantly harder to elicit than others. Specifically, for any target next-token distribution over the vocabulary, we attempt to find a prompt that induces the LM to output a distribution as close as possible to the target, using either soft (Li and Liang, 2021) or hard (Wallace et al., 2019) gradient-based prompt tuning. We find that (1) in general, distributions with very low or very high entropy are easier to approximate than those with moderate entropy; (2) among distributions with the same entropy, those containing “outlier tokens” are easier to approximate; (3) target distributions generated by LMs-even LMs with different tokenizers-are easier to approximate than randomly chosen targets. These results offer insights into the expressiveness of LMs and the challenges of using them as probability distribution proposers.
AB - Autoregressive neural language models (LMs) generate a probability distribution over tokens at each time step given a prompt. In this work, we attempt to systematically understand the probability distributions that LMs can produce, showing that some distributions are significantly harder to elicit than others. Specifically, for any target next-token distribution over the vocabulary, we attempt to find a prompt that induces the LM to output a distribution as close as possible to the target, using either soft (Li and Liang, 2021) or hard (Wallace et al., 2019) gradient-based prompt tuning. We find that (1) in general, distributions with very low or very high entropy are easier to approximate than those with moderate entropy; (2) among distributions with the same entropy, those containing “outlier tokens” are easier to approximate; (3) target distributions generated by LMs-even LMs with different tokenizers-are easier to approximate than randomly chosen targets. These results offer insights into the expressiveness of LMs and the challenges of using them as probability distribution proposers.
UR - https://www.scopus.com/pages/publications/105040268773
UR - https://www.scopus.com/pages/publications/105040268773#tab=citedBy
U2 - 10.18653/v1/2025.emnlp-main.1057
DO - 10.18653/v1/2025.emnlp-main.1057
M3 - Conference contribution
AN - SCOPUS:105040268773
T3 - EMNLP 2025 - 2025 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference
SP - 20904
EP - 20917
BT - EMNLP 2025 - 2025 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference
A2 - Christodoulopoulos, Christos
A2 - Chakraborty, Tanmoy
A2 - Rose, Carolyn
A2 - Peng, Violet
Y2 - 4 November 2025 through 9 November 2025
ER -