TY - GEN
T1 - Building Self-Awareness of LLMs Over Weight Quantization
AU - Poon, Alexander
AU - Xu, Zhaozhuo
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Large Language Models (LLMs) are increasingly deployed in resource-constrained settings using lossy compression techniques such as weight quantization. While effective for reducing memory and computational demands, these techniques often degrade performance. This paper investigates a novel question: Can an LLM recognize that it has been lossy-compressed? We propose a probing framework that combines few-shot prompting with introspective queries to assess whether a model can identify its own quantization state. Experiments across multiple LLMs, including Qwen, Mistral, LLaMA, and GPT-3.5, demonstrate that many models can accurately detect their compression status, especially when provided with reference examples. These findings introduce a new perspective on model self-awareness and have broader implications for research on transparency and interpretability in compressed LLMs.
AB - Large Language Models (LLMs) are increasingly deployed in resource-constrained settings using lossy compression techniques such as weight quantization. While effective for reducing memory and computational demands, these techniques often degrade performance. This paper investigates a novel question: Can an LLM recognize that it has been lossy-compressed? We propose a probing framework that combines few-shot prompting with introspective queries to assess whether a model can identify its own quantization state. Experiments across multiple LLMs, including Qwen, Mistral, LLaMA, and GPT-3.5, demonstrate that many models can accurately detect their compression status, especially when provided with reference examples. These findings introduce a new perspective on model self-awareness and have broader implications for research on transparency and interpretability in compressed LLMs.
KW - compression-aware computing
KW - Large language model
KW - quantization
UR - https://www.scopus.com/pages/publications/105035378420
UR - https://www.scopus.com/pages/publications/105035378420#tab=citedBy
U2 - 10.1109/ICDMW69685.2025.00297
DO - 10.1109/ICDMW69685.2025.00297
M3 - Conference contribution
AN - SCOPUS:105035378420
T3 - IEEE International Conference on Data Mining Workshops, ICDMW
SP - 2421
EP - 2424
BT - Proceedings - 25th IEEE International Conference on Data Mining Workshops, ICDMW 2025
T2 - 25th IEEE International Conference on Data Mining Workshops, ICDMW 2025
Y2 - 12 November 2025 through 15 November 2025
ER -