Skip to main navigation Skip to search Skip to main content

Building Self-Awareness of LLMs Over Weight Quantization

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Large Language Models (LLMs) are increasingly deployed in resource-constrained settings using lossy compression techniques such as weight quantization. While effective for reducing memory and computational demands, these techniques often degrade performance. This paper investigates a novel question: Can an LLM recognize that it has been lossy-compressed? We propose a probing framework that combines few-shot prompting with introspective queries to assess whether a model can identify its own quantization state. Experiments across multiple LLMs, including Qwen, Mistral, LLaMA, and GPT-3.5, demonstrate that many models can accurately detect their compression status, especially when provided with reference examples. These findings introduce a new perspective on model self-awareness and have broader implications for research on transparency and interpretability in compressed LLMs.

Original languageEnglish
Title of host publicationProceedings - 25th IEEE International Conference on Data Mining Workshops, ICDMW 2025
Pages2421-2424
Number of pages4
ISBN (Electronic)9798331581329
DOIs
StatePublished - 2025
Event25th IEEE International Conference on Data Mining Workshops, ICDMW 2025 - Washington, United States
Duration: 12 Nov 202515 Nov 2025

Publication series

NameIEEE International Conference on Data Mining Workshops, ICDMW
ISSN (Print)2375-9232
ISSN (Electronic)2375-9259

Conference

Conference25th IEEE International Conference on Data Mining Workshops, ICDMW 2025
Country/TerritoryUnited States
CityWashington
Period12/11/2515/11/25

Keywords

  • compression-aware computing
  • Large language model
  • quantization

Fingerprint

Dive into the research topics of 'Building Self-Awareness of LLMs Over Weight Quantization'. Together they form a unique fingerprint.

Cite this