Skip to main navigation Skip to search Skip to main content

DeepEBC: Compressing the Pre-Trained LLMs with Error-Bounded Lossy Compression

  • Jiaqi Xu
  • , Zhaorui Zhang
  • , Gaolin Wei
  • , Sheng Di
  • , Benben Liu
  • , Xiaodong Yu
  • , Xiaoyi Lu
  • Hong Kong Polytechnic University
  • Argonne National Laboratory
  • The University of Hong Kong
  • University of California Merced

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

With the rapid evolution of LLMs, the size of the model has increased dramatically to achieve high accuracy, which brings strong memory and storage overhead for both LLM training, fine-tuning, and inference. Efficiently compressing the model parameter with a high compression ratio while guaranteeing the model accuracy becomes a challenging issue and is still largely under exploration. In this study, we employ the error-bounded lossy compressor, cuSZp, to compress model parameters. The algorithm achieves a compression ratio exceeding threefold while maintaining model inference accuracy loss within. Furthermore, we have enhanced the algorithm to better accommodate the data characteristics of LLMs, since LLM's model parameters may occasionally contain outliers. We evaluate each data block prior to compression; if the proportion of outliers within a block falls below a specified threshold, we apply specialized processing to that block to further improve compression efficiency. Additionally, we have extended the original algorithm to directly compress float16 tensors, offering greater convenience for models utilizing float16 parameters compared to existing quantization methods. Finally, we select distinct hyperparameter error bounds for each tensor based on the magnitude of the gradients and the variance of the model parameters, thereby further enhancing inference accuracy. Extensive experiment results based on the most popular benchmarks show that our proposed approach can achieve a compression ratio of up to 3.67 × while ensuring that the loss of inference accuracy is kept below.

Original languageEnglish
Title of host publicationProceedings of Supercomputing Asia and International Conference on High Performance Computing in Asia Pacific Region Workshops, SCA/HPCAsia 2026 Workshops
Pages274-283
Number of pages10
ISBN (Electronic)9798400723285
DOIs
StatePublished - 25 Jan 2026
EventSupercomputing Asia and International Conference on High Performance Computing in Asia Pacific Region Workshops, SCA/HPCAsia 2026 Workshops - Osaka, Japan
Duration: 26 Jan 202629 Jan 2026

Publication series

NameProceedings of Supercomputing Asia and International Conference on High Performance Computing in Asia Pacific Region Workshops, SCA/HPCAsia 2026 Workshops

Conference

ConferenceSupercomputing Asia and International Conference on High Performance Computing in Asia Pacific Region Workshops, SCA/HPCAsia 2026 Workshops
Country/TerritoryJapan
CityOsaka
Period26/01/2629/01/26

Keywords

  • Accuracy
  • cuSZp
  • Error-Bounded Lossy Compression
  • LLMs

Fingerprint

Dive into the research topics of 'DeepEBC: Compressing the Pre-Trained LLMs with Error-Bounded Lossy Compression'. Together they form a unique fingerprint.

Cite this