TY - GEN
T1 - DeepEBC
T2 - Supercomputing Asia and International Conference on High Performance Computing in Asia Pacific Region Workshops, SCA/HPCAsia 2026 Workshops
AU - Xu, Jiaqi
AU - Zhang, Zhaorui
AU - Wei, Gaolin
AU - Di, Sheng
AU - Liu, Benben
AU - Yu, Xiaodong
AU - Lu, Xiaoyi
N1 - Publisher Copyright:
© 2026 Copyright held by the owner/author(s).
PY - 2026/1/25
Y1 - 2026/1/25
N2 - With the rapid evolution of LLMs, the size of the model has increased dramatically to achieve high accuracy, which brings strong memory and storage overhead for both LLM training, fine-tuning, and inference. Efficiently compressing the model parameter with a high compression ratio while guaranteeing the model accuracy becomes a challenging issue and is still largely under exploration. In this study, we employ the error-bounded lossy compressor, cuSZp, to compress model parameters. The algorithm achieves a compression ratio exceeding threefold while maintaining model inference accuracy loss within. Furthermore, we have enhanced the algorithm to better accommodate the data characteristics of LLMs, since LLM's model parameters may occasionally contain outliers. We evaluate each data block prior to compression; if the proportion of outliers within a block falls below a specified threshold, we apply specialized processing to that block to further improve compression efficiency. Additionally, we have extended the original algorithm to directly compress float16 tensors, offering greater convenience for models utilizing float16 parameters compared to existing quantization methods. Finally, we select distinct hyperparameter error bounds for each tensor based on the magnitude of the gradients and the variance of the model parameters, thereby further enhancing inference accuracy. Extensive experiment results based on the most popular benchmarks show that our proposed approach can achieve a compression ratio of up to 3.67 × while ensuring that the loss of inference accuracy is kept below.
AB - With the rapid evolution of LLMs, the size of the model has increased dramatically to achieve high accuracy, which brings strong memory and storage overhead for both LLM training, fine-tuning, and inference. Efficiently compressing the model parameter with a high compression ratio while guaranteeing the model accuracy becomes a challenging issue and is still largely under exploration. In this study, we employ the error-bounded lossy compressor, cuSZp, to compress model parameters. The algorithm achieves a compression ratio exceeding threefold while maintaining model inference accuracy loss within. Furthermore, we have enhanced the algorithm to better accommodate the data characteristics of LLMs, since LLM's model parameters may occasionally contain outliers. We evaluate each data block prior to compression; if the proportion of outliers within a block falls below a specified threshold, we apply specialized processing to that block to further improve compression efficiency. Additionally, we have extended the original algorithm to directly compress float16 tensors, offering greater convenience for models utilizing float16 parameters compared to existing quantization methods. Finally, we select distinct hyperparameter error bounds for each tensor based on the magnitude of the gradients and the variance of the model parameters, thereby further enhancing inference accuracy. Extensive experiment results based on the most popular benchmarks show that our proposed approach can achieve a compression ratio of up to 3.67 × while ensuring that the loss of inference accuracy is kept below.
KW - Accuracy
KW - cuSZp
KW - Error-Bounded Lossy Compression
KW - LLMs
UR - https://www.scopus.com/pages/publications/105032151650
UR - https://www.scopus.com/pages/publications/105032151650#tab=citedBy
U2 - 10.1145/3784828.3784832
DO - 10.1145/3784828.3784832
M3 - Conference contribution
AN - SCOPUS:105032151650
T3 - Proceedings of Supercomputing Asia and International Conference on High Performance Computing in Asia Pacific Region Workshops, SCA/HPCAsia 2026 Workshops
SP - 274
EP - 283
BT - Proceedings of Supercomputing Asia and International Conference on High Performance Computing in Asia Pacific Region Workshops, SCA/HPCAsia 2026 Workshops
Y2 - 26 January 2026 through 29 January 2026
ER -