Collapse-Aware Clipping for Low-Bit Quantization of Post-Softmax Activations in Vision Transformers

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Post-training quantization (PTQ) has emerged as an important model compression technique for the efficient deployment of Vision Transformers (ViTs). However, directly applying existing PTQ methods developed for convolutional neural networks (CNNs) to ViTs often leads to performance degradation due to quantization-sensitive operations within the ViT architecture. In particular, post-softmax activations exhibit a highly asymmetric distribution in which most values are concentrated near zero while only a few take relatively large values, for which logarithmic quantizers have been widely adopted. Nevertheless, severe performance collapse can still occur in some ultralow-bit logarithmic quantization settings. In this paper, we show that such collapse cannot be sufficiently explained solely by an increase in element-wise reconstruction error, but is closely related to distortion of the row-normalized structure of post-softmax activations. To mitigate this issue, we propose a collapse-aware clipping framework composed of Entropy Cover Interval Selection (ECIS) and Collapse-Aware Selective Calibration (CASC). ECIS generates clipping boundary candidates that reflect the information structure of post-softmax activations, while CASC determines the clipping parameter by selecting an appropriate calibration objective according to whether a collapse signature is observed. Experimental results on ImageNet with various ViT-based models show that the proposed method alleviates accuracy degradation in ultralow-bit logarithmic quantization settings where performance collapse occurs, while maintaining comparable performance in non-collapse cases. These results demonstrate that the proposed framework can mitigate performance collapse in post-softmax activation quantization while preserving compatibility with existing logarithmic quantizers without redesigning the quantizer itself or introducing additional inference-time operations.

키워드

Quantization (signal)ModelingCalibrationVision transformersAccuracyConferencesEntropyComputersComputer visionTrainingClippinglogarithmic quantizationpost-softmax activationpost-training quantizationvision transformer
제목
Collapse-Aware Clipping for Low-Bit Quantization of Post-Softmax Activations in Vision Transformers
저자
Kim, Seong MinSong, Byung Cheol
DOI
10.1109/ACCESS.2026.3706454
발행일
2026
유형
Article
저널명
IEEE Access
14
페이지
95989 ~ 96003