-
摘要: Transformer模型具有庞大的参数规模和计算开销, 难以直接部署于资源受限的边缘设备, 限制了其在开放环境中的实际应用.低比特量化能够显著降低模型存储需求并提升推理效率, 是实现Transformer模型轻量化部署的重要技术路径.然而, 现有量化方法通常受固定量化区间约束, 在处理非线性激活产生的长尾分布时, 难以有效平衡裁剪误差与舍入误差, 导致模型性能明显下降. 为此, 提出一种面向低比特Transformer模型的无约束训练后量化方法(CFQuant), 通过建模神经网络激活值的概率分布实现自适应量化. 首先, CFQuant在校准过程中自适应估计激活分布密度, 并通过迭代搜索最小化量化误差. 其次, 设计高效尺度偏移算法动态调整模型分布, 以减小校准阶段与推理阶段之间的分布偏移, 进一步提升量化稳定性. 此外, 构建基于查找表的矩阵乘法加速机制, 将低比特推理中的浮点乘法转换为预计算查找操作, 从而降低推理计算开销.最后, 通过在多类Transformer模型以及视觉、语言和多模态任务上的大量实验, 充分验证了所提方法的灵活性和有效性.
-
关键词:
- 深度学习 /
- 训练后量化 /
- Transformer /
- 量化误差 /
- 自注意力
Abstract: The massive parameter scales and high computational overhead of Transformer models make them difficult to deploy directly on resource-constrained edge devices, limiting their practical applications in open environments. Low-bit quantization can significantly reduce model storage requirements and improve inference efficiency, making it an important technical approach for lightweight deployment of Transformer models. However, existing quantization methods are usually constrained by fixed quantization intervals. When processing long-tailed distributions produced by nonlinear activations, they have difficulty balancing clipping and rounding errors effectively, leading to considerable performance degradation. To address this problem, we propose a constraint-free post-training quantization method for low-bit Transformer models (CFQuant), which performs adaptive quantization by modeling the probability distributions of neural-network activations. First, CFQuant adaptively estimates activation-distribution density during calibration and minimizes quantization error through iterative search. Second, an efficient scale-shift algorithm dynamically adjusts model distributions to reduce the distribution shift between calibration and inference stages, thereby further improving quantization stability. Moreover, an acceleration mechanism based on matrix multiplication with a lookup table converts floating-point multiplications in low-bit inference into precomputed lookup operations, thereby reducing inference overhead. Finally, extensive experiments on multiple types of Transformer models and on vision, language, and multimodal tasks demonstrate the flexibility and effectiveness of the proposed method.-
Key words:
- deep learning /
- post-training quantization /
- Transformer /
- quantization error /
- self-attention
-
表 1 不同方法在ImageNet-1k数据集上分类任务的性能对比(%)
Table 1 Performance comparison of different methods on classification tasks on ImageNet-1k dataset (%)
方法 比特数(W/A) ViT-T ViT-S ViT-B DeiT-T DeiT-S DeiT-B Swin-T Swin-S Swin-B 全精度 32/32 75.47 81.39 84.54 72.21 79.85 81.80 81.37 83.23 85.27 PTQ4ViT[19] 4/4 16.53 42.57 30.69 36.96 34.08 64.39 73.06 76.09 74.02 APQ-ViT[33] 4/4 17.56 47.95 41.41 47.94 43.55 67.48 — 77.15 76.48 RepQ-ViT[20] 4/4 26.72 65.05 68.48 57.43 69.03 75.61 72.37 79.45 80.74 CFQuant 4/4 55.41 70.95 73.72 62.57 72.57 78.02 73.80 80.65 81.84 表 2 不同方法在MS-COCO数据集上目标检测和实例分割任务的性能对比(%)
Table 2 Performance comparison of different methods on object detection and instance segmentation tasks on the MS-COCO dataset (%)
方法 比特数(W/A) Mask R-CNN Cascade Mask R-CNN Swin-T Swin-S Swin-T Swin-S 边界框AP 掩码AP 边界框AP 掩码AP 边界框AP 掩码AP 边界框AP 掩码AP 全精度 32/32 46.0 41.6 48.5 43.3 50.4 43.7 51.9 45.0 PTQ4ViT[19] 4/4 6.9 7.0 26.7 26.6 14.7 13.5 0.5 0.5 APQ-ViT[33] 4/4 23.7 22.6 44.7 40.1 27.2 24.4 47.7 41.1 RepQ-ViT[20] 4/4 36.1 36.0 44.2 40.2 47.0 41.4 49.3 43.1 CFQuant 4/4 36.6 36.5 43.4 40.6 47.7 42.0 50.1 43.6 表 3 在多个零样本任务上LLaMA系列模型不同方法的性能对比(%)
Table 3 Performance comparison of different methods on multiple zero-shot tasks across LLaMA models (%)
模型 方法 比特数(W/A) PIQA ARC-e WinoGrande BoolQ ARC-c HellaSwag 平均值 LLaMA-7B 全精度 16/16 77.47 52.48 67.07 73.08 41.46 73.00 64.09 SmoothQuant[15] 4/4 49.80 30.40 48.00 49.10 25.80 27.40 38.41 OmniQuant[36] 4/4 66.15 45.20 53.43 63.51 31.14 56.44 52.65 AffineQuant[60] 4/4 69.37 42.55 55.33 63.73 31.91 57.65 53.42 CFQuant 4/4 67.91 43.90 55.36 64.45 32.19 57.14 53.49 LLaMA-13B 全精度 16/16 79.10 59.89 70.31 68.01 44.45 76.21 66.33 SmoothQuant[15] 4/4 61.04 39.18 51.06 61.80 30.80 52.29 49.36 OmniQuant[36] 4/4 69.69 47.39 55.80 62.84 33.10 58.96 54.63 AffineQuant[60] 4/4 66.32 43.90 54.70 64.10 29.61 56.88 52.58 CFQuant 4/4 71.20 45.61 56.39 64.17 33.35 59.51 55.04 表 4 在多个零样本任务上Qwen3系列模型不同方法的性能对比(%)
Table 4 Performance comparison of different methods on multiple zero-shot tasks across Qwen3 models (%)
模型 方法 比特数(W/A) PIQA ARC-e WinoGrande BoolQ ARC-c HellaSwag 平均值 Qwen3-8B 全精度 16/16 77.80 80.93 67.72 86.61 56.66 74.92 74.11 SmoothQuant[15] 4/8 73.42 71.01 62.16 79.11 44.32 51.68 63.62 PrefixQuant[61] 4/8 71.98 70.50 64.09 82.26 46.93 65.09 66.81 CFQuant 4/8 76.17 71.13 65.82 79.79 47.61 71.56 68.68 SmoothQuant[15] 4/4 50.53 25.51 52.17 39.58 25.73 26.72 36.71 PrefixQuant[61] 4/4 56.42 38.55 48.07 45.69 26.79 34.82 41.72 CFQuant 4/4 72.74 71.84 64.09 78.81 47.87 67.93 67.21 Qwen3-14B 全精度 16/16 79.82 82.87 72.93 89.33 60.24 78.82 77.34 SmoothQuant[15] 4/4 51.22 25.81 50.18 38.47 26.51 25.79 36.33 PrefixQuant[61] 4/4 57.89 43.81 51.30 56.73 29.10 39.83 46.44 CFQuant 4/4 65.07 58.71 56.91 78.84 40.10 49.39 58.17 表 5 在WikiText-2和C4数据集上语言生成任务的性能比较
Table 5 Performance comparison on language generation tasks on WikiText-2 and C4 datasets
模型 方法 比特数(W/A) WikiText-2 $ \downarrow $ C4 $ \downarrow $ LLaMA-7B 全精度 16/16 5.68 7.08 SmoothQuant[15] 4/4 25.25 32.32 OmniQuant[36] 4/4 11.26 14.51 CFQuant 4/4 10.36 13.91 LLaMA-13B 全精度 16/16 5.09 6.61 SmoothQuant[15] 4/4 40.05 47.18 OmniQuant[36] 4/4 10.87 13.78 CFQuant 4/4 10.14 13.28 Qwen3-8B 全精度 16/16 9.72 13.29 SmoothQuant[15] 4/4 $ 3.36\times10^{4} $ $ 2.29\times10^{4} $ PrefixQuant[61] 4/4 155.78 119.93 CFQuant 4/4 13.78 17.26 Qwen3-14B 全精度 16/16 8.64 12.01 SmoothQuant[15] 4/4 $ 2.16\times10^{5} $ $ 1.99\times10^{5} $ PrefixQuant[61] 4/4 187.28 129.05 CFQuant 4/4 32.30 35.76 表 6 在MMBench数据集上Qwen3-VL-4B模型的 不同方法性能对比(%)
Table 6 Performance comparison of different methods for the Qwen3-VL-4B model on the MMBench dataset (%)
表 7 在ImageNet-1k数据集上查找表额外内存开销的消融实验
Table 7 Ablation study of the additional lookup-table memory overhead on the ImageNet-1k dataset
模型 CFQuant 模型尺寸(MB) Top-1准确率(%) DeiT-S × 11.18 68.27 √ 11.19 (+0.1%) 72.57 (+4.30) Swin-S × 24.78 79.03 √ 24.80 (+0.2%) 80.65 (+1.62) 表 8 在ImageNet-1k数据集上不同激活函数 量化策略的消融实验
Table 8 Ablation study of quantization strategies for different activation functions on the ImageNet-1k dataset
模型 Softmax GeLU Top-1准确率(%) DeiT-S × × 68.27 √ × 68.82 (+0.55) × √ 71.33 (+3.06) √ √ 72.57 (+4.30) Swin-S × × 79.03 √ × 79.28 (+0.25) × √ 80.36 (+1.33) √ √ 80.65 (+1.62) 表 9 在ImageNet-1k数据集上高效尺度偏移算法的 消融实验
Table 9 Ablation studies of efficient scale-shift algorithm on the ImageNet-1k dataset
模型 ESA Top-1准确率(%) DeiT-S × 71.08 √ 72.57 (+1.49) Swin-S × 78.93 √ 80.65 (+1.72) 表 10 在ImageNet-1k数据集上不同量化方法校准时间的对比
Table 10 Comparison of calibration time for different quantization methods on the ImageNet-1k dataset
表 11 在ImageNet-1k数据集上的推理效率对比
Table 11 Comparison on inference efficiency on the ImageNet-1k dataset
方法 模型尺寸(MB) 位运算量(G) 延迟(s) 全精度 89.44 4 741 10.39 RepQ-ViT[20] 11.18 141 2.67 CFQuant 11.18 346 4.33 CFQuant $ _{+\text{MM-LUT}} $ 11.19 144 2.80 -
[1] Dosovitskiy A, Beyer L, Kolesnikov A, Weissenborn D, Zhai X H, Unterthiner T, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In: Proceedings of the 9th International Conference on Learning Representations. Virtual Event: OpenReview.net, 2021. [2] Dai Y, Chen X J, Wang X H, Pang M H, Gao L L, Shen H T. ReSParser: Fully convolutional multiple human parsing with representative sets. IEEE Transactions on Multimedia, 2024, 26: 1384−1394 doi: 10.1109/TMM.2023.3281070 [3] Zhang S S, Roller S, Goyal N, Artetxe M, Chen M Y, Chen S H, et al. OPT: Open pre-trained Tansformer language models. arXiv preprint arXiv: 2205.01068, 2022. [4] Li S S, Xu X, Meng W X, Song J K, Peng C, Shen H T. Mitigating hallucinations in large vision-language models via reasoning uncertainty-guided refinement. IEEE Transactions on Multimedia, 2025, 27: 7380−7391 doi: 10.1109/TMM.2025.3599076 [5] Touvron H, Martin L, Stone K, Albert P, Almahairi A, Babaei Y, et al. LLaMA 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv: 2307.09288, 2023. [6] Zitkovich B, Yu T H, Xu S C, Xu P, Xiao T, Xia F, et al. RT-2: Vision-language-action models transfer web knowledge to robotic control. In: Proceedings of the Conference on Robot Learning. Atlanta, USA: PMLR, 2023. 2165−2183 [7] 李浩然, 陈宇辉, 崔文博, 刘卫恒, 刘锴, 周明才, 等. 面向具身操作的视觉-语言-动作模型综述. 自动化学报, 2026, 52(1): 18−51Li Hao-Ran, Chen Yu-Hui, Cui Wen-Bo, Liu Wei-Heng, Liu Kai, Zhou Ming-Cai, et al. Survey of vision-language-action models for embodied manipulation. Acta Automatica Sinica, 2026, 52(1): 18−51 [8] Qu D L, Song H M, Chen Q Z, Wang D, Yao Y Q, Ye X Y, et al. SpatialVLA: Exploring spatial representations for visual-language-action model. In: Proceedings of the Robotics: Science and Systems 2025. Los Angeles, USA: University of Southern California, 2025. Article No. 011 [9] Liu Z, Lin Y T, Cao Y, Hu H, Wei Y X, Zhang Z, et al. Swin Tansformer: Hierarchical vision Tansformer using shifted windows. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Montreal, Canada: IEEE, 2021. 9992−10002 [10] 王文晟, 谭宁, 黄凯, 张雨浓, 郑伟诗, 孙富春. 基于大模型的具身智能系统综述. 自动化学报, 2025, 51(1): 1−19Wang Wen-Sheng, Tan Ning, Huang Kai, Zhang Yu-Nong, Zheng Wei-Shi, Sun Fu-Chun. Embodied intelligence systems based on large models: A survey. Acta Automatica Sinica, 2025, 51(1): 1−19 [11] Gu Y X, Dong L, Wei F R, Huang M L. MiniLLM: Knowledge distillation of large language models. In: Proceedings of the 12th International Conference on Learning Representations. Vienna, Austria: OpenReview.net, 2024. [12] Chen G B, Choi W, Yu X, Han T, Chandraker M. Learning efficient object detection models with knowledge distillation. In: Proceedings of the 31st International Conference on Neural Information Processing Systems. Long Beach, USA: Curran Associates, Inc., 2017. 742−751 [13] Ashkboos S, Croci M L, do Nascimento M G, Hoefler T, Hensman J. SliceGPT: Compress large language models by deleting rows and columns. In: Proceedings of the 12th International Conference on Learning Representations. Vienna, Austria: OpenReview.net, 2024. [14] Sun M J, Liu Z, Bair A, Kolter J Z. A simple and effective pruning approach for large language models. In: Proceedings of the 12th International Conference on Learning Representations. Vienna, Austria: OpenReview.net, 2024. [15] Xiao G X, Lin J, Seznec M, Wu H, Demouth J, Han S. SmoothQuant: Accurate and efficient post-training quantization for large language models. In: Proceedings of the 40th International Conference on Machine Learning. Honolulu, USA: PMLR, 2023. 38087−38099 [16] Liu Y J, Yang H R, Dong Z, Keutzer K, Du L, Zhang S H. NoisyQuant: Noisy bias-enhanced post-training activation quantization for vision Tansformers. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Vancouver, Canada: IEEE, 2023. 20321−20330 [17] Yin P, Zhu X S, Song J K, Gao L L, Shen H T. SI-BiViT: Binarizing vision Tansformers with spatial interaction. In: Proceedings of the 32nd ACM International Conference on Multimedia. Melbourne, Australia: ACM, 2024. 8169−8178 [18] Li Y H, Gong R H, Tan X, Yang Y, Hu P, Zhang Q, et al. BRECQ: Pushing the limit of post-training quantization by block reconstruction. In: Proceedings of the 9th International Conference on Learning Representations. Virtual Event: OpenReview.net, 2021. [19] Yuan Z H, Xue C H, Chen Y Q, Wu Q, Sun G Y. PTQ4ViT: Post-training quantization for vision Tansformers with twin uniform quantization. In: Proceedings of the 17th European Conference on Computer Vision. Tel Aviv, Israel: Springer, 2022. 191−207 [20] Li Z K, Xiao J R, Yang L W, Gu Q Y. RepQ-ViT: Scale reparameterization for post-training quantization of vision Tansformers. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Paris, France: IEEE, 2023. 17181−17190 [21] Jacob B, Kligys S, Chen B, Zhu M L, Tang M, Howard A, et al. Quantization and training of neural networks for efficient integer-arithmetic-only inference. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City, USA: IEEE, 2018. 2704−2713 [22] Cai J Y, Takemoto M, Nakajo H. A deep look into logarithmic quantization of model parameters in neural networks. In: Proceedings of the 10th International Conference on Advances in Information Technology. Bangkok, Thailand: ACM, 2018. Article No. 6 [23] Lloyd S. Least squares quantization in PCM. IEEE Transactions on Information Theory, 1982, 28(2): 129−137 [24] Max J. Quantizing for minimum distortion. IRE Transactions on Information Theory, 1960, 6(1): 7−12 [25] Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez A N, et al. Attention is all you need. In: Proceedings of the 31st International Conference on Neural Information Processing Systems. Long Beach, USA: Curran Associates, Inc., 2017. 6000−6010 [26] Devlin J, Chang M W, Lee K, Toutanova K. BERT: Pre-training of deep bidirectional Tansformers for language understanding. In: Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). Minneapolis, USA: Association for Computational Linguistics, 2019. 4171−4186 [27] Radford A, Wu J, Child R, Luan D, Amodei D, Sutskever I. Language models are unsupervised multitask learners [Online], available: https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf, July 3, 2026 [28] Touvron H, Cord M, Douze M, Massa F, Sablayrolles A, Jégou H. Training data-efficient image Tansformers & distillation through attention. In: Proceedings of the 38th International Conference on Machine Learning. Virtual Event: PMLR, 2021. 10347−10357 [29] Zhang Z Z, Zhang H, Zhao L, Chen T, Arik S Ö, Pfister T. Nested hierarchical Tansformer: Towards accurate, data-efficient and interpretable visual understanding. In: Proceedings of the 36th AAAI Conference on Artificial Intelligence. Virtual Event: AAAI Press, 2022. 3417−3425 [30] Strudel R, Garcia R, Laptev I, Schmid C. Segmenter: Transformer for semantic segmentation. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Montreal, Canada: IEEE, 2021. 7242−7252 [31] Brown T B, Mann B, Ryder N, Subbiah M, Kaplan J, Dhariwal P, et al. Language models are few-shot learners. In: Proceedings of the 34th International Conference on Neural Information Processing Systems. Vancouver, Canada: Curran Associates, Inc., 2020. Article No. 159 [32] Sun Y X, Liu R K, Bai H L, Bao H, Zhao K, Li Y N, et al. FlatQuant: Flatness matters for LLM quantization. In: Proceedings of the 42nd International Conference on Machine Learning. Vancouver, Canada: PMLR, 2025. 57587−57613 [33] Ding Y F, Qin H T, Yan Q H, Chai Z H, Liu J J, Wei X L, et al. Towards accurate post-training quantization for vision Tansformer. In: Proceedings of the 30th ACM International Conference on Multimedia. Lisbon, Portugal: ACM, 2022. 5380−5388 [34] Jeon Y, Lee C, Cho E, Ro Y. Mr.BiQ: Post-training non-uniform quantization based on minimizing the reconstruction error. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New Orleans, USA: IEEE, 2022. 12319−12328 [35] Frantar E, Ashkboos S, Hoefler T, Alistarh D. OPTQ: Accurate quantization for generative pre-trained Tansformers. In: Proceedings of the 11th International Conference on Learning Representations. Kigali, Rwanda: OpenReview.net, 2023. [36] Shao W Q, Chen M Z, Zhang Z Y, Xu P, Zhao L R, Li Z Q, et al. OmniQuant: Omnidirectionally calibrated quantization for large language models. In: Proceedings of the 12th International Conference on Learning Representations. Vienna, Austria: OpenReview.net, 2024. [37] Liu Z C, Zhao C S, Fedorov I, Soran B, Choudhary D, Krishnamoorthi R, et al. SpinQuant: LLM quantization with learned rotations. In: Proceedings of the 13th International Conference on Learning Representations. Singapore: OpenReview.net, 2025. [38] Oh S, Sim H, Kim J, Lee J. Non-uniform step size quantization for accurate post-training quantization. In: Proceedings of the 17th European Conference on Computer Vision. Tel Aviv, Israel: Springer, 2022. 658−673 [39] Lee E H, Miyashita D, Chai E, Murmann B, Wong S S. LogNet: Energy-efficient neural networks using logarithmic computation. In: Proceedings of the International Conference on Acoustics, Speech and Signal Processing (ICASSP). New Orleans, USA: IEEE, 2017. 5900−5904 [40] Lin Y, Zhang T Y, Sun P Q, Li Z, Zhou S C. FQ-ViT: Post-training quantization for fully quantized vision Tansformer. In: Proceedings of the 31st International Joint Conference on Artificial Intelligence. Vienna, Austria: IJCAI, 2022. 1173−1179 [41] Berger T. Rate-distortion theory. Wiley Encyclopedia of Telecommunications. Hoboken: John Wiley & Sons, 2003. [42] Yao Z W, Dong Z, Zheng Z C, Gholami A, Yu J L, Tan E, et al. HAWQ-V3: Dyadic neural network quantization. In: Proceedings of the 38th International Conference on Machine Learning. Virtual Event: PMLR, 2021. 11875−11886 [43] Song H Y, Dharmapurikar S, Turner J, Lockwood J. Fast Hash table lookup using extended bloom filter: An aid to network processing. ACM SIGCOMM Computer Communication Review, 2005, 35(4): 181−192 [44] Deng J, Dong W, Socher R, Li L J, Li K, Fei-Fei L. ImageNet: A large-scale hierarchical image database. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Miami, USA: IEEE, 2009. 248−255 [45] Lin T Y, Maire M, Belongie S, Hays J, Perona P, Ramanan D, et al. Microsoft COCO: Common objects in context. In: Proceedings of the 13th European Conference on Computer Vision. Zurich, Switzerland: Springer, 2014. 740−755 [46] Bisk Y, Zellers R, le Bras R, Gao J F, Choi Y. PIQA: Reasoning about physical commonsense in natural language. In: Proceedings of the 34th AAAI Conference on Artificial Intelligence. New York, USA: AAAI Press, 2020. 7432−7439 [47] Clark P, Cowhey I, Etzioni O, Khot T, Sabharwal A, Schoenick C, et al. Think you have solved question answering? Try ARC, the AI2 reasoning challenge. arXiv preprint arXiv: 1803.05457, 2018. [48] Clark C, Lee K, Chang M W, Kwiatkowski T, Collins M, Toutanova K. BoolQ: Exploring the surprising difficulty of natural yes/no questions. In: Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). Minneapolis, USA: Association for Computational Linguistics, 2019. 2924−2936 [49] Zellers R, Holtzman A, Bisk Y, Farhadi A, Choi Y. HellaSwag: Can a machine really finish your sentence? In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Florence, Italy: Association for Computational Linguistics, 2019. 4791−4800 [50] Sakaguchi K, le Bras R, Bhagavatula C, Choi Y. WinoGrande: An adversarial Winograd schema challenge at scale. Communications of the ACM, 2021, 64(9): 99−106 [51] Merity S, Xiong C, Bradbury J, Socher R. Pointer sentinel mixture models. In: Proceedings of the 5th International Conference on Learning Representations. Toulon, France: OpenReview.net, 2017. [52] Raffel C, Shazeer N, Roberts A, Lee K, Narang S, Matena M, et al. Exploring the limits of transfer learning with a unified text-to-text Tansformer. The Journal of Machine Learning Research, 2020, 21(1): Article No. 140 [53] Liu Y, Duan H D, Zhang Y H, Li B, Zhang S Y, Zhao W B, et al. MMBench: Is your multi-modal model an all-around player? In: Proceedings of the 18th European Conference on Computer Vision. Milan, Italy: Springer, 2025. 216−233 [54] Yang A, Li A F, Yang B S, Zhang B C, Hui B Y, Zheng B, et al. Qwen3 technical report. arXiv preprint arXiv: 2505.09388, 2025. [55] Bai S, Cai Y X, Chen R Z, Chen K Q, Chen X H, Cheng Z S, et al. Qwen3-VL technical report. arXiv preprint arXiv: 2511.21631, 2025. [56] Zheng X Y, Li Y Y, Chu H R, Feng Y, Ma X D, Wang Z N, et al. An empirical study of Qwen3 quantization. Visual Intelligence, 2025, 4(1): Article No. 11 [57] Liu Z H, Wang Y H, Han K, Zhang W, Ma S W, Gao W. Post-training quantization for vision Tansformer. In: Proceedings of the 35th International Conference on Neural Information Processing Systems. Virtual Event: Curran Associates, Inc., 2021. Article No. 2152 [58] Li R D, Wang Y, Liang F, Qin H W, Yan J J, Fan R. Fully quantized network for object detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Long Beach, USA: IEEE, 2019. 2805−2814 [59] Wightman R. PyTorch Image Models [Online], available: https://github.com/huggingface/pytorch-image-models, July 3, 2026 [60] Ma Y X, Li H X, Zheng X W, Ling F, Xiao X F, Wang R, et al. AffineQuant: Affine transformation quantization for large language models. In: Proceedings of the 12th International Conference on Learning Representations. Vienna, Austria: OpenReview.net, 2024. [61] Chen M Z, Liu Y, Wang J H, Bin Y, Shao W Q, Luo P. PrefixQuant: Eliminating outliers by prefixed tokens for large language models quantization. IEEE Transactions on Pattern Analysis and Machine Intelligence, DOI: 10.1109/TPAMI.2026.3711802 [62] Lin J, Tang J M, Tang H T, Yang S, Chen W M, Wang W C, et al. AWQ: Activation-aware weight quantization for on-device LLM compression and acceleration. In: Proceedings of the 7th Annual Conference on Machine Learning and Systems. Santa Clara, USA: MLSys, 2024. -
计量
- 文章访问数: 13
- HTML全文浏览量: 8
- 被引次数: 0
下载: