• 中文核心
  • EI
  • 中国科技核心
  • Scopus
  • CSCD
  • 英国科学文摘

留言板

尊敬的读者、作者、审稿人, 关于本刊的投稿、审稿、编辑和出版的任何问题, 您可以本页添加留言。我们将尽快给您答复。谢谢您的支持!

姓名
邮箱
手机号码
标题
留言内容
验证码

图像数据集知识效能评估方法

丁昕苗 李好男 刘雨帆 刘彤飞 张克卿 原春锋 李兵 王坚 胡卫明

丁昕苗, 李好男, 刘雨帆, 刘彤飞, 张克卿, 原春锋, 李兵, 王坚, 胡卫明. 图像数据集知识效能评估方法. 自动化学报, xxxx, xx(x): x−xx doi: 10.16383/j.aas.c250457
引用本文: 丁昕苗, 李好男, 刘雨帆, 刘彤飞, 张克卿, 原春锋, 李兵, 王坚, 胡卫明. 图像数据集知识效能评估方法. 自动化学报, xxxx, xx(x): x−xx doi: 10.16383/j.aas.c250457
Ding Xin-Miao, Li Hao-Nan, Liu Yu-Fan, Liu Tong-Fei, Zhang Ke-Qing, Yuan Chun-Feng, Li Bing, Wang Jian, Hu Wei-Ming. Knowledge efficiency evaluation method for image datasets. Acta Automatica Sinica, xxxx, xx(x): x−xx doi: 10.16383/j.aas.c250457
Citation: Ding Xin-Miao, Li Hao-Nan, Liu Yu-Fan, Liu Tong-Fei, Zhang Ke-Qing, Yuan Chun-Feng, Li Bing, Wang Jian, Hu Wei-Ming. Knowledge efficiency evaluation method for image datasets. Acta Automatica Sinica, xxxx, xx(x): x−xx doi: 10.16383/j.aas.c250457

图像数据集知识效能评估方法

doi: 10.16383/j.aas.c250457 cstr: 32138.14.j.aas.c250457
基金项目: 北京市自然科学基金(JQ24022), 国家自然科学基金(61876100, 62072286, 62372451, 62192785, 62372082, 62192782, 62532015), 中国人工智能学会−蚂蚁集团联合研究基金(CAAI-MYJJ2024-02), 中国科协青年人才托举工程(2024QNRC001)资助
详细信息
    作者简介:

    丁昕苗:山东工商学院信息与电子工程学院教授. 2013年获得中国矿业大学(北京)博士学位. 主要研究方向为图像与视频理解, 机器学习和网络安全. E-mail: dingxinmiao@126.com

    李好男:山东工商学院信息与电子工程学院硕士研究生. 主要研究方向为机器学习, 模式识别. E-mail: 2023420059@sdtbu.edu.cn

    刘雨帆:中国科学院自动化研究所副研究员. 2023年获得中国科学院大学博士学位. 主要研究方向为计算机视觉, 模型压缩和视频理解. 本文通信作者. E-mail: liuyufan@ia.ac.cn

    刘彤飞:中国科学院自动化研究所博士研究生. 主要研究方向为数据集压缩. E-mail: liutongfei@ia.ac.cn

    张克卿:中国科学院自动化研究所博士研究生. 主要研究方向为大语言模型安全, 知识蒸馏. E-mail: zhangkeqing2025@ia.ac.cn

    原春锋:中国科学院自动化研究所研究员. 2010年获得中国科学院自动化研究所博士学位. 主要研究方向为模式识别, 计算机视觉. E-mail: yuanchunfeng@ia.ac.cn

    李兵:中国科学院自动化研究所研究员. 2009年获得北京交通大学博士学位. 主要研究方向为视频理解, 颜色恒常性, 视觉显著性, 多实例学习和网络内容安全. E-mail: libing@ia.ac.cn

    王坚:中国科学院自动化研究所工程师. 主要研究方向为机器学习, 计算机视觉. E-mail: wangjian@ia.ac.cn

    胡卫明:中国科学院自动化研究所研究员. 1998年获得浙江大学博士学位. 主要研究方向为视觉运动分析, 网络不良信息识别和网络入侵检测. E-mail: huweiming@ia.ac.cn

Knowledge Efficiency Evaluation Method for Image Datasets

Funds: Supported by Beijing Natural Science Foundation (JQ24022), National Natural Science Foundation of China (61876100, 62072286, 62372451, 62192785, 62372082, 62192782, 62532015), CAAI-Ant Group Research Fund (CAAI-MYJJ 2024-02), and Young Elite Scientists Sponsorship Program by CAST (2024QNRC001)
More Information
    Author Bio:

    DING Xin-Miao Professor at the School of Information and Electronic Engineering, Shandong Technology and Business University. She received her Ph.D. degree from China University of Mining and Technology, Beijing, in 2013. Her research interests include image and video understanding, machine learning, and internet security

    LI Hao-Nan Master student at the School of Information and Electronic Engineering, Shandong Technology and Business University. His research interests include machine learning and pattern recognition.This authorcontributed equally to the first author

    LIU Yu-Fan Associate researcher at the Institute of Automation, Chinese Academy of Sciences. She received her Ph.D. degree from University of Chinese Academy of Sciences in 2023. Her research interests include computer vision, model compression, and video understanding.Corresponding author of this paper

    LIU Tong-Fei Ph.D. candidate at the Institute of Automation, Chinese Academy of Sciences. His research interest is dataset compression

    ZHANG Ke-Qing Ph.D. candidate at the Institute of Automation, Chinese Academy of Sciences. His research interests include large language model safety and knowledge distillation

    YUAN Chun-Feng Researcher at the Institute of Automation, Chinese Academy of Sciences. She received her Ph.D. degree from the Institute of Automation, Chinese Academy of Sciences in 2010. Her research interests include pattern recognition and computer vision

    LI Bing Researcher at the Institute of Automation, Chinese Academy of Sciences. He received his Ph.D. degree from Beijing Jiaotong University in 2009. His research interests include video understanding, color constancy, visual saliency, multi-instance learning, and web content security

    WANG Jian Engineer at the Institute of Automation, Chinese Academy of Sciences. His research interests include machine learning and computer vision

    HU Wei-Ming Researcher at the Institute of Automation, Chinese Academy of Sciences. He received his Ph.D. degree from Zhejiang University in 1998. His research interests include visual motion analysis, recognition of web objectionable information, and network intrusion detection

  • 摘要: 当前数据交易过程面临价值评估难题, AI模型训练缺乏数据选择标准. 其原因在于不同数据集在信息密度、质量结构与任务适配性方面呈现出显著差异, 而现有方法多聚焦于数据质量或模型性能的单一维度, 缺乏对数据集信息密度与多维价值特性的系统性量化框架, 限制了其在任务导向场景下数据集的科学选择与高效利用. 为此, 提出数据集知识效能评估框架, 即一种综合衡量数据集有效知识含量及其在特定任务中适应能力的量化方法, 并基于知识提炼、知识质量分析与知识任务适配构建三维评估框架: 1) 知识提炼: 将无损数据集蒸馏算法引入效能评估, 通过知识密度比量化数据集有效知识承载能力; 2) 知识质量分析: 构建多维质量指标体系, 深入分析和评估数据的知识质量; 3) 知识任务适配: 创新性提出任务驱动的评估适配机制, 通过契合度分析与权重优化实现对特定任务适应性的评估. 在CIFAR-10/100、VGGFace2、ImageNet-1K和NABirds等典型数据集上的实验结果表明, 所提方法能够较为准确地评估知识密度与任务适配性, 为数据集选择与应用提供理论依据和技术支撑.
  • 图  1  知识效能评估框架

    Fig.  1  Knowledge efficiency evaluation framework

    图  2  数据集每类蒸馏前后IPC变化

    Fig.  2  Per-class IPC changes of datasets before and after distillation

    图  3  语义层面标注错误

    Fig.  3  Semantic-level annotation errors

    图  4  数据集类内多样性统计

    Fig.  4  Intra-class diversity statistics of the dataset

    图  5  数据集灰度均值与标准差分布

    Fig.  5  Distribution of grayscale mean and standard deviation of the dataset

    图  6  数据集信息熵分布

    Fig.  6  Information entropy distribution of the dataset

    表  1  不同数据评估范式的对比与本文方法定位

    Table  1  Comparison of different data evaluation paradigms and positioning of our method

    评估范式 评估对象 核心问题与边界 知识承载 多维质量 任务适配
    数据质量评估 数据本身 关注数据是否规范及单一维度是否达标, 通常难以刻画数据集整体知识结构
    数据价值评估 经济层 关注数据的经济或业务收益, 但强依赖具体场景与收益假设, 可复用性弱
    样本贡献度 单样本 刻画单样本对训练目标的边际贡献, 计算代价高且难形成数据集层整体结论
    迁移适配预测 表征层 评估预训练表征的可迁移性, 主要反映模型与任务匹配, 难解释数据本体与质量结构
    DKEE(本文) 数据集层 统一刻画数据集中可被提炼的有效知识含量, 并在不同任务场景下进行多维质量分析,
    从而支持训练前的数据选择与治理决策
    下载: 导出CSV

    表  2  数据集蒸馏前后整体规模知识密度对比

    Table  2  Overall-scale knowledge density comparison of datasets before and after distillation

    数据集 类别数$ K $ 原始样本 蒸馏后样本 $ g(D,\; T) $
    MNIST[60] 10 60000 250 0.004
    CIFAR-10[3] 10 50000 8750 0.175
    CIFAR-100[3] 100 50000 7500 0.150
    TinyImageNet[61] 200 100000 8000 0.080
    ImageNet[4] 1000 1300000 1100000 0.846
    NABirds[65] 555 44400 33300 0.750
    Places365[62] 365 1788500 1168000 0.653
    VGGFace2[63] 9, 131 3287160 2739300 0.833
    下载: 导出CSV

    表  3  数据集完整性与标签一致性检测结果

    Table  3  Dataset completeness and label consistency detection results consistency detection results

    数据集 总图像数 重复图像数 有效率(%) 正确标签数 标签一致性(%)
    MNIST 70000 0 100.00 70000 100.00
    CIFAR-10 60000 0 100.00 60000 100.00
    CIFAR-100 60000 26 99.96 60000 100.00
    TinyImageNet 130000 0 100.00 130000 100.00
    Places365 2204960 0 100.00 2204960 100.00
    VGGFace2 3311286 0 100.00 3311286 100.00
    ImageNet 1431167 0 100.00 1431167 100.00
    NABirds 48562 0 100.00 48562 100.00
    下载: 导出CSV

    表  4  数据集噪声检测结果

    Table  4  Dataset noise detection results

    数据集 总样本数 正常样本数 噪声样本数 噪声比例(%)
    MNIST 60000 42238 17762 29.60
    CIFAR10 50000 33078 16922 33.84
    CIFAR100 50000 33269 16731 33.46
    TinyImageNet 100000 89462 10538 10.54
    Places365 1803460 1801968 1492 0.08
    VGGFace2 3141890 3140174 1716 0.05
    ImageNet 1281167 1279738 1429 0.12
    NABirds 23912 6021 17891 74.82
    下载: 导出CSV

    表  5  数据集类间多样性测量结果

    Table  5  Dataset inter-class diversity measurement results

    数据集 类间多样性均值
    MNIST 0.004162
    CIFAR-10 0.002321
    CIFAR-100 0.009125
    TinyImageNet 0.004166
    Places365 0.011322
    ImageNet 0.012894
    VGGFace2 0.029938
    NABirds 0.009781
    下载: 导出CSV

    表  6  经典模型在不同数据集上的测试准确率(%)

    Table  6  Test accuracy of classic models on different datasets (%)

    模型 MNIST CIFAR-10 CIFAR-100 TinyImageNet Places365 VGGFace2 ImageNet NABirds
    VGG16[55] 99.20 93.25 73.50 56.80 55.00 89.65 71.50 74.90
    ResNet18[54] 99.50 93.02 77.10 58.70 54.70 91.20 69.80 76.21
    ResNet50[54] 99.60 93.62 79.20 61.50 55.20 92.75 76.15 79.55
    DenseNet121[56] 99.70 95.04 79.30 64.20 56.20 93.10 74.90 80.30
    MobileNetV2[58] 99.40 94.50 75.30 59.10 54.00 90.50 71.80
    EfficientNet-B0[57] 99.70 96.00 80.10 66.00 57.00 94.00 77.10 63.70
    PyramidNet[59] 99.75 97.25 83.54 68.00 58.70 95.50 82.00
    平均准确率 99.55 94.67 78.29 62.04 55.83 92.39 74.75 74.93$ ^* $
    −− 表示该数据集未公开基准结果
    下载: 导出CSV

    表  7  歧义性、领域偏移与泄露比测量结果(%)

    Table  7  Measurement results of ambiguity domain shift and leakage ratio (%)

    数据集 歧义性 领域偏移 泄露比
    MNIST 99.88 0.120 4.99
    CIFAR-10 97.75 0.170 5.15
    CIFAR-100 99.04 0.210 4.49
    TinyImageNet 95.36 0.010 8.74
    Places365 99.69 0.036 73.35
    ImageNet 99.64 0.020 62.82
    VGGFace2 99.85 0.040 78.43
    NABirds 99.67 0.370 2.39
    下载: 导出CSV

    表  8  预训练模型在不同数据集上的迁移准确率(%)

    Table  8  Transfer accuracy of pre-trained models on different datasets (%)

    源模型 目标数据集
    MNIST CIFAR-10 CIFAR-100 TinyImageNet Places365 ImageNet VGGFace2 NABirds LFW[64] CUB[66]
    MNIST 83.68 7.42 10.98 0 0 0 0 0 0.81 0.78
    CIFAR-10 25.19 91.90 72.81 25.82 61.65 79.41 52.47 65.00 1.59 0.67
    CIFAR-100 5.30 8.35 76.92 10.17 31.55 56.00 23.96 42.66 2.66 0.86
    TinyImageNet 1.39 3.91 20.52 44.14 26.26 62.34 20.83 36.87 4.01 1.85
    Places365 0.45 0.68 2.51 8.81 45.60 39.30 8.78 26.19 5.18 2.59
    ImageNet 0.13 0.29 1.50 7.85 27.61 69.15 5.56 38.71 8.80 28.39
    VGGFace2 0.22 0.26 0.33 1.28 9.64 23.59 98.00 17.95 98.46 1.69
    NABirds 0.24 0.33 0.42 1.35 2.60 26.85 1.59 73.63 18.29 80.84
    下载: 导出CSV

    表  9  不同任务类型下的数据集适配效能

    Table  9  Dataset adaptation efficiency under different task types

    数据集 通用任务$ V_{\text{task}}(D,\; T) $ 特定任务$ V_{\text{task}}(D,\; T) $
    MNIST 0.420 0.420
    CIFAR-10 0.439 0.439
    CIFAR-100 0.456 0.456
    TinyImageNet 0.519 0.519
    Places365 0.537 0.537
    ImageNet 0.546 0.546
    VGGFace2 0.420 0.477
    NABirds 0.510 0.536
    下载: 导出CSV

    A1  NABirds数据集中困难样本剔除前后的准确率对比

    A1  Accuracy comparison before and after removing hard samples from NABirds dataset

    实验设置 分类准确率(%)
    原始数据集 56.0
    剔除20% 困难样本 44.0
    下载: 导出CSV

    B1  基于ResNet18的跨数据集迁移分类准确率(%)

    B1  Cross-dataset transfer classification accuracy based on ResNet18 (%)

    源模型 目标数据集
    MNIST CIFAR-10 CIFAR-100 TinyImageNet Places365 ImageNet VGGFace2 NABirds
    MNIST 99.36 17.37 3.31 1.10 0.27 0.10 0.29 0.26
    CIFAR-10 44.11 81.13 14.64 4.57 0.73 0.29 0.32 0.48
    CIFAR-100 62.41 67.56 79.26 17.80 2.34 1.29 0.30 0.42
    TinyImageNet 31.38 55.80 33.59 99.36 17.97 16.04 2.38 1.52
    Places365 91.60 68.84 39.47 41.81 52.17 32.92 16.50 3.69
    ImageNet 94.64 77.74 52.80 58.05 38.80 63.61 25.24 22.14
    VGGFace2 94.59 78.20 52.56 57.83 38.91 63.51 25.65 0.23
    NABirds 54.28 27.35 8.23 3.50 4.47 1.81 1.23 22.41
    下载: 导出CSV

    C1  COCO数据集在通用分类任务下的多维评估结果

    C1  Multi-dimensional evaluation results of COCO dataset for general classification task

    指标 结果
    知识密度 95%
    完整度 100%
    标签一致性 100%
    噪声比 93.08%
    类内多样性 0.034485
    类间多样性 0.01
    均值 0.0018
    方差 0.0022
    0.7269
    基线模型性能 36.28%
    歧义性 0.9984
    领域偏移 0.059
    泄露比 11.03%
    知识效能 4.47
    下载: 导出CSV

    D1  单次评估流程的耗时统计(h)

    D1  Time cost statistics of single evaluation process (h)

    数据集 知识提炼 质量分析
    MNIST 1 0.5
    CIFAR-10 2 0.5
    TinyImageNet 8 1
    Places365 15 15
    VGGFace2 30 20
    ImageNet 20 20
    NABirds 5 1
    下载: 导出CSV
  • [1] 关于促进数据产业高质量发展的指导意见. 国家发展改革委, 国家数据局, 教育部, 财政部, 金融监管总局, 中国证监会.

    Guiding opinions on promoting the high-quality development of the data industry. [Online], available: https://www.ndrc.gov.cn/xxgk/zcfb/tz/202412/t20241230_1395341.html, September 1 2025.
    [2] 关于完善数据流通安全治理更好促进数据要素市场化价值化的实施方案. 国家发展改革委, 国家数据局, 中央网信办, 工业和信息化部, 公安部, 市场监管总局.

    Implementation plan for improving data circulation security governance and better promoting the marketization and valorization of data elements. [Online], available: https://www.ndrc.gov.cn/xxgk/zcfb/tz/202501/t20250115_1395692.html, September 1 2025.
    [3] Alex Krizhevsky and Geoffrey Hinton. Learning multiple layers of features from tiny images. Technical report, Department of Computer Science, University of Toronto, Canada, 2009.
    [4] Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A largescale hierarchical image database. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 248–255, Miami, USA, 2009. IEEE.
    [5] Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. arXiv: 1503.02531, 2015.
    [6] Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio. Fitnets: Hints for thin deep nets. arXiv: 1412.6550, 2014.
    [7] Sergey Zagoruyko and Nikos Komodakis. Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer. arXiv: 1612.03928, 2016.
    [8] Frederick Tung and Greg Mori. Similaritypreserving knowledge distillation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1365–1374, Seoul, Korea, 2019. IEEE.
    [9] Yufan Liu, Jiajiong Cao, Bing Li, Chunfeng Yuan, Weiming Hu, and Yangxi Li. Knowledge distillation via instance relationship graph. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7096–7104, Long Beach, USA, 2019. IEEE.
    [10] Yonglong Tian, Dilip Krishnan, and Phillip Isola. Contrastive representation distillation. arXiv: 1910.10699, 2019.
    [11] Byeongho Heo, Jeesoo Kim, Sangdoo Yun, Hyojin Park, Nojun Kwak, and Jin Young Choi. A comprehensive overhaul of feature distillation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1921–1930, Seoul, Korea, 2019. IEEE.
    [12] Yufan Liu, Jiajiong Cao, Bing Li, Weiming Hu, Jingting Ding, and Liang Li. Cross-architecture knowledge distillation. International Journal of Computer Vision, 2024, 132(10): 3396−3411
    [13] Zhengyuan Yang, Jingen Liu, Jing Huang, Xiaodong He, Tao Mei, and Chenliang Xu. Crossmodal contrastive distillation for instructional activity anticipation. In Proceedings of the 26th International Conference on Pattern Recognition, pages 5002–5009, Montreal, Canada, 2022. IEEE.
    [14] Jiahui Yu, Linjie Yang, Ning Xu, Jianchao Yang, and Thomas Huang. Slimmable neural networks. arXiv: 1812.08928, 2018.
    [15] Yufan Liu, Jiajiong Cao, Bing Li, Weiming Hu, and Stephen Maybank. Learning to explore distillability and sparsability: A joint framework for model compression. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023, 45(3): 3378−3395
    [16] Yufan Liu, Jiajiong Cao, Weiming Bai, Bing Li, and Weiming Hu. Learning from the raw domain: Cross modality distillation for compressed video action recognition. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, pages 1–5, Rhodes Island, Greece, 2023. IEEE.
    [17] Zhiyuan Fang, Jianfeng Wang, Xiaowei Hu, Lijuan Wang, Yezhou Yang, and Zicheng Liu. Compressing visual-linguistic model via knowledge distillation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1428–1438, Virtual Event, 2021. IEEE.
    [18] Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. Advances in Neural Information Processing Systems, 2012, 25: 1097−1105
    [19] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, and Aidan N Gomez. Attention is all you need. Advances in Neural Information Processing Systems, 2017, 30: 5998−6008 doi: 10.1007/978-3-031-84300-6_13
    [20] Tongzhou Wang, Jun-Yan Zhu, Antonio Torralba, and Alexei A Efros. Dataset distillation. arXiv: 1811.10959, 2018.
    [21] Bo Zhao, Konda Reddy Mopuri, and Hakan Bilen. Dataset condensation with gradient matching. arXiv: 2006.05929, 2020.
    [22] Timothy Nguyen, Zhourong Chen, and Jaehoon Lee. Dataset meta-learning from kernel ridgeregression. arXiv: 2011.00050, 2020.
    [23] Timothy Nguyen, Roman Novak, Lechao Xiao, and Jaehoon Lee. Dataset distillation with infinitely wide convolutional networks. Advances in Neural Information Processing Systems, 2021, 34: 5186−5198
    [24] Bo Zhao and Hakan Bilen. Dataset condensation with distribution matching. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 6514–6523, Waikoloa, USA, 2023. IEEE.
    [25] Ganlong Zhao, Guanbin Li, Yipeng Qin, and Yizhou Yu. Improved distribution matching for dataset condensation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7856–7865, Vancouver, Canada, 2023. IEEE.
    [26] George Cazenavette, Tongzhou Wang, Antonio Torralba, Alexei A Efros, and Jun-Yan Zhu. Dataset distillation by matching training trajectories. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4750–4759, New Orleans, USA, 2022. IEEE.
    [27] Zeyuan Yin, Eric Xing, and Zhiqiang Shen. Squeeze, recover and relabel: dataset condensation at imagenet scale from a new perspective. In Advances in Neural Information Processing Systems, volume 36, pages 73582–73603, 2023.
    [28] Xin Zhang, Jia-Wei Du, Ping Liu, and Joey Tianyi Zhou. Breaking class barriers: efficient dataset distillation via inter-class feature compensator. arXiv: 2408.06927, 2024.
    [29] Ming-Yang Chen, Jia-Wei Du, Bo Huang, Yi Wang, Xiao-Bo Zhang, and Wei Wang. Influence-guided diffusion for dataset distillation. In Proceedings of the 13th International Conference on Learning Representations, Singapore, Singapore, 2025. OpenReview.net.
    [30] Ding Qi, Jian Li, Jun-Yao Gao, Shu-Guang Dou, Ying Tai, and Jian-Long Hu. Towards universal dataset distillation via task-driven diffusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10557–10566, Nashville, USA, 2025. IEEE.
    [31] Ziyao Guo, Kai Wang, George Cazenavette, Hui Li, Kaipeng Zhang, and Yang You. Towards lossless dataset distillation via diffculty-aligned trajectory matching. arXiv: 2310.05773, 2023.
    [32] Shitong Shao, Zikai Zhou, Huanran Chen, and Zhiqiang Shen. Elucidating the design space of dataset condensation. Advances in Neural Information Processing Systems, 2024, 37: 99161−99201
    [33] Andrea Schioppa. Effcient sketches for training data attribution and studying the loss landscape. Advances in Neural Information Processing Systems, 2024, 37: 37692−37735
    [34] Tongfei Liu, Yufan Liu, Bing Li, Weiming Hu, Yuming Li, and Chenguang Ma. Noise-optimized distribution distillation for dataset condensation. In Proceedings of the 33rd ACM International Conference on Multimedia, pages 10352–10360, Melbourne, Australia, 2025. ACM.
    [35] Ping Liu and Jia-Wei Du. The evolution of dataset distillation: toward scalable and generalizable solutions. arXiv: 2502.05673, 2025.
    [36] Curtis G Northcutt, Anish Athalye, and Jonas Mueller. Pervasive label errors in test sets destabilize machine learning benchmarks. arXiv: 2103.14749, 2021.
    [37] Bingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan, Albert Gordo, and Jiashi Feng. Decoupling representation and classifier for longtailed recognition. arXiv: 1910.09217, 2019.
    [38] Shahadat Uddin and Haohui Lu. Dataset metalevel and statistical features affect machine learning performance. Scientific Reports, 2024, 14(1): 1670. doi: 10.1038/s41598-024-51825-x
    [39] Tin Kam Ho. A data complexity analysis of comparative advantages of decision forest constructors. Pattern Analysis and Applications, 2002, 5(2): 102−112
    [40] Cuong Nguyen, Tal Hassner, Matthias Seeger, and Cedric Archambeau. Leep: A new measure to evaluate transferability of learned representations. In Proceedings of the International Conference on Machine Learning, pages 7294–7305, Virtual Event, 2020. PMLR.
    [41] Kaichao You, Yong Liu, Jianmin Wang, and Mingsheng Long. Logme: Practical assessment of pre-trained models for transfer learning. In Proceedings of the International Conference on Machine Learning, pages 12133–12143, Virtual Event, 2021. PMLR.
    [42] Swabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, and Noah A Smith. Dataset cartography: Mapping and diagnosing datasets with training dynamics. arXiv: 2009.10795, 2020.
    [43] Mengyu Dai, Amir Hossein Raffee, Aashish Jain, and Joshua Correa. Evaluating transferability in retrieval tasks: An approach using mmd and kernel methods. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22390–22400, Seattle, USA, 2024. IEEE.
    [44] 丁文文王戈王坛 王飞跃李娟娟, 秦蕊. 基于web3的去中心化自治组织与运营新框架. 自动化学报, 49(AAS-CN-2022-0753): 985, 2023. doi: 10.16383/j.aas.c220753.
    [45] 毛建旭别克扎提·巴赫提杜锐 王耀南张辉, 严星雨. 面向源网荷的智能化数据协同推断技术研究综述. 自动化学报, 51(AAS-CN-2025-0203): 2387, 2025. doi: 10.16383/j.aas.c250203.
    [46] 袁勇倪晓春 王飞跃欧阳丽炜, 王帅. 智能合约: 架构及进展. 自动化学报, 45(zdhxb-45-3-445): 445, 2019. doi: 10.16383/j.aas.c180586.
    [47] Harold Hotelling. Analysis of a complex of statistical variables into principal components. Journal of Educational Psychology, 1933, 24(6): 417−441 doi: 10.1037/h0070888
    [48] Laurens Van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. Journal of Machine Learning Research, 2008, 9(11): 2579−2605
    [49] Martin Ester, Hans-Peter Kriegel, Jorg Sander, and Xiaowei Xu. A density-based algorithm for discovering clusters in large spatial databases with noise. In Proceedings of the 2nd International Conference on Knowledge Discovery and Data Mining, pages 226–231, Portland, USA, 1996. AAAI Press.
    [50] Jianhua Lin. Divergence measures based on the shannon entropy. IEEE Transactions on Information Theory, 1991, 37(1): 145−151 doi: 10.1109/18.61115
    [51] Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, and Sandhini Agarwal. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning, pages 8748–8763, Virtual Event, 2021. PMLR.
    [52] Diederik P Kingma and Max Welling. Autoencoding variational bayes. arXiv: 1312.6114, 2013.
    [53] Solomon Kullback and Richard A Leibler. On information and suffciency. The Annals of Mathematical Statistics, 1951, 22(1): 79−86
    [54] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 770–778, Las Vegas, USA, 2016. IEEE.
    [55] Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv: 1409.1556, 2014.
    [56] Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger. Densely connected convolutional networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4700–4708, Honolulu, USA, 2017. IEEE.
    [57] Mingxing Tan and Quoc Le. Effcientnet: Rethinking model scaling for convolutional neural networks. In Proceedings of the International Conference on Machine Learning, pages 6105–6114, Long Beach, USA, 2019. PMLR.
    [58] Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. Mobilenetv2: Inverted residuals and linear bottlenecks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4510–4520, Salt Lake City, USA, 2018. IEEE.
    [59] Dongyoon Han, Jiwhan Kim, and Junmo Kim. Deep pyramidal residual networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 5927–5935, Honolulu, USA, 2017. IEEE.
    [60] Yann LeCun, Leon Bottou, Yoshua Bengio, and Patrick Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 1998, 86(11): 2278−2324 doi: 10.1109/5.726791
    [61] Yann Le and Xuan Yang. Tiny imagenet visual recognition challenge. Technical report, Stanford University, USA, 2015.
    [62] Alejandro Lopez-Cifuentes, Marcos EscuderoVinolo, Jesus Bescos, and Alvaro Garcia-Martin. Semantic-aware scene recognition. Pattern Recognition, 2020, 102: 107256 doi: 10.1016/j.patcog.2020.107256
    [63] Qiong Cao, Li Shen, Weidi Xie, Omkar M Parkhi, and Andrew Zisserman. Vggface2: A dataset for recognising faces across pose and age. In Proceedings of the 13th IEEE International Conference on Automatic Face and Gesture Recognition, pages 67–74, Xi’an, China, 2018. IEEE.
    [64] Gary B Huang, Marwan Mattar, Tamara Berg, and Eric Learned-Miller. Labeled faces in the wild: A database for studying face recognition in unconstrained environments. In Workshop on Faces in ’Real-Life’ Images: Detection, Alignment, and Recognition, pages 1–11, Marseille, France, 2008.
    [65] Grant Van Horn, Steve Branson, Ryan Farrell, Scott Haber, Jessie Barry, and Panos Ipeirotis. Building a bird recognition app and large scale dataset with citizen scientists: The fine print in fine-grained dataset collection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 595–604, Boston, USA, 2015. IEEE.
    [66] Catherine Wah, Steve Branson, Peter Welinder, Pietro Perona, and Serge Belongie. The caltechucsd birds-200-2011 dataset. Technical Report CNS-TR-2011-001, California Institute of Technology, USA, 2011.
  • 加载中
计量
  • 文章访问数:  7
  • HTML全文浏览量:  5
  • 被引次数: 0
出版历程
  • 收稿日期:  2025-09-18
  • 录用日期:  2026-03-04
  • 网络出版日期:  2026-08-04

目录

    /

    返回文章
    返回