|国家预印本平台
| 注册
首页|大视觉Transformer参数高效微调中的静默训练崩溃:ln(N)特征及其被最优精度跟踪机制的掩盖

大视觉Transformer参数高效微调中的静默训练崩溃:ln(N)特征及其被最优精度跟踪机制的掩盖

兰奥 陈俊言

istic_logo国家预印本平台

大视觉Transformer参数高效微调中的静默训练崩溃:ln(N)特征及其被最优精度跟踪机制的掩盖

Silent Training Collapse in Parameter-Efficient Fine-Tuning of Large Vision Transformers: ln(N) Signature and Its Masking by Best-Accuracy Tracking

兰奥 1陈俊言2

作者信息

  • 1. 南华大学
  • 2. 湖南大学
  • 折叠

摘要

研究目的:刻画大视觉Transformer(ViT)在小样本工业计算机断层扫描(CT)缺陷识别任务中,采用参数高效微调(PEFT)时出现的静默训练崩溃失效模式,并找出会掩盖该失效模式的评估方案。 研究方法:基于包含673张图像的肺部CT数据集,使用LoRA对DINOv2 ViT‑L/14(参数量3.04亿)开展微调;通过单变量消融实验定位根本成因;在LIDC‑IDRI数据集(6691张切片)上复现该失效特征;采用2×2匹配学习率对照组,检验平均绝对误差(MAE)域自适应出现的表观性能下降现象。 研究结果:模型发生静默崩溃,退化为不具备学习能力的多数类预测器(准确率32.4%,交叉熵1.370,略低于均匀基准值ln(4)≈1.3863)。该崩溃现象被最优准确率追踪机制掩盖:最优轮次快照准确率为51.11%,掩盖了最终模型仅达到32.4%的先验基准水平。仅将学习率从10⁻³下调至2×10⁻⁴即可避免崩溃(冻结归一化层的ViT‑L模型准确率可达94.20%);文献中报道的LayerNorm带来的11.15%表观性能优势,实际缩减为无统计学意义的‑1.19%。在LIDC‑IDRI数据集上,ViT‑B/14在三分之二的交叉验证折上发生崩溃(准确率62.24%,F1值为0,AUC为0.5),修正评估方案后模型性能恢复至96.85%。匹配学习率对照实验表明,文献报道的MAE指标下降28.7个百分点属于学习率导致的伪现象;采用稳定学习率后,基于MAE自适应的LoRA模型准确率可达92.6%。 研究局限性:ln(N)基准阈值仅适用于分类任务;本研究仅测试了LoRA方法;所用数据集样本规模为数百至数千级别;工业场景验证依赖合成仿真数据。 研究结论:公开报道的归一化效果是训练不稳定造成的伪现象,并非LayerNorm带来的真实增益。研究人员应当同时报告最终轮次指标与最优轮次指标;将损失收敛至ln(N)视作训练崩溃的判断依据;并使用适配模型规模的学习率。

Abstract

Objective: To characterize a silent training-collapse failure mode in parameter-efficient fine-tuning (PEFT) of large vision transformers (ViTs) for small-sample industrial computed tomography (CT) defect recognition, and to identify the evaluation protocol that masks it. Methods: A DINOv2 ViT-L/14 (304M) was fine-tuned with LoRA on a 673-image lung CT dataset; a single-variable ablation isolated the root cause; the signature was reproduced on LIDC-IDRI (6,691 slices); and a matched 2×2 learning-rate control tested an apparent MAE domain-adaptation degradation. Results: The model silently collapsed to a non-learning majority-class predictor (accuracy 32.4%, cross-entropy 1.370, just below the uniform floor ln(4) ≈ 1.3863). The collapse was masked by best-accuracy tracking: a best-epoch snapshot of 51.11% hid a final model at the 32.4% prior. Lowering the learning rate from 10⁻³ to 2×10⁻⁴ alone prevented collapse (frozen-norm ViT-L reached 94.20%), and the apparent +11.15% LayerNorm advantage reduced to an insignificant −1.19%. On LIDC-IDRI, a ViT-B/14 collapsed on 2/3 folds (62.24%, F1 = 0, AUC = 0.5), and fixing the protocol recovered 96.85%. A matched learning-rate control showed a reported −28.7 percentage-point MAE degradation to be a learning-rate artifact, with MAE-adapted LoRA reaching 92.6% at the stabilized rate. Limitations: The ln(N) floor is specific to classification; only LoRA was tested; datasets number hundreds to thousands of samples; and the industrial validation relies on synthetic simulated data. Conclusions: The headline normalization effects are artifacts of training instability, not genuine LayerNorm effects. Practitioners should report final-epoch metrics alongside best-epoch metrics, treat loss convergence to ln(N) as a collapse diagnostic, and use a scale-aware learning rate.

关键词

工业计算机断层扫描/无损检测/核燃料包壳/参数高效微调/训练崩溃/训练稳定性/视觉Transformer

Key words

Industrial Computed Tomography/ Non-Destructive Testing/ Nuclear Fuel Cladding/ Parameter-Efficient Fine-Tuning/ Training Collapse/ Training Stability/ Vision Transformers

引用本文复制引用

兰奥,陈俊言.大视觉Transformer参数高效微调中的静默训练崩溃:ln(N)特征及其被最优精度跟踪机制的掩盖[EB/OL].(2026-09-01)[2026-09-01].https://sinoxiv.napstic.cn/article/26161450.

学科分类

计算技术、计算机技术
首发时间 2026-09-01 08:37:45
下载量:0
|
点击量:5
段落导航相关论文