元认知门控式动态检索的领域相关收益边界与可迁移诊断框架
Controlled Evaluation of Metacognitive-Gated Retrieval: Domain-Dependent Payoff Boundaries and a Transferable Diagnostic Framework
摘要
元认知门控式动态检索广泛用于大语言模型幻觉治理,但其依赖三个未经系统检验的假设:自评置信度与事实正确性正相关(H1)、阈值调优后可用于决策(H2)、与上下文锚定分层组合具备可加性(H3)。本文在 4 个中小开源指令模型(Qwen2.5-1.5B/3B/7B、Llama-3.2-3B)上,以全量检索为基线,在合成虚构语料与按首次公开披露日分层的真实 SEC 去污染语料上做全阈值精确重放(t=0…101),并配同触发率随机对照与 oracle 上界。结果表明:自陈置信度对"是否知道"的判别力 AUC 仅 0.341–0.500(两个模型显著反向),102 个阈值中无一内点同时优于两个端点;换用基于采样的语义熵后层内判别力升至0.886/0.750,但端到端仍不超越"总是检索"。以信息论形式化:门控优于"总是检索"当且仅当自信作答精度高于检索精度(定理 1),而自陈信道对"是否知道"的互信息仅 0.000–0.069 bit,瓶颈在读出环节而非带宽(定理 2)。然而该失效是域依赖的:在开放域事实问答(TriviaQA)上,同一自陈置信度确有判别力(AUC 0.64–0.69),且门控在检索不完美的开放域稳定优于 RAG(跨 14B 至满血模型七组对照全胜)。据此本文给出领域无关的门控"诊断—部署"框架,并揭示两项独立的方法学缺陷——弃答置信度的指代歧义与置信度派生指标的跨运行不可复现。本文建议一切置信度派生指标均应报告 k/n。
Abstract
Metacognitive-gated dynamic retrieval is widely used for hallucination mitigation in large language models, yet it rests on three untested assumptions: self-reported confidence correlates with correctness (H1), a tuned threshold enables decision-making (H2), and layering gating on grounding is additive (H3). On four small-to-mid open-weight instruction models, with full retrieval as the baseline, we run exhaustive threshold replay over synthetic and decontaminated real SEC corpora, together with a same-trigger-rate random control and an oracle upper bound. Self-reported confidence yields AUCs of only 0.3410.500 for predicting correctness, and no interior threshold dominates both endpoints; switching to semantic entropy raises within-stratum AUC to 0.886/0.750, yet end-to-end performance never beats always-retrieve. Formally, a gate beats always-retrieve iff the precision of its confident answers exceeds retrieval accuracy (Theorem 1), while the self-report channel carries only 0.0000.069 bits of knowledge informationa readout bottleneck (Theorem 2). This failure is domain-dependent: on open-domain factoid QA the same confidence is discriminative (AUC 0.640.69), and gating reliably beats RAG under imperfect retrieval. We distill a domain-agnostic diagnosis-and-deployment framework, and report two methodological defects: referent ambiguity of abstention confidence and cross-run irreproducibility. All confidence-derived metrics should be reported with k/n.关键词
大语言模型/幻觉治理/元认知门控/自适应检索/不确定性量化/弃答/可复现性Key words
large language models/hallucination mitigation/metacognitive gating/adaptive retrieval/uncertainty quantification/abstention/reproducibility引用本文复制引用
刘海熠.元认知门控式动态检索的领域相关收益边界与可迁移诊断框架[EB/OL].(2026-09-25)[2026-09-29].https://chinaxiv.org/abs/202609.00353.学科分类
计算技术、计算机技术