首页|MMD-Flagger: Leveraging Maximum Mean Discrepancy to Detect Hallucinations

MMD-Flagger: Leveraging Maximum Mean Discrepancy to Detect Hallucinations

来源：

英文摘要

Large language models (LLMs) have become pervasive in our everyday life. Yet, a fundamental obstacle prevents their use in many critical applications: their propensity to generate fluent, human-quality content that is not grounded in reality. The detection of such hallucinations is thus of the highest importance. In this work, we propose a new method to flag hallucinated content, MMD-Flagger. It relies on Maximum Mean Discrepancy (MMD), a non-parametric distance between distributions. On a high-level perspective, MMD-Flagger tracks the MMD between the generated documents and documents generated with various temperature parameters. We show empirically that inspecting the shape of this trajectory is sufficient to detect most hallucinations. This novel method is benchmarked on two machine translation datasets, on which it outperforms natural competitors.

作者：Kensuke Mitsuzawa、Damien Garreau

作者单位：

学科分类：计算技术、计算机技术

推荐引用：Kensuke Mitsuzawa,Damien Garreau.MMD-Flagger: Leveraging Maximum Mean Discrepancy to Detect Hallucinations[EB/OL].(2025-06-02)[2025-07-02].https://arxiv.org/abs/2506.01367.点此复制

MMD-Flagger: Leveraging Maximum Mean Discrepancy to Detect Hallucinations

MMD-Flagger: Leveraging Maximum Mean Discrepancy to Detect Hallucinations

评论