|国家预印本平台
首页|MMD-Flagger: Leveraging Maximum Mean Discrepancy to Detect Hallucinations

MMD-Flagger: Leveraging Maximum Mean Discrepancy to Detect Hallucinations

MMD-Flagger: Leveraging Maximum Mean Discrepancy to Detect Hallucinations

来源:Arxiv_logoArxiv
英文摘要

Large language models (LLMs) have become pervasive in our everyday life. Yet, a fundamental obstacle prevents their use in many critical applications: their propensity to generate fluent, human-quality content that is not grounded in reality. The detection of such hallucinations is thus of the highest importance. In this work, we propose a new method to flag hallucinated content, MMD-Flagger. It relies on Maximum Mean Discrepancy (MMD), a non-parametric distance between distributions. On a high-level perspective, MMD-Flagger tracks the MMD between the generated documents and documents generated with various temperature parameters. We show empirically that inspecting the shape of this trajectory is sufficient to detect most hallucinations. This novel method is benchmarked on two machine translation datasets, on which it outperforms natural competitors.

Kensuke Mitsuzawa、Damien Garreau

计算技术、计算机技术

Kensuke Mitsuzawa,Damien Garreau.MMD-Flagger: Leveraging Maximum Mean Discrepancy to Detect Hallucinations[EB/OL].(2025-06-02)[2025-07-02].https://arxiv.org/abs/2506.01367.点此复制

评论