首页|GD-Retriever: Controllable Generative Text-Music Retrieval with Diffusion Models

GD-Retriever: Controllable Generative Text-Music Retrieval with Diffusion Models

来源：

英文摘要

Multimodal contrastive models have achieved strong performance in text-audio retrieval and zero-shot settings, but improving joint embedding spaces remains an active research area. Less attention has been given to making these systems controllable and interactive for users. In text-music retrieval, the ambiguity of freeform language creates a many-to-many mapping, often resulting in inflexible or unsatisfying results. We introduce Generative Diffusion Retriever (GDR), a novel framework that leverages diffusion models to generate queries in a retrieval-optimized latent space. This enables controllability through generative tools such as negative prompting and denoising diffusion implicit models (DDIM) inversion, opening a new direction in retrieval control. GDR improves retrieval performance over contrastive teacher models and supports retrieval in audio-only latent spaces using non-jointly trained encoders. Finally, we demonstrate that GDR enables effective post-hoc manipulation of retrieval behavior, enhancing interactive control for text-music retrieval tasks.

作者：Elio Quinton、GyÃ¶rgy Fazekas、Julien Guinot

作者单位：

学科分类：计算技术、计算机技术

推荐引用：Elio Quinton,GyÃ¶rgy Fazekas,Julien Guinot.GD-Retriever: Controllable Generative Text-Music Retrieval with Diffusion Models[EB/OL].(2025-06-24)[2025-07-16].https://arxiv.org/abs/2506.17886.点此复制

GD-Retriever: Controllable Generative Text-Music Retrieval with Diffusion Models

GD-Retriever: Controllable Generative Text-Music Retrieval with Diffusion Models

评论