首页|$S^2M^2$: Scalable Stereo Matching Model for Reliable Depth Estimation

$S^2M^2$: Scalable Stereo Matching Model for Reliable Depth Estimation

来源：

英文摘要

The pursuit of a generalizable stereo matching model, capable of performing well across varying resolutions and disparity ranges without dataset-specific fine-tuning, has revealed a fundamental trade-off. Iterative local search methods achieve high scores on constrained benchmarks, but their core mechanism inherently limits the global consistency required for true generalization. However, global matching architectures, while theoretically more robust, have historically been rendered infeasible by prohibitive computational and memory costs. We resolve this dilemma with $S^2M^2$: a global matching architecture that achieves state-of-the-art accuracy and high efficiency without relying on cost volume filtering or deep refinement stacks. Our design integrates a multi-resolution transformer for robust long-range correspondence, trained with a novel loss function that concentrates probability on feasible matches. This approach enables a more robust joint estimation of disparity, occlusion, and confidence. $S^2M^2$ establishes a new state of the art on Middlebury v3 and ETH3D benchmarks, significantly outperforming prior methods in most metrics while reconstructing high-quality details with competitive efficiency.

作者：Junhong Min、Youngpil Jeon、Jimin Kim、Minyong Choi

作者单位：

学科分类：计算技术、计算机技术

推荐引用：Junhong Min,Youngpil Jeon,Jimin Kim,Minyong Choi.$S^2M^2$: Scalable Stereo Matching Model for Reliable Depth Estimation[EB/OL].(2025-07-30)[2025-08-10].https://arxiv.org/abs/2507.13229.点此复制

$S^2M^2$: Scalable Stereo Matching Model for Reliable Depth Estimation

$S^2M^2$: Scalable Stereo Matching Model for Reliable Depth Estimation

评论