Fus3D: Decoding Consolidated 3D Geometry from Feed-Forward Geometry Transformer Latents

Fink L, Franke L, Kopanas G, Stamminger M, Hedman P (2026)


Publication Type: Conference contribution

Publication year: 2026

Journal

Original Authors: Laura Fink, Linus Franke, George Kopanas, Marc Stamminger, Peter Hedman

Pages Range: 578-597

Event location: Malmö SE

ISBN: 9783032371669

DOI: 10.1007/978-3-032-37167-6_33

Abstract

We propose a feed-forward method for dense Signed Distance Field (SDF) regression from unstructured image collections in less than three seconds, without camera calibration or post-hoc fusion. Our key insight is that the intermediate feature space of pretrained multi-view feed-forward geometry transformers (FFGT) already encodes a powerful joint world representation; yet, existing pipelines discard it, routing features through per-view prediction heads before assembling 3D geometry post-hoc, which discards valuable completeness information and accumulates inaccuracies. We instead perform 3D extraction directly from FFGT features via learned volumetric extraction: voxelized canonical embeddings that progressively absorb multi-view geometry information through interleaved cross- and self-attention into a structured volumetric latent grid. A simple convolutional decoder then maps this grid to a dense SDF. We additionally propose a scalable, validity-aware supervision scheme directly using SDFs derived from depth maps or 3D assets, tackling practical issues like non-watertight meshes. Our approach yields complete and well-defined distance values across sparse- and dense-view settings and demonstrates geometrically plausible completions.

Authors with CRIS profile

How to cite

APA:

Fink, L., Franke, L., Kopanas, G., Stamminger, M., & Hedman, P. (2026). Fus3D: Decoding Consolidated 3D Geometry from Feed-Forward Geometry Transformer Latents. In Proceedings of the European Conference on Computer Vision (ECCV 2026) (pp. 578-597). Malmö, SE.

MLA:

Fink, Laura, et al. "Fus3D: Decoding Consolidated 3D Geometry from Feed-Forward Geometry Transformer Latents." Proceedings of the European Conference on Computer Vision (ECCV 2026), Malmö 2026. 578-597.

BibTeX: Download