NeurIPS 2026

RelationVGGT

Visual Geometry Transformers for 3D Spatial Relation Segmentation

Minsu Kim1 Jaesung Choe2 Jiwoo Lee1 Yu-Chiang Frank Wang2 Seon Joo Kim1

1Yonsei University2NVIDIA

TL;DR

Given a subject mask in one view and a relational query such as “resting on”, RelationVGGT segments the target object in every other view without being told its category, in a single forward pass on unposed images.

Task overview. A relational query (What is MASK lying on?) and a subject image with its mask condition RelationVGGT; unposed multi-view target images are the query. The output is the target, the sofa, masked in red in every target view.

Reference view with the subject highlighted in green

Reference view. The subject is given only as a mask in this one view.

Predicted target masks in four target views for each method

Target views. Each row is a different target view; red marks the predicted target. Baselines need posed images and per-scene optimization, RelationVGGT runs feed-forward on unposed images.

Abstract

Recent advances in 3D reconstruction have progressed from per-scene optimization to feed-forward inference, and semantic scene understanding has followed suit—yet existing methods remain confined to object-centric perception, neglecting spatial relations between objects. We formulate 3D spatial relation segmentation in a feed-forward, pose-free multi-view setting: given a visually specified subject and a relational text query, the model segments the target across views without receiving its category name. To this end, we propose RelationVGGT, a novel feed-forward framework that integrates semantic features from a visual foundation model with geometry-aware representations from a 3D geometry foundation model and leverages a relation transformer for subject-conditioned, cross-view relation prediction—requiring neither per-scene optimization nor known camera poses. We additionally provide a fully automated annotation pipeline built on ScanNet++ with VLMs and LLMs, enabling scalable training data generation for this new task.

The task

A scene is captured by a set of images. One of them is the reference view, where a mask marks the subject: the object the relation starts from. A relational text query names the spatial relation, and the target is the object that satisfies it with respect to the subject. For every other view, the model must segment that target. If the subject is a monitor and the query is “resting on”, the answer is the desk under it, in every view where the desk appears.

Given

  • Unposed multi-view images of a scene
  • A subject mask in a single reference view
  • A relation phrase, e.g. “resting on”

Predict

  • The target object’s mask in every other view

Never given

  • The target’s category name
  • Camera poses or per-scene optimization

Key insight

3D relational reasoning does not require volumetric search: 3D-aware features and pixel-wise geometric grounding let us localize relational targets directly in image space.

What changed? From volumetric search to pixel-grid prediction

The closest prior work, RelationField, models relations inside a radiance field. We keep the idea of a relation feature but move where it is predicted.

Prior work: volumetric search

z x evaluate gθ at every candidate x scene volume (per-scene, posed)

gθ(x, z) ↦ r

x
3D position of a candidate target point
z
3D position of the query subject
r
relation feature between x and z, compared with the text embedding of the query

Finding the target means evaluating gθ at every candidate x in the scene volume. That requires an explicit reconstruction optimized per scene from posed images, and scene-specific positional encodings tie what is learned to one coordinate system.

RelationVGGT: pixel-grid prediction

Ir,Ms reference view with subject mask fθ u It rt(u) for every pixel, every target view, one pass

fθ(Ir, Ms, It, u) → rt(u)

Ir
reference image, the one view where the subject is marked
Ms
binary mask of the subject in Ir
It
a target view: any of the other images
u
a pixel location in It
rt(u)
relation feature between pixel u and the subject, aligned with the language embedding space

With the subject fixed by Ir and Ms, the search moves onto each target view’s image grid. Comparing rt(u) with the query embedding e(q) gives a dense relevance map directly: no volume to sweep, no per-scene optimization, and the model transfers to unseen scenes.

The task also differs from 3D referring segmentation, where a text expression can name the target outright. Here the subject is specified visually and only the relation is given in words, so the target’s identity must be inferred from the subject–relation pair.

Method

RelationVGGT turns the reference subject and a set of unposed views into language-aligned relation features for every target-view pixel, in a single forward pass.

RelationVGGT architecture overview
Overview of RelationVGGT. Multi-view images are encoded by frozen 2D and 3D visual foundation models and fused by an input mixer. The subject mask conditions the relation transformer, whose output is decoded into per-view relation feature maps. The target is segmented by feature–text similarity with the encoded relational query.

Semantics meet geometry

Dense DINOv2 features supply object-level semantics; Pi3 encoder tokens supply multi-view geometry. An input mixer fuses both into joint tokens for every view.

Relation transformer

Target tokens alternate cross-attention to the subject-masked reference tokens with self-attention across all target views, making them subject-conditioned and multi-view aware.

Open-vocabulary masks

A DPT head upsamples tokens to pixel-aligned relation features. A sigmoid over feature–text similarity gives the mask; Pi3’s point decoder lifts it to 3D.

One spatial arrangement can be described many ways (“standing on”, “on”, “supported by”), so each target mask is supervised with several compatible phrasings at once, using a multi-query sigmoid focal loss and a Dice loss.

Automatic relation annotation

Multi-view consistent relation labels are scarce and costly to annotate by hand. We build them automatically on ScanNet++: a VLM proposes relations per frame with Set-of-Mark prompting, proposals are kept only if at least three viewpoints agree, geometric checks on 3D boxes remove implausible edges, and an LLM rephrases each triplet to diversify the queries.

Relation annotation pipeline
Relation annotation pipeline. (a) Instance masks are projected onto frames with SoM prompting to query a VLM for pairwise relations. (b) Proposals are aggregated into a scene graph and refined by geometry-aware post-processing and LLM rephrasing. (c) Multi-view consistent subject (green) and target (red) annotations.

Results

We evaluate on manually annotated relation benchmarks built on Replica, LERF, and ScanNet++: 119 queries covering 28 unique predicates. mAcc is the fraction of queries with IoU above 0.25.

MethodSettingReplicaLERFScanNet++Average
mIoUmAccmIoUmAccmIoUmAccmIoUmAcc
LangSplatPer-scene, posed0.1910.2560.1590.2540.0740.0640.1410.191
OpenGaussianPer-scene, posed0.3490.4700.2560.3780.1620.2540.2550.367
RelationFieldPer-scene, posed0.3260.5370.2760.4100.3070.4910.3030.480
RelationField†Per-scene, posed0.2190.3360.2320.3460.2400.3950.2300.359
RelationVGGT (Ours)Feed-forward, pose-free0.4460.6400.2540.3320.4450.6750.3820.549

Best Second Third. Average is the macro average over the three benchmarks. RelationField† removes target-object retrieval and uses the relation embedding alone. RelationField also receives the ground-truth target category at test time; RelationVGGT does not.

Wide camera baselines

We also compare with video foundation models, feeding the multi-view frames as a video with the subject region and relation, but not the target category. The 119 queries are sorted by camera baseline into three near-equal groups. Video models degrade as viewpoints spread apart, while the 3D geometry backbone keeps RelationVGGT stable.

MethodSmall baseline (40)Medium baseline (39)Large baseline (40)
mIoUmAccmIoUmAccmIoUmAcc
VideoLISA0.1370.1710.1050.1430.0380.064
Sa2VA0.2670.3210.2250.2890.1920.225
UniPixel0.4230.4960.3280.4510.3540.475
RelationVGGT (Ours)0.4170.5500.4040.6400.3660.575
Comparison with UniPixel across views
RelationVGGT vs. UniPixel. Our masks stay consistent across views, whereas UniPixel loses the target as the viewpoint changes.

More analysis

Every component pulls its weight: removing any one of them lowers ScanNet++ mIoU by 0.023 to 0.075.

  1. Input features. Replacing Pi3 tokens with positionally encoded, scale-standardized pointmaps drops mIoU from 0.445 to 0.370: raw coordinates carry geometry but none of the semantics already embedded in learned features. Removing DINOv2 (Pi3 only) costs 0.023 mIoU as object-level semantics are lost. Removing Pi3 (DINOv2 only) costs 0.075, since DINOv2 alone has no multi-view awareness and cannot capture cross-view correspondence.
  2. Relation transformer. Our block alternates subject-conditioned cross-attention with global self-attention across target views. A cross-attention-only variant (0.399) and VGGT’s alternating attention, which interleaves global and per-frame self-attention (0.394), both fall short.
  3. Training objective. Supervising each mask with a single predicate instead of several compatible ones drops mIoU to 0.406. Spatial relations are many-to-one, and multi-query supervision avoids penalizing semantically plausible alternatives.
  4. Annotation pipeline. Training on labels without geometry-based edge filtering drops mIoU to 0.398, because the model then learns from physically implausible relations, such as contact between distant objects.
DesignmIoUmAcc
RelationVGGT0.4450.675
1. Input features
Pointmap PE + DINOv20.3700.568
Pi3 only0.4220.579
DINOv2 only0.3700.579
2. Relation transformer
Cross-attention only0.3990.579
Alternating attention (VGGT)0.3940.571
3. Training objective
Single-label focal loss0.4060.604
4. Annotation pipeline
w/o geometric edge filtering0.3980.625

ScanNet++ benchmark, 40 queries, threshold τ = 0.3.

BibTeX

@inproceedings{kim2026relationvggt,
  title     = {RelationVGGT: Visual Geometry Transformers for
               3D Spatial Relation Segmentation},
  author    = {Kim, Minsu and Choe, Jaesung and Lee, Jiwoo and
               Wang, Yu-Chiang Frank and Kim, Seon Joo},
  booktitle = {Advances in Neural Information Processing Systems},
  year      = {2026}
}