Get Started
Research questionHow can text-guided 3D scene synthesis be evaluated against fine-grained spatial constraints?Text-guided scene evaluations often overlook details such as how an object should be placed or how objects should relate spatially. This makes it difficult to detect partial or subtle misalignment between a description and its synthesized scene.
Computer Vision
Evaluation & Benchmarks
Latest papersRecent research connected to this question, newest first.Apples on the Table? Evaluating Text-Guided 3D Scene Synthesis via Fine-Grained Constraint VerificationThe source evaluates this problem with descriptions paired with human-annotated constraints and reference scenes. Its framework decomposes descriptions into atomic constraints, grounds textual references to 3D objects, and reasons about their spatial relationships; reported results indicate that current synthesis approaches satisfy at most 10% of the evaluated constraints.research paper · Sep 1, 2026
Related questions
How reliably can open-set vision-language 3D scene graphs support outdoor navigation and object retrieval?How can satellite imagery support large, explorable 3D urban scenes when high-quality scans are scarce?How can we complete unseen 3D scene regions from sparse, unconstrained views without dense 3D supervision?How can reference-guided image generation control distinct attributes of multiple objects in complex scenes?
Home
Topics
Search
Library