Get Started
Research questionHow can reference-guided image generation control distinct attributes of multiple objects in complex scenes?A single reference image can contain several objects whose attributes need to be used independently. Global conditioning and a single identifier token can blend these cues, making it difficult to control each object's contribution in a new image.
Computer Vision
Diffusion Models
Image Generation
Latest papersRecent research connected to this question, newest first.RefDiT: Local Attribute Guidance in Reference-Based Image GenerationThe source describes RefDiT, a DiT-based diffusion framework that takes a reference image, text prompt, and optional user-provided guidance context. It decomposes identifier tokens at the attribute level and learns correspondence between tokens and local reference regions using LoRA blocks and inference-prompt context adjustment. The supplied abstract does not report quantitative results or specify annotation requirements.research paper · Sep 4, 2026
Related questions
How can composed image retrieval distinguish changed, preserved, and removed visual details?How can text-to-video generation continuously edit local attributes while preserving unrelated content?How can generative image systems preserve subject identity under viewpoint changes, degradation, and iterative edits?How can single-image diffusion personalization preserve identity while enabling controllable attribute-level context changes?
Home
Topics
Search
Library