Research questionHow can single-pass image captioning capture fine-grained visual details without multi-stage latency?Image captioning systems often produce fluent descriptions while omitting attributes, counts, textures, materials, and spatial relations. Recovering these details through multi-stage generation and verification can substantially increase inference latency.