Research questionHow can multimodal models integrate evidence across deeply interleaved text and images?Many multimodal evaluations use images with shallow textual instructions, so they may not test reasoning when text and visual clues depend on one another. Real tasks can require recovering facts from evidence distributed throughout a deeply interleaved context.