Research questionHow can we trace which visual, question, or prior-token signals drive VLM generation at each decoding step?Final-answer metrics cannot show how visual input, question text, and already generated tokens shape a response as it unfolds. This obscures whether generation remains visually grounded or increasingly follows the question or its own prior output.