Research questionWhen can frozen vision-language models reliably reason about counterfactual scenes from object-token edits alone?Object-token edits change a scene representation without providing a new image, so it is unclear whether a VLM uses the altered objects or relies on learned priors. Evaluation is difficult when post-edit answers have not been annotated.