Research questionHow can pixel-text encoders robustly represent grounded multilingual text across resolutions under tight visual-token budgets?Fixed image resolutions can miss fine text, while visual shortcuts and weak grounding undermine generalization across languages and image contexts. Preserving useful text representations also becomes harder when visual tokens must be heavily compressed.