Get Started
Home
Topics
Search
Library
Research questionHow can vision-language models faithfully transcribe corrupted document text instead of rewriting it?Imperfect document text can prompt models to replace unusual words with more plausible ones, producing fluent but incorrect transcriptions. Clean-text OCR benchmarks may miss this silent rewriting.
AI
Computer Vision
Evaluation & Benchmarks
Multimodal Models
Natural Language Processing
Latest papersRecent research connected to this question, newest first.Do VLMs Read or Rewrite? On Transcription Faithfulness in Vision-Language ModelsEvidence comes from FaithC4, a multilingual benchmark of 1,455 single-page English, Chinese, and Korean documents with scrambled, randomly substituted, and visually similar characters. It evaluates 15 general-purpose and OCR-specialized VLMs alongside traditional OCR pipelines, with additional layer analysis of Qwen3-VL-4B and word-length effects.research paper · Sep 2, 2026
Related questions
How can OCR diagnose and correct case-level errors while preserving rendering-equivalent outputs?How can zero-shot vision-language models adapt online to image corruption using only unlabeled test images?How can language models improve accessibility-focused text simplification for low-resource languages when cross-lingual transfer is unreliable?How can retrieval-augmented language models resist ordinary-looking GEO-optimized documents that distort synthesized answers?