Get Started
Home
Topics
Search
Library
Research questionHow can document-image PII redaction prevent page-level leakage despite OCR errors and visual noise?OCR errors, layout structure, and visual noise can hide identifiers in document images. A page remains unsafe for release when even one PII item is missed.
AI
Alignment & Safety
Computer Vision
Evaluation & Benchmarks
Image & Video Processing
Machine Learning
Multimodal Models
Natural Language Processing
Technology
Latest papersRecent research connected to this question, newest first.LeakageBench: Document-Level Leakage Risk for Redacting Personally Identifiable Information in Document ImagesLeakageBench provides evidence from 500 document images containing 11,954 GDPR-aligned annotations covering direct identifiers, linkage keys, and contextual re-identification surfaces. It evaluates OCR-dependent detectors and OCR-free vision-language models with entity-level F1, group-wise leakage, and document-level leakage; reported tool assistance improved localization F1 from 0.090 to 0.249 while critical page-level leakage remained 0.968.research paper · Sep 2, 2026
Related questions
How can OCR diagnose and correct case-level errors while preserving rendering-equivalent outputs?How can PII detectors avoid missed entities when deployment data shifts from benchmark conditions?How can text-guided image editing explain edits while keeping provenance resistant to white-box removal?How can image editors infer edit regions and preserve non-target content without introducing artifacts?