Get Started
Home
Topics
Search
Library
Research questionHow can automated vision-language systems reliably describe defects and answer questions about NDE images?NDE images can contain subtle defect features that inspectors must translate into descriptions and targeted answers. Automated systems may generate fluent captions while omitting or misrepresenting diagnostically important details.
AI
Computer Vision
Evaluation & Benchmarks
Machine Learning
Multimodal Models
Natural Language Processing
Research Paper
Latest papersRecent research connected to this question, newest first.An Integrated Vision-and-Language Pretraining (VLP) and Visual Question Answering (VQA) model to Automate Nondestructive Evaluation Image AnalysisThe source describes a vision-language system trained on annotated NDE images, using ResNet50 for image features and GPT2 for caption generation, with a VQA component for image-based questions. Its evidence includes BLEU comparisons against expert descriptions; reported accuracy is low, although captions mention some important features and are shorter. Field performance and actual reductions in inspection errors are not established.research paper · Sep 4, 2026
Related questions
How can industrial vision systems grade defect severity from imperfect detections while respecting ordinal levels and morphology?How can vision-language-action robots be evaluated for execution quality and decision confidence beyond binary task success?How can training-free visual anomaly detection catch local defects and global rule violations without category-specific modeling?How should vision-language models answer valid parts of compound queries while withholding unsafe or unanswerable parts?