Get Started
Topic · 0 recaps

Image & Video Processing

Image and video signal processing — restoration, compression, super-resolution, and the lower-level vision techniques behind modern computer-vision pipelines.
PostsQuestions
Home
Topics
Search
Library
Questions researchers are working onFollow a question through Rcap’s explanations and the latest papers addressing it.
Can compact pretrained brain MRI models transfer across Alzheimer’s tasks and cohorts without task-specific retraining?Limited labeled neuroimaging data makes task-specific deep learning difficult. It remains uncertain whether features learned for one brain MRI task generalize to different Alzheimer’s-related tasks and cohorts.Can internal photogrammetric validation certify metric accuracy without external survey or control-point measurements?A reconstruction can be internally geometrically consistent while still containing coherent global distortion. This makes metric correctness difficult to establish when external survey or control-point measurements are unavailable.Can preprocessing defenses detect adversarial attacks in depthwise-separable edge vision CNNs when they cannot restore predictions?Preprocessing defenses are often assumed to transfer across model architectures, but depthwise-separable CNNs may respond differently to adversarial perturbations than residual or Inception-style networks. Their failure to recover predictions may still produce measurable differences between clean and adversarial inputs, while image-quality scores may not reflect defensive value.Can quantitative attribution metrics show where facial-video rPPG models read pulse signals without indicating heart-rate accuracy?Facial-video remote photoplethysmography models estimate pulse from short clips, while attribution maps are often interpreted as evidence of what the model uses. Localization of skin regions may not correspond to accurate heart-rate estimation.How can 3D brain MRI inpainting reconstruct healthy tissue in pathological regions without changing observed anatomy?Pathological or masked regions remove information needed for automated brain MRI analysis. A reconstruction must appear anatomically plausible without modifying the surrounding anatomy that remains visible.How can 3D Gaussian splatting reconstruct large aerial surfaces without cross-region seams or local geometric inconsistencies?Aerial scenes may be divided into independently optimized regions, breaking continuous surfaces and creating stitching artifacts. Constraints centered on individual Gaussians and fixed regularization can also fail when local geometry varies between structured and unstructured areas.How can 3D Gaussian Splatting training stay efficient as high-resolution scenes require more Gaussian primitives?During 3D Gaussian Splatting optimization, the number of Gaussian primitives can keep increasing, raising the cost of each training run. Higher-resolution scenes intensify this burden and can slow convergence.How can 3D occupancy models learn from noisy 2D pseudo-labels without 3D annotations?2D pseudo-labels may contain both depth errors and semantic mistakes. Projecting these imperfect targets into 3D can propagate inaccuracies through the predicted occupancy field.How can a transferable image-style representation separate overall identity from fine-grained visual attributes?Image style combines high-level identity with many visual factors that are entangled with image content. Without an explicit shared representation, styles are difficult to compare, transfer, and model consistently.How can adversarial perturbations transfer to unseen semantic-segmentation models while accounting for dense spatial and class-wise structure?A perturbation crafted on a surrogate segmentation model must mislead an unseen target model. This is difficult because dense prediction depends on spatially organized and class-specific representations, not only output scores.How can affective AI be evaluated using long-term, naturalistic, passively sensed workplace data?Most affective-computing systems are evaluated on short, controlled datasets, making it difficult to separate person-specific, team-level, and seasonal variation in real workplaces. Conventional classification metrics can also miss demographic bias and poorly represent outcomes such as employee turnover.How can agents find people across cameras from vague witness clues under spatial-temporal and turn constraints?Witness accounts may be partial or ambiguous, while relevant observations are distributed across camera locations and time. An agent must choose questions and searches before its interaction budget runs out.How can agricultural robotics reduce the manual pixel-level labeling needed for accurate plant and fruit segmentation?Agricultural robotics depends on pixel-level plant and fruit segmentation, but creating those masks requires costly, labor-intensive annotation. Reducing annotator input without substantially degrading training-label quality remains difficult.How can AI augment computational design while preserving domain grounding, verification, and scientific judgment?AI can expand the search space of computational design, but its contributions may be difficult to ground in domain requirements or verify and reproduce. Researchers therefore need ways to keep AI-assisted design decisions traceable and scientifically accountable.How can AI generate attractive graphics with accurate text and editable layers?Bitmap generation often flattens designs, making text unreliable and later edits difficult. Code-based generation preserves structure but can struggle with aesthetic judgment and complex visual assets.How can athletes be localized in world coordinates from a single ultra-high-resolution calibrated broadcast frame?A single broadcast frame must support athlete localization despite extreme differences in apparent scale. Perspective distortion also makes it difficult to convert image locations into accurate world coordinates.How can attackers evade black-box AIGC detectors using a frozen diffusion model without source images or detector-aware retraining?AIGC detectors must identify synthetic images even when generation can be adjusted to reduce detectable signals. This is difficult when the attacker can observe only detector responses and cannot alter the generator or rely on source images.How can automated aortic segmentation in 4D flow MRI remain time-resolved without dense annotations or excessive computation?Reproducible hemodynamic measurements require accurate segmentation throughout the cardiac cycle. However, dense 4D labels are scarce, while processing volumetric time-series data is computationally demanding.How can automated lunar-crater detection remain reliable across crater sizes, illumination, and rugged terrain?Lunar imagery contains craters with widely varying sizes and shapes, while illumination changes and rugged terrain can obscure their boundaries. Missed or mislocalized craters can complicate assessment of potential landing sites.How can automated palynomorph detection scale to whole-slide multifocal microscopy images?Manual inspection of multifocal whole-slide images is too slow for large-scale palynological studies. The images must be made tractable for automated analysis without losing detections across the full slide.
Previous
1 / 15
Next