Get Started
Topic · 32 recaps
Machine Learning
General machine learning — algorithms, architectures, optimization, theory. The broad arXiv category that papers without a more specific home tend to land in.
Play all
...
Posts
Questions
Home
Topics
Search
Library
Questions researchers are working on
Follow a question through Rcap’s explanations and the latest papers addressing it.
Search
Can autoencoder parameter spectra serve as usable representations of their training data?
The statistical properties of training data may be reflected in the singular values of an autoencoder’s parameter matrices. The difficulty is determining whether this information remains sufficiently distinctive and usable as a data representation.
Can automated alignment research mitigate multiple measurable safety failures without sacrificing general model capability?
Alignment failures such as deception, sycophancy, and jailbreaks can be measured, but reducing several simultaneously may interfere with a model’s broader capabilities. It is also unclear whether automated researchers can develop effective interventions without extensive human guidance.
Can causal fairness constraints transfer across synthetic-data generators and privacy levels without sacrificing fidelity?
Synthetic data releases must suppress unfair causal pathways while retaining enough statistical structure for downstream use. It is unclear whether these controls remain effective when the generator family or formal privacy guarantee changes.
Can compact pretrained brain MRI models transfer across Alzheimer’s tasks and cohorts without task-specific retraining?
Limited labeled neuroimaging data makes task-specific deep learning difficult. It remains uncertain whether features learned for one brain MRI task generalize to different Alzheimer’s-related tasks and cohorts.
Can language-based models replace specialized architectures for structured data without sacrificing structural representation and computation?
Task-level accuracy can look competitive even when a model does not preserve or compute the structure that makes structured-data problems tractable. This makes architectural replacement difficult to judge from predictive performance alone.
Can multimodal chest-radiograph triage trained on NLP-derived labels reliably match expert severity judgments?
Chest-radiograph triage must distinguish urgent examinations from routine ones, but labels extracted from reports may not capture radiologists’ severity judgments. Strong benchmark performance can also coexist with visual explanations that do not localize clinically relevant findings.
Can non-smooth or quantized activations support stable echo-state dynamics beyond conventional spectral-radius expectations?
Echo state network stability is often analyzed using smooth activations and conservative spectral-radius conditions. Irregular or quantized activations may change how reservoir states contract, remain distinct, or converge, but their stability behavior is not fully understood.
Can one squared-loss estimator achieve both minimax and universal exponential rates for finite versus countably infinite hypothesis classes?
Model-selection aggregation seeks minimax excess-risk guarantees, while universal learning seeks exponential rates. Whether one estimator can provide both depends on whether the hypothesis class is finite or countably infinite.
Can phase-transition counts during fine-tuning predict final test accuracy across architectures and distribution shifts?
During fine-tuning, class separability may change in discrete jumps, but the relationship between the number of jumps and eventual test accuracy may depend on architecture and whether evaluation data are i.i.d. or corrupted.
Can post-training ternarization make language models smaller without unacceptable capability loss or slower inference?
Ultra-low-bit weights can shrink model storage, but nominal bit counts may not reflect the stored representation, uneven task degradation, or actual inference speed. Compression may therefore improve footprint without improving end-to-end deployment performance.
Can pre-generalization interventions reveal when training constrains which equally fitting neural-network solutions will later generalize?
Overparameterized networks can fit the same training data while differing substantially on unseen examples. During grokking, generalization emerges after a plateau, making it difficult to determine when training has begun constraining the eventual solution.
Can preprocessing defenses detect adversarial attacks in depthwise-separable edge vision CNNs when they cannot restore predictions?
Preprocessing defenses are often assumed to transfer across model architectures, but depthwise-separable CNNs may respond differently to adversarial perturbations than residual or Inception-style networks. Their failure to recover predictions may still produce measurable differences between clean and adversarial inputs, while image-quality scores may not reflect defensive value.
Can prompt phrasing reliably improve LLM-derived chemical features for drug-toxicity prediction?
Minor changes in prompt phrasing can alter LLM outputs, making it unclear whether prompt optimization produces stable chemical features for toxicity models. This variability complicates the use of LLM-generated features in a costly drug-development process.
Can quantitative attribution metrics show where facial-video rPPG models read pulse signals without indicating heart-rate accuracy?
Facial-video remote photoplethysmography models estimate pulse from short clips, while attribution maps are often interpreted as evidence of what the model uses. Localization of skin regions may not correspond to accurate heart-rate estimation.
Can reducing the complexity of a linear generative prior improve expected reconstruction error in noiseless Gaussian compressed sensing?
Compressed sensing can use a family of linear generative priors with different effective complexities, but restricting that prior may affect reconstruction accuracy in ways that differ from ordinary denoising. The key issue is whether lower-complexity priors reduce expected error in the noiseless setting.
Can scaling vision-language models overcome their limitations in neurosurgical tool detection?
Neurosurgical tool detection requires specialized data and expert labeling, while larger models and longer training demand substantial computational resources. It remains unclear whether adding these resources produces meaningful gains or leaves important limitations unchanged.
Can stationary-point ELBO values be expressed as entropy sums across variational generative models?
The ELBO is optimized in unsupervised latent-variable learning, yet its value at convergence can be difficult to interpret across different generative models. A common entropy-based form could clarify what stationary-point values represent.
Can stochastic weight averaging improve equivariance in augmented classification without repeated ensemble training?
Data augmentation incorporates task symmetries into neural networks, while deep ensembles can require many separate training runs. The practical difficulty is determining whether weight averaging can provide stronger symmetry handling without that repeated cost.
Can tabular foundation models learn transferable physical laws with units and noiseless mechanisms, not just interpolate data?
High predictive accuracy on equation-generated tables may reflect interpolation rather than representation of governing physics. Physical modeling also requires handling units and noiseless mechanisms, which table completion may not capture.
Can text-to-image models match web-scale performance using smaller, reproducible datasets and models?
Billion-scale web-scraped datasets can make text-to-image results difficult to reproduce because their contents and availability change. The central difficulty is determining whether a much smaller, standardized image collection can retain the capabilities associated with those models.
Previous
1 / 50
Next