Get Started
Topic · 0 recaps
Statistical Machine Learning
Theoretical and statistical foundations of learning — generalization, uncertainty, kernels, Bayesian methods, and the math underneath modern empirical work.
...
Posts
Questions
Home
Topics
Search
Library
Questions researchers are working on
Follow a question through Rcap’s explanations and the latest papers addressing it.
Search
Can autoencoder parameter spectra serve as usable representations of their training data?
The statistical properties of training data may be reflected in the singular values of an autoencoder’s parameter matrices. The difficulty is determining whether this information remains sufficiently distinctive and usable as a data representation.
Can causal fairness constraints transfer across synthetic-data generators and privacy levels without sacrificing fidelity?
Synthetic data releases must suppress unfair causal pathways while retaining enough statistical structure for downstream use. It is unclear whether these controls remain effective when the generator family or formal privacy guarantee changes.
Can Gaussian-width restricted-eigenvalue guarantees survive heavy-tailed measurements under only a uniform small-ball condition?
Restricted eigenvalue bounds support stable recovery, but heavy-tailed measurements can make empirical control depend on simultaneous threshold occupancy rather than geometric width alone. This creates a gap between Gaussian-design behavior and what uniform small-ball assumptions can guarantee.
Can one squared-loss estimator achieve both minimax and universal exponential rates for finite versus countably infinite hypothesis classes?
Model-selection aggregation seeks minimax excess-risk guarantees, while universal learning seeks exponential rates. Whether one estimator can provide both depends on whether the hypothesis class is finite or countably infinite.
Can phase-transition counts during fine-tuning predict final test accuracy across architectures and distribution shifts?
During fine-tuning, class separability may change in discrete jumps, but the relationship between the number of jumps and eventual test accuracy may depend on architecture and whether evaluation data are i.i.d. or corrupted.
Can pre-generalization interventions reveal when training constrains which equally fitting neural-network solutions will later generalize?
Overparameterized networks can fit the same training data while differing substantially on unseen examples. During grokking, generalization emerges after a plateau, making it difficult to determine when training has begun constraining the eventual solution.
Can reducing the complexity of a linear generative prior improve expected reconstruction error in noiseless Gaussian compressed sensing?
Compressed sensing can use a family of linear generative priors with different effective complexities, but restricting that prior may affect reconstruction accuracy in ways that differ from ordinary denoising. The key issue is whether lower-complexity priors reduce expected error in the noiseless setting.
Can stationary-point ELBO values be expressed as entropy sums across variational generative models?
The ELBO is optimized in unsupervised latent-variable learning, yet its value at convergence can be difficult to interpret across different generative models. A common entropy-based form could clarify what stationary-point values represent.
Can stochastic weight averaging improve equivariance in augmented classification without repeated ensemble training?
Data augmentation incorporates task symmetries into neural networks, while deep ensembles can require many separate training runs. The practical difficulty is determining whether weight averaging can provide stronger symmetry handling without that repeated cost.
Can tabular foundation models learn transferable physical laws with units and noiseless mechanisms, not just interpolate data?
High predictive accuracy on equation-generated tables may reflect interpolation rather than representation of governing physics. Physical modeling also requires handling units and noiseless mechanisms, which table completion may not capture.
For non-separable logistic regression, when does constant-step gradient descent converge globally rather than enter a stable cycle?
A step size can keep the solution locally stable while trajectories from other initializations fail to reach it. Instead, the iterates may settle into a stable periodic cycle.
How can active preference learning obtain scalable, calibrated uncertainty for neural reward models without full Bayesian inference?
Active preference learning must choose which comparisons to request, but reliable uncertainty estimates become expensive for neural reward models when inference considers all parameters. Poorly calibrated uncertainty can lead to less informative queries and inefficient reward learning.
How can Adam’s error be bounded for strongly convex stochastic optimization without assuming bounded iterates?
Analyses of Adam on strongly convex stochastic problems have often assumed that its iterates remain uniformly bounded. Without that premise, it is unclear whether the optimizer’s error can be controlled unconditionally.
How can adversarial online maximization of non-monotone DR-submodular functions achieve the best offline approximation factor?
Adversarially changing objectives make it difficult to retain the approximation quality available when optimizing a fixed function offline. The difficulty is sharper for non-monotone objectives, where additional decisions can reduce value and feedback may be available only through an oracle.
How can agnostic learning handle low-intrinsic-dimensional concepts under arbitrary distributions with Gaussian-robust benchmarks?
Worst-case agnostic learning can be computationally hard even for simple concept classes, while existing tractable results often rely on highly structured instance distributions. Restricting the comparator to classifiers robust to small Gaussian perturbations changes the benchmark without requiring the data distribution itself to be structured.
How can archetypal profiles be identified from three-way asymmetric dissimilarities?
Standard multidimensional scaling methods generally assume symmetric relationships, making it difficult to represent directional, non-reflexive relationships across multiple occasions. This limits the extraction of meaningful archetypal profiles from three-way asymmetric data.
How can associative-memory capacity be compared across Hopfield and attention-like models without conflating assumptions?
Hopfield-style memories combine recurrent dynamics, energy landscapes, and pattern storage, but retrieval capacity depends on the disorder ensemble, scaling limit, and success criterion. Connections to attention and biological interpretation introduce further assumptions that can make superficially similar results incomparable.
How can attribute hypergraphs refine frozen graph-clustering assignments without labels or altering the model or original graph?
After training, correcting graph-clustering assignments without retraining leaves few available signals. Higher-order relations in node attributes may help, but indiscriminate updates can introduce erroneous changes.
How can autonomous vehicles anticipate target-gap choice and model longitudinal preparation before mandatory lane changes?
A vehicle may need to choose a target gap and reposition longitudinally before lateral movement begins, even when the eventual gap is not yet geometrically obvious. Treating lane changing as only a lateral maneuver can miss this preparatory phase.
How can B2B marketing teams resolve fragmented company records to predict conversion over long sales cycles?
B2B buying processes can span months or years, while one company may appear across fragmented contact records and inconsistent company names. This makes it difficult to construct a reliable customer-level view and determine which prospects are likely to convert.
Previous
1 / 14
Next