Research questionHow can deep-network momentum adapt forgetting to unevenly sampled input directions?A single exponential decay treats frequently and rarely activated directions alike. As a result, stale gradient information may persist in some directions while other directions are forgotten too quickly.