Research questionHow do architecture and optimization shape accessible representations and scaling at finite budgets on the same data?Models trained on identical data can follow different loss-versus-budget curves because architecture and optimization may make different task-relevant representations accessible. Existing scaling descriptions do not fully explain how task geometry, architectural support, and finite-budget acquisition combine to determine those curves.