Research questionHow can we generate scalable, photorealistic labeled data for monocular 3D human mesh estimation?Monocular images provide ambiguous depth, making accurate 3D mesh annotation expensive and difficult to scale. Rendered synthetic data offers precise labels but often lacks the photorealism and diversity needed for effective training.