Research questionCan post-training ternarization make language models smaller without unacceptable capability loss or slower inference?Ultra-low-bit weights can shrink model storage, but nominal bit counts may not reflect the stored representation, uneven task degradation, or actual inference speed. Compression may therefore improve footprint without improving end-to-end deployment performance.