Research questionHow can we interpret multimodal video model performance when benchmark label reliability is unknown?Many video benchmarks publish labels without reliability statistics, making it unclear how much observed error reflects the model versus annotation noise. Dense classroom coding further requires interpreting many behavior labels over time, including codes that may be only partly observable from video.