Research questionHow can conversational speech emotion recognition use cross-speaker context without confusing each speaker’s emotional trajectory?Emotion in dialogue depends on both a speaker’s own prior turns and what other participants say. Treating every adjacent utterance as one sequence can blur these distinct sources of temporal evidence across different time scales.