Research questionHow can fully time-to-first-spike spiking language models encode normalization and attention while preserving language-modeling quality?In time-to-first-spike coding, each neuron emits at most one spike within a time window, making operations such as layer normalization and attention difficult to represent. Supporting these operations is necessary for a fully time-to-first-spike language model that retains conventional language capabilities.