Two speech models can sustain a conversation while leaving longer gaps than human speakers. Coupled but Late measures that distinction using two PersonaPlex-7B instances exchanging audio tokens on a shared clock.
Across three ways of locating speech boundaries, their median floor-transfer offsets range from 400 to 560 milliseconds, compared with 137 milliseconds in the full-call Switchboard human reference. The study also tests whether that delay comes from the communication channel or the models' response policy.
Measure a conversation rather than one isolated model
A full-duplex model listens while it speaks. In this experiment, two instances exchange audio tokens at 80-millisecond steps, with one model per GPU. There is no live recognition or synthesis service between them adding its own timing.
The baseline consists of 36 seeded, 90-second conversations in five shared scenes. The authors keep personas, voices, scenes and seeds fixed across matched conditions.
A floor-transfer offset measures the interval between the end of the current speaker's speech unit and the beginning of the next speaker's unit. Positive values indicate a gap; negative values indicate overlap. The authors apply the same extraction rule to model conversations and Switchboard word alignments, while using three boundary estimators for generated speech.
Use controls to distinguish coupling from coincident activity
Re-pairing speakers from different conversations preserves each speaker's activity and the common clock, but removes the original partner. That change collapses the median offset toward zero and increases overlap. Re-pairing within the same scene produces similar results, supporting the interpretation that timing belongs to the original interaction.
A separate control replaces partner audio with encoded silence. The agents tend to give an opening line and then stop, rather than continuing independent monologues. Partner input therefore sustains their speech.
These tests support coupling between the models. They do not establish human-like turn prediction. Only about 1% of their transfers occur during the final 120 milliseconds of the partner's turn, compared with about 10% in the human reference.
Separate channel delay from a reactive waiting pattern
The communication setup adds up to 240 milliseconds of delay. Subtracting that maximum delay leaves a median gap of 240 milliseconds, still above the human references. Adding delay in one direction shifts the pooled response mode in that direction while leaving the opposite direction's mode unchanged.
The authors interpret the remaining wait and sparse anticipatory starts as consistent with reacting to the perceived end of a turn. Timing alone does not identify the model's internal mechanism.
The comparison also has limits. Switchboard contains telephone conversations between strangers, while the agents share prompted scenes. Replacing the live model partner with recorded human speech changes acoustics and content as well as interaction, so that condition cannot isolate a closed-loop effect by itself.
The findings concern one checkpoint and this experimental protocol. They warn that self-play dialogue or a model-based spoken examiner can inherit timing patterns different from those of human conversation. The study does not show that every full-duplex model has the same delay.
Comments