Abstract
This disclosure describes techniques for improving speech-to-text transcription and speaker labeling (also known as diarization or voice recognition) when an audio recording exceeds a transcription model's processing limit or context window. A computing system may buffer received audio into segments, using a target duration or a conversation-termination signal to determine segment length, and select segment boundaries at pauses in speech when a conversation spans multiple segments. To establish continuity across segments, the computing system may prepend a portion of a preceding segment to a subsequent segment, causing the transcription model to independently transcribe the portion of overlapping audio. The computing system may splice together two or more of the transcribed segments of a single conversation using the portion of overlapping audio to reconcile speaker identifiers, propagate speaker names, and construct a merged transcript with consistent identifiers throughout. The disclosed buffering and splicing may be performed together or separately. Aspects of this disclosure include audio buffering, conversation-termination detection, pause-based segmentation, overlapping segments, transcription, diarization, transcript splicing and alignment, speaker-identifier reconciliation, speaker-name propagation, and merged-transcript construction.
Creative Commons License

This work is licensed under a Creative Commons Attribution 4.0 License.
Recommended Citation
Kogan, David, "BUFFERING AND SPLICING AUDIO FOR TRANSCRIPTION", Technical Disclosure Commons, ()
https://www.tdcommons.org/dpubs_series/11789