Abstract
This disclosure describes techniques for detecting and mitigating large mistakes, often called block errors, in automated speech-to-text (STT) or voice recognition machine-generated transcriptions. A computing system may transcribe audio (such as creating closed captions or meeting notes) using a primary artificial intelligence (AI) transcription model and, in parallel, a secondary transcription model. In one implementation, the secondary transcription model is a computationally less expensive model that may provide a “checksum” or verification signal for the primary transcription model. The computing system may tokenize and align the resulting transcriptions side-by-side to identify regions containing unusually large insertions or deletions. Such regions may indicate transcription block errors of fabricated or omitted content. Upon detecting a potential block error, the computing system may generate one or more additional transcriptions using one or both models. The computing system may adjust transcription conditions, e.g., AI temperature settings, between attempts and select a final transcription based on alignment quality. Aspects of this disclosure include parallel transcription, comparison of independently generated transcriptions, transcription alignment, block error detection, repeated transcription, and selection among candidate transcriptions.
Creative Commons License

This work is licensed under a Creative Commons Attribution 4.0 License.
Recommended Citation
Kogan, David; Au, Lawrence; and Sun, Libin, "DETECTING TRANSCRIPTION ERRORS USING PARALLEL MODEL OUTPUTS", Technical Disclosure Commons, ()
https://www.tdcommons.org/dpubs_series/11788