Abstract
This publication describes a system and methodology for improving speech recognition accuracy in smart home devices by transitioning from early-stage transcription biasing to late-stage semantic disambiguation. Conventional automatic speech recognition architectures prematurely resolve phonetic ambiguities by selecting only the single highest-ranked transcript, which often fails when processing niche media requests, regional dialects, or voice commands in noisy environments. The described framework preserves a complete set of alternative transcript hypotheses, packages them with their respective confidence metrics into a contextual container, and forwards them to a downstream semantic processing engine. By analyzing the phonetic relationships among these hypotheses alongside contextual signals such as user history, preferences, and content popularity, the semantic model resolves the user’s intent with significantly higher accuracy. Experimental results demonstrate substantial improvements across multiple conversational and execution metrics, demonstrating the robustness of this multi-hypothesis late-stage disambiguation paradigm in noisy acoustic environments. Keywords: automatic speech recognition, semantic disambiguation, speech hypothesis, smart home devices, and phonetic processing.
Creative Commons License

This work is licensed under a Creative Commons Attribution 4.0 License.
Recommended Citation
St. Marie, Linda; Mishra, Mubaraq; and , "Late-Stage Semantic Disambiguation of Multi-Hypothesis Automatic Speech Recognition in Digital Voice Platforms", Technical Disclosure Commons, ()
https://www.tdcommons.org/dpubs_series/11482