Abstract
This disclosure describes a new framework for isolating a specific person's voice from background noise, often called Target Speaker Extraction (TSE), speech isolation, voice separation, audio unmixing, voice filtering, target audio filtering, speaker segregation, acoustic filtering, vocal isolation, target voice retrieval, or speech decoupling for devices equipped with artificial intelligence. Traditional voice recognition systems require users to manually record their voice in a quiet room, which is frustrating and often fails to recognize natural changes in a person's voice like sickness or morning voice. Rather than requiring the user to perform this manual process, the new framework described herein enables the computing device to passively learn a user's unique speech patterns (a “voiceprint”, "vocal profile", “acoustic signature”, or "phonation signature") in the background during secure moments, like right after the user unlocks their phone. This approach, also known as zero enrollment, zero-shot registration, passive setup, automatic profiling, background onboarding, non-intrusive registration, seamless enrollment, or effortless setup, eliminates the need for manual training. When the user later speaks into their device using speech-to-text (such as a voice keyboard) or a voice assistant, the system uses an on-device lightweight machine learning model to separate their voice from any overlapping background chatter or other speakers. The isolated voice data is then turned into highly accurate text. By continuously updating the voiceprint in the background, this new framework eliminates setup friction and adapts to how the user sounds in real-time. Moreover, all sensitive biometric data and voiceprint creation happens entirely on the user's personal device, ensuring strict privacy and preventing data leaks.
Creative Commons License

This work is licensed under a Creative Commons Attribution 4.0 License.
Recommended Citation
Li, Zhi and Huang, Xuelin, "ZERO-ENROLLMENT TARGET SPEAKER EXTRACTION", Technical Disclosure Commons, (August 18, 2026)
https://www.tdcommons.org/dpubs_series/11402