Abstract
Multimodal LLMs implemented on a server can support remote user interaction over a network but require continuous real-time streaming of raw, uncompressed video and/or audio. This can be infeasible due to bandwidth constraints or other reasons, and can prevent real-time feedback in time-sensitive interactive settings. This disclosure describes an asymmetric edge and cloud collaborative architecture that can enable full-duplex and seamlessly interruptible multimodal LLM orchestration at low latency. The architecture decouples sensor processing tasks from cognitive reasoning tasks involved in interaction between a user and a multimodal LLM via modalities such as audio, video, or image. Computationally heavy processing of perceptual input (e.g., audio and/or visual input) is performed within a localized sandbox on the edge user device. An out-of-band control plane supports interruption of ongoing generated output based on non-verbal contextual cues. Upon interruption, token-level context can be rolled back to within milliseconds of when the interruption is detected. The architecture supports user interaction with multimodal LLMs in instructional contexts such as learning to play the piano.
Creative Commons License

This work is licensed under a Creative Commons Attribution 4.0 License.
Recommended Citation
Zhu, Longlong and Huang, Weiyue, "Low Bandwidth Full-Duplex Multimodal LLM Interaction with Contextual Interruption", Technical Disclosure Commons, (September 16, 2026)
https://www.tdcommons.org/dpubs_series/11763