Abstract
This disclosure describes an operating-system-level passenger copilot, also referred to as an in-car assistant, passenger companion, or media helper, for vehicles equipped with passenger displays, rear-seat entertainment screens, or secondary displays. The system coordinates video screen image capture (frame ingestion), voice isolation (seat-isolated routing), cross-device watch history syncing (watch graph synchronization), and on-screen landmark navigation (point of interest extraction) to enhance the in-vehicle entertainment experience without distracting the driver. When a user taps an on-screen icon or speaks into a dedicated microphone, the vehicle operating system captures a video frame, screenshot, or image from a streaming video application. A multimodal artificial intelligence model, vision model, or image recognition system analyzes the captured image and the passenger’s isolated voice query to identify on-screen items, such as actors, clothing, landmarks, or geographic locations. The system routes answers or product links exclusively to the passenger output devices - such as localized headphones, headrest speakers, or a dedicated text overlay – better ensuring the driver is not disturbed. Additionally, the passenger profile connects to a cloud-based video discovery service or smart TV application programming interface to pull associated viewing history from home televisions or mobile devices. The system compares available movies, shows, or media content runtimes against a vehicle’s remaining estimated time of arrival (ETA) or travel duration to suggest matched recommendations that finish exactly when the trip ends. Finally, identified on-screen locations, destinations, or points of interest are packaged as a secure waypoint payload and sent over an internal vehicle network to a central navigation screen, allowing the driver to accept or reject the route stop via steering wheel buttons or physical console controls.
Creative Commons License

This work is licensed under a Creative Commons Attribution 4.0 License.
Recommended Citation
Gupta, Richa, "CONTEXT-AWARE MULTIMODAL GENERATIVE COPILOT AND CROSS-SURFACE DISCOVERY ENGINE FOR AUTOMOTIVE PASSENGER DISPLAYS", Technical Disclosure Commons, ()
https://www.tdcommons.org/dpubs_series/11521