Abstract

Finding specific moments in media can be challenging, as certain methods like timeline scrubbing or keyword searches may not align with a user's descriptive memory of a scene. A system can facilitate media navigation through natural language queries. For example, a system may create a time-stamped, multimodal index of a media asset by analyzing its visual, auditory, and textual content. When a user provides a descriptive query, a generative artificial intelligence model can semantically evaluate the request against the index to identify a corresponding media segment and its timecode. The system could then provide the timecode to a media player to execute a seek operation. This approach can provide a mechanism for locating content based on a conversational description of a scene, offering a potential alternative to time-based or keyword-based navigation techniques.

Creative Commons License

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.

Share

COinS