Abstract
Generative world models may produce virtual environments that can be generic and lack grounding in specific, pre-existing contexts, or require significant resources and effort to enrich and articulate. A technique is described for generating a contextually grounded interactive environment from a source video. The method can involve a hierarchical analysis of the video, deconstructing it into scenes, shots, and frames and extracting multi-level contextual information using machine learning models. A representative composite image can be synthesized from distinct frames within a scene. This image, combined with textual context, may form a structured, multi-modal prompt. This prompt can then guide a foundational world model to generate a detailed, navigable simulation. This process can facilitate the transformation of passive video content into an explorable digital space, providing a localized representation of the world depicted in the video, which can be useful for prototyping interactive experiences.
Creative Commons License

This work is licensed under a Creative Commons Attribution 4.0 License.
Recommended Citation
Lui, Anthony, "Generating Contextually Grounded Virtual Worlds Using Hierarchical Video Analysis", Technical Disclosure Commons, ()
https://www.tdcommons.org/dpubs_series/12047