Abstract

Existing artificial intelligence (AI) presentation generation tools produce static, monolithic outputs that lack post-generation editability, forcing complete regeneration for minor modifications. Additionally, raw generative AI application programming interfaces (APIs) impose strict rate limits and require complex backend orchestration to coordinate text, image, and audio synthesis at scale. The disclosed technology provides a componentized multi-modal presentation generation platform that produces editable presentation bundles comprising structured text layouts, context-aware AI-generated imagery, and synchronized text-to-speech narration. The platform employs a coordinated pipeline with task queues and automatic retry mechanisms to manage API rate limits. Each presentation element remains decoupled and independently editable at the slide level through a WYSIWYG interface. A hybrid storage architecture streams content from volatile memory buffers while persisting assets to user-owned cloud storage via OAuth-based authentication, enabling non-destructive session restoration.

Creative Commons License

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.

Share

COinS