Inventor(s)

Abstract

The central challenge of camera-based gesture input in video meetings is distinguishing deliberate interaction gestures from the natural hand motion that accompanies speech. This paper describes a system that meets that challenge through an intent classification engine evaluating kinematic, temporal, postural, spatial, and speech-aligned features over sliding temporal windows. Qualified gestures pass to a gesture-to-structure transformation pipeline that produces clean, labelled diagram elements—not freeform ink—and renders them as first-class shared content within the conferencing platform, visible to all participants in real time, without requiring specialised hardware.

Creative Commons License

Creative Commons License
This work is licensed under a Creative Commons Attribution-Share Alike 4.0 License.

Share

COinS