Inventor(s)

Abstract

This disclosure details a vision-centric multimodal system for mitigating false-positive face detections on portable computing devices. The system distinguishes physical human users from 2D media displays—including high-resolution photographs and digital screens—by cross-referencing a primary computer vision module with lightweight acoustic source localization and Wi-Fi channel state analysis. By evaluating the spatiotemporal synchronization among visual facial features, auditory signals, and wireless ambient micro-movements, the system identifies and suppresses non-physical facial entities in real-time. Crucially, this multi-modal pipeline operates locally on low-power edge hardware without invoking the primary application processor or relying on neural network acceleration, thereby maximizing energy efficiency and ensuring user privacy.

Creative Commons License

Creative Commons License
This work is licensed under a Creative Commons Attribution-Share Alike 4.0 License.

Share

COinS