Abstract

An active, hardware-software co-designed detachable accessory cryptographically couples with a host mobile device to securely offload and accelerate on-device machine learning execution, such as Large Language Models (LLMs). Governed by a kernel-level host execution scheduler, the system partitions model inference across host and auxiliary Neural Processing Units (NPUs) using asymmetrical speculative decoding to reduce inter-device communication overhead. To resolve mobile thermal and electrical limitations, the architecture integrates active Peltier cooling, dynamic tensor quantization scaling, and prefill-synchronized power shaving to protect host battery health and maintain sustained performance. System integrity and privacy are protected via a host secure enclave key escrow tied to a hardware clock-cycle heartbeat, which triggers immediate register zeroization upon physical detachment, and an active side-channel noise masking system that dynamically obfuscates NPU electromagnetic and thermal signatures.

Creative Commons License

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.

Share

COinS