Abstract

A system is proposed herein for distributing artificial intelligence (AI) inference across endpoint, edge, and cloud tiers. The system routes each semantic portion of a request only to an execution tier permitted by enterprise policy. A local model processes the request first, and intent, local-model confidence, and live network and compute telemetry determine whether selected portions remain local or are promoted with compressed intermediate state. Promoted results are verified before acceptance, and observed outcomes are used to improve future routing and expert placement.

Creative Commons License

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.

Share

COinS