Abstract

Large artificial intelligence (AI) models consume substantial memory. Although reducing the numerical precision of model weights through quantization can help, choosing the appropriate precision for each layer is an exponentially complex problem complicated by unpredictable interactions between layers. Proposed herein is a profiling framework that explicitly measures cross-layer interactions and uses a multi-stage search pipeline to reduce trillions of candidate configurations to a few hundred targeted evaluations. Built-in safety margins and repair mechanisms help ensure that the final quantization plan meets a user-specified quality threshold. Experiments across multiple architectures demonstrated memory savings of 29-75% while maintaining strict quality limits and consistently outperforming random search.

Creative Commons License

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.

Share

COinS