Abstract
This publication details a hardware architecture and data processing methodology for dynamic exponent and mantissa repurposing, tailored to facilitate adaptive microscaling within machine learning accelerators. The described methodology addresses the quantization errors and memory bandwidth constraints that frequently emerge during the deployment of large-scale artificial intelligence models. By dynamically repurposing redundant exponent bits into additional mantissa bits specifically for outlier values within a given tensor block, the system dramatically enhances numerical precision without expanding the core memory footprint of the data block. A dedicated hardware decoder can be positioned immediately prior to a dot-product engine to re-route electrical signals and dequantize values on the fly, thereby maintaining seamless compatibility with existing arrays of multiply-accumulate units.
Keywords: microscaling, quantization, artificial intelligence accelerator, hardware decoder, floating-point format, dot-product engine, exponent repurposing, signal-to-quantization-noise ratio, block floating-point, tensor processing.
Creative Commons License

This work is licensed under a Creative Commons Attribution 4.0 License.
Recommended Citation
N/A, "Dynamic Exponent-Mantissa Repurposing for Adaptive Microscaling (DEMR-AM) Architecture", Technical Disclosure Commons, (August 14, 2026)
https://www.tdcommons.org/dpubs_series/11363