Inventor(s)

Abstract

As semiconductor technology scales deep into nanoscale dimensions (sub-65 nm), subthreshold leakage current threatens the viability of tiled spatial microprocessors. While spatial architectures—exemplified by the MIT/UCSD RAW processor—elegantly eliminate global wire delays and expose on-chip functional units, interconnects, and pins directly to software, their massive replication of execution resources results in substantial static power waste when application workloads exhibit spatial or temporal underutilization. Traditional reactive dynamic power management (DPM) schemes, such as OS-directed ACPI states or hardware activity timers, fail in spatial architectures due to unpredictable wake-up transition penalties, loss of cycle-level synchronization across scalar operand networks (SONs), and inability to decouple communication relays from idle computation cores. In this paper, we propose a synergistic compilermicroarchitecture framework that exploits the static, cycle-accurate determinism of spatial computing to perform fine-grained, proactive energy orchestration. First, we introduce an asymmetric Split-Island Tile Microarchitecture that decouples each identical tile into two independent, multi-threshold CMOS (MTCMOS) powergated domains: a Computation Island (C-Island) housing the 8-stage MIPS core, FPU, and cache memories, and a Routing Island (R-Island) housing the static switch processor, dynamic routers (MDN/GDN), and crossbars. Second, we present an on-chip Closed-Loop Power Infrastructure comprising distributed header switches, inter-island isolation rings, a centralized Power Manager, and a Clock Control Unit enforcing a global homogeneous frequency invariant to maintain synchronous interconnect timing. Third, we develop a CompilerOrchestrated Power Gating strategy that profiles spatial compute/routing duty cycles, schedules look-ahead wake-up hints to completely hide the RC rail settling time, and executes cascade power-gating in pipelined stream graphs with state retention on a designated terminal tile. Circuit-level SPICE simulations and cycleaccurate architectural evaluations across Mediabench, StreamIt, and SPLASH-2 benchmarks at 45 nm demonstrate that our framework eliminates up to 78.4% of static leakage in routing/idle tiles and achieves an average 41.8% reduction in total chip energy with less than 1.2% execution time degradation and 3.4% silicon area overhead.

Creative Commons License

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.

Share

COinS