Inventor(s)

Abstract

Tiled manycore architectures established a modular paradigm that conquered the global wire-delay wall by decomposing monolithic processors into replicated, network-connected compute elements. Early academic prototypes, such as the MIT Raw microprocessor, and commercial implementations, including Tilera TILE64, TILEPro64, and Tile-Gx 100, proved the viability of pipelined scalar operand networks and 2D mesh on-chip interconnects. However, scaling tiled architectures beyond 100 tiles toward the multi-thousand-tile horizon encounters four crippling bottlenecks: (1) the Cartesian Diameter and Hop-Count Wall in 2D Manhattan meshes, which incurs severe routing latency and packet contention; (2) the Technology Utilization and Dark Silicon Wall, where uniform homogeneous tiles cannot respect thermal limits without severe underclocking; (3) the Directory Coherence Scalability Wall, where flat directory tables exhaust storage and bottleneck memory gateways; and (4) the Reliability and Multi-Tenancy Wall, where deep-submicron physical defects and cloud co-tenancy demand hardware-isolated failure containment without hypervisor serialization. This paper presents HexaTile, a next-generation heterogeneous tiled manycore architecture designed to dismantle these scaling barriers. HexaTile makes four foundational contributions: First, it introduces a non-Cartesian Beehive Honeycomb Topology with degree-6 connectivity per tile. We provide an isometric coordinate formulation (u,v,w) and prove that HexaTile achieves an analytical 28.9% reduction in network diameter and a 31.2% reduction in average hop count over an equalarea 2D mesh, while demonstrating that an offset-row layout synthesizes onto standard orthogonal Manhattan metal layers with zero design rule violations. Second, we formulate HexDOR, a provably deadlock-free dimension-ordered routing algorithm on isometric axes, alongside Hex-Adapt, a fault-resilient turn-restricted adaptive protocol. Third, HexaTile architects an advanced Decoupled 5-Plane Physical Interconnect (iMeshX)—segregating Memory (MDN-NG), Coherence (TDN-NG), Direct Register-Mapped Operands (UDN-NG), Diffused I/O (IDN-NG), and Reconfigurable Streaming (RSN)—proving that physical network replication outperforms virtual channel multiplexing in sub-5nm nodes by saving buffer power and eliminating head-of-line blocking. Fourth, HexaTile implements Adaptive Multicore Hardwall 2.0 (AMH-2), providing hardwarelevel line-speed link filtering for sub-cycle fault containment and secure spatial multi-tenancy. Full-system cycle-accurate gem5+BookSim simulations of 64, 256, and 1024-tile configurations in a 3nm GAA process across PARSEC 3.0, SPLASH-3, graph analytics, and Transformer workloads show that HexaTile reduces average network latency by 31.4%, improves saturation throughput by 27.8%, achieves an average 1.24× full-system speedup (1.48× on communication-intensive kernels), and cuts the energy-delay product (EDP) by 34.2% compared to 2D mesh baselines.

Creative Commons License

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.

Share

COinS