
d-Matrix Unveils Raptor 3D-DRAM: A High-Bandwidth Path for Generative Inference
At Hot Chips 2026, d-Matrix unveils Raptor, a 3D-DRAM accelerator that stacks a logic die on a DRAM layer to deliver ultra-high bandwidth for generative AI inference. With a 1-Hi 32GB per card and a dense bank/channel layout, it uses stream-blocking and stream-flipping to achieve around 100 TB/s I/O and 0.37 pJ/bit, claiming ~32.6 GB/s per mm2 and 2.96 mW per GB/s versus HBM4. The system targets roughly 1,000 tokens/sec per user for frontier 3T-class models at 1M context, but faces intertwined challenges in bank-to-channel mapping, I/O power, and thermal reliability that will influence real-world viability.








