d-Matrix Raptor 3D-DRAM accelerator targets HBM limits at Hot Chips 2026
d-Matrix at Hot Chips 2026, d-Matrix unveiled its Raptor, a 3D-DRAM accelerator for AI inference, claiming a major memory bandwidth and power efficiency boost over HBM4. The design stacks logic directly on DRAM to overcome HBM's bandwidth ceiling. Raptor achieves roughly 100 TB/s bandwidth at 0.37 pJ/bit, about 10 times more energy-efficient than HBM4 systems. It uses a 36-micron face-to-face stacking process, with techniques like stream blocking and stream flipping to eliminate bandwidth waste and reduce I/O power. The accelerator targets low-latency inference, claiming 1,000 tokens per second per user for a 3-trillion-parameter model with 1M context. It offers 32.6 GB/s per mm², about 20 times higher than HBM4, though real-world performance remains unverified.