Microsoft Details Maia 200 AI Chip's Software-Defined Dataflow Design
Microsoft's paper introduces Maia 200, an AI inference accelerator claiming 145 Tflop/s FP4 and 5072 Tflop/s FP8 within a 750W power envelope and 7 TB/s HBM bandwidth. The chip exemplifies what the authors call a Software Defined Locally Accessed Dataflow Architecture (SDLA), shifting design focus from thread-centric to data-movement-centric execution to improve efficiency and scalability for large-scale AI inference, backed by a Flynn-inspired taxonomy of data management strategies. Microsoft presented the work at HotChips26, positioning Maia 200 as delivering top inference performance for production LLMs across its fleet.
- Maia 200: A Software Defined Dataflow System for Large-scale AI Acceleration
- Maia 200: Software-Defined Dataflow and All-Ethernet Networking
Discussion: 2 tweets from 2 authors · @thoefler, @Vengineer