Items tagged with d-matrix
by
Zak Killian - Mon, Aug 24, 2026
In AI inference, you have two phases of the workload: prefill and decode. In most cases, decode is overwhelmingly the more time-consuming portion of the workload because it's strictly memory-bandwidth bound. You can have all the compute in...
Read more...