題組題幹(本題:第29題,共 4 小題)點擊展開
In recent years, general purpose processors have increasingly been supplemented, and in some cases replaced, by specialized accelerators such as GPUs and TPUs. These accelerators depart from traditional CPU designs by prioritizing high throughput computation, data level parallelism, and memory bandwidth over single thread latency and complex control flow. While such designs can deliver orders of magnitude improvements for selected workloads, they also introduce new architectural tradeoffs at the system level. The following questions examine the architectural principles that distinguish accelerators from CPUs, and the system-level consequences of these design choices.
Consider a multi-accelerator system executing a workload that performs frequent collective operations (e.g., all-reduce) on large tensors. Which interconnect property is most critical for maximizing sustained throughput?