Modern GPUs often use GDDR (Graphics Double Data Rate) or HBM. How does the memory subsystem of a GPU differ from a CPU's DDR subsystem regarding latency and parallelism?
A CPU memory controllers are optimized for Low Latency to ensure quick response times for serial tasks and OS interrupts.
B GPU memory controllers are optimized for Throughput and are designed to tolerate high latency by switching between thousands of active threads (latency hiding).
C Memory Coalescing: GPU hardware attempts to combine memory accesses from adjacent threads into a single transaction to maximize bus utilization, a feature less critical in standard CPU scalar execution.
D CPUs typically use massive L1/L2/L3 caches to reduce the effective memory latency, whereas GPUs devote more die area to ALUs and use smaller caches relative to their compute throughput.
E GDDR memory is pin-compatible with standard DDR memory, meaning you can plug GDDR6 chips directly into a standard PC motherboard DDR4 slot.