Consider two compute systems, Machine A and Machine B, designed to accelerate linear-algebra workloads. Machine A can issue 3 multiply operations per cycle and sustains a memory bandwidth of 0.1 words per cycle. Machine B can issue 1 multiply operation per cycle and sustains a memory bandwidth of 0.5 words per cycle. Both machines have a multiplier latency of 1 cycle and operate at the same clock frequency.
Assume the following throughout this problem: instructions execute in program order; in each cycle, all multiplications whose dependencies are satisfied are issued, up to the number of available multipliers; only multiplications and loads from main memory incur cost; load latency is fully hidden once data is available; all kernels execute in steady state on large inputs; results are kept on-chip and are not written back to main memory; and at any time at most one kernel is executing on a machine.
Consider the following kernels:
- Matrix–vector multiplication of an N×N matrix with an N-element vector.
- Matrix–matrix multiplication of two N×N matrices.
Which of the following statements is (are) correct?