In Modern GPU Architectures (SIMT - Single Instruction, Multiple Threads), how are memory divergence and control divergence handled?