Do you mean the coupled dies on stuff like the B200? An NVidia chip die has many SMs if so.
Do you mean TMEM MMA cooperative execution? I'm guessing that must be it given what the paper is about.
cooperative execution yeah
as you can tell I do not do CUDA for a living :D