I don’t think we’re even close to exhausting the potential of transformer architectures. gpt4o shows that a huge amount can be gained by implementing work done on understanding other media modalities. There’s a lot of audio that they can continue to train on still and the voice interactions they collect will go into further fine tuning. Even after that plays out there will be video to integrate next and thanks to physics simulations and 3D rendering there is a potentially endless and readily generated license free supply of it, at least for the simpler examples. For more complex real world video they could just set up web cams in public areas around the world where consent isn’t required by law and collect masses of data every second. Given that audio seems to have enabled emotional understanding and possibly even humour, I can’t imagine what all might fall out of video. At the least it’s going to improve reasoning since it will involve predicting cause and effect. There are probably a lot of others you could add though we don’t have large datasets for them.