Thoughts
- It's fast (~3 seconds on my RTX 4090)
- Surprisingly capable of maintaining image integrity even at high resolutions (1536x1024, sometimes 2048x2048)
- The adherence is impressive for a 6B parameter model
Some tests (2 / 4 passed):
Personally I find it works better as a refiner model downstream of Qwen-Image 20b which has significantly better prompt understanding but has an unnatural "smoothness" to its generated images.
https://github.com/Tongyi-MAI/Z-Image
Screenshot of site with network tools open to indicate link
EDIT: It's possible that this issue might have existed in an old cached version. I'll purge the cache just to make sure.