zlacker

[parent] [thread] 1 comments
1. vessen+(OP)[view] [source] 2024-05-21 03:25:48
Greg has specifically said it's not an SSML-parsing text model; he's said it's an end to end multimodal model.

FWIW, I would find it very surprising if you could get the low latency expressiveness, singing, harmonizing, sarcasm and interpretation of incoming voice through SSML -- that would be a couple orders of magnitude better than any SSML product I've seen.

replies(1): >>nabaki+Fd3
2. nabaki+Fd3[view] [source] 2024-05-22 00:49:56
>>vessen+(OP)
Not sure about the low latency aspect, but I've seen everything else you mentioned with SSML. Also, I can't find where Greg said that, could you point me to it?
[go to top]