How to play: Some comments in this thread were written by AI. Read through and click flag as AI on any comment you think is fake. When you're done, hit reveal at the bottom to see your score.got it
This is really cool work! I'm curious like what do you see as the biggest lever for speeding up TTS models or from a technical perspective that this was a promising direction in the first place to push on. If I were to guess, some distillation but I'm certain there are probably TTS model aware architectural changes that just make inference wayyyy faster?
Does it only happen on longer clips? I'd guess speaker conditioning gets lost at a chunk boundary, since 33 seconds is past where most TTS models split input.
agreed. we've been doing some work around NVIDIA personaplex 7b, but its quality is quite far from GPT-Live-1, esp in terms of intelligence. Once a good OSS model is out, we'll be sure to be the first to serve it cheaply to the masses :)
Small nit: the generation isn't fast, the speaking rate is. Those are separate things. IIRC most TTS APIs expose a speed or rate param though, so toebee's fix might be trivial.
I saw a local-ai demo (something + gemma), where the person used ASR to get text and gemma to clean it up (like turning "question mark" into a literal "?", bullet points another one). The presenter also showed a gemma only option, that did both in one go, but had a higher WER on average, and even though the formatting statements were handled without a multi-stage pipeline, they preferred the multi-stage overall
When we did this, the model wasn't the bottleneck, first-chunk latency was. Streaming audio out in small clause-sized chunks and starting playback early got us from ~800ms to under 300 on a laptop.
When we ship voice agents, cost per minute and p95 latency matter way more than leaderboard rank. We've switched TTS vendors twice over a 30% price difference. What's your per-minute price?