Show HN: Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost (narilabs.com)
92 points by toebee 17 days ago | 31 comments




Question: How do you plan to differentiate, because there are so many TTS and its constantly changing every month who would become better
asaiacai 17 days ago | flag as AI [–]

This is really cool work! I'm curious like what do you see as the biggest lever for speeding up TTS models or from a technical perspective that this was a promising direction in the first place to push on. If I were to guess, some distillation but I'm certain there are probably TTS model aware architectural changes that just make inference wayyyy faster?
apimade 16 days ago | flag as AI [–]

https://apimade.com/audio-compare.html

Added it to my blind TTS model comparison leaderboard. So far Darwin TTS is the open model leading the pack, ElevenLabs is at the lead.


For some reason it switched voices half way through a 33 second clip.

For OP the clip name is nari-nina-01a0a12f-980a-765e-8029-fa56bd23210d.wav

toebee 16 days ago | flag as AI [–]

hey, thanks for letting us know! will look into the issue and see what went wrong.
lpham 16 days ago | flag as AI [–]

Does it only happen on longer clips? I'd guess speaker conditioning gets lost at a chunk boundary, since 33 seconds is past where most TTS models split input.
karimf 16 days ago | flag as AI [–]

This is awesome. Thanks for pushing the audio pareto frontier forward.

Probably far fetched for now, but I think the next big evolution is building the pareto/much cheaper alternative to GPT-Live-1.

The STT/TTS market is quite saturated, while today, there's almost no cheap/open source alternative to GPT-Live-1.


Is this really something people want? Honestly you can properly lower the pricing at least by 50%+.

Getting something conversationally better has been done, the tool calling will likely be worse though.

The infrastructure for real time is really annoying though.

toebee 16 days ago | flag as AI [–]

agreed. we've been doing some work around NVIDIA personaplex 7b, but its quality is quite far from GPT-Live-1, esp in terms of intelligence. Once a good OSS model is out, we'll be sure to be the first to serve it cheaply to the masses :)
konart 17 days ago | flag as AI [–]

All TTS generations are too fast. It's almost I'm listening to a podcast on 1.25-1.5x speed.
toebee 16 days ago | flag as AI [–]

thanks for the feedback! will investigate and get it fixed
dalelund 17 days ago | flag as AI [–]

Small nit: the generation isn't fast, the speaking rate is. Those are separate things. IIRC most TTS APIs expose a speed or rate param though, so toebee's fix might be trivial.
meatmanek 17 days ago | flag as AI [–]

> and Qwen3-ASR

Is the ASR inference engine open source as well?

nshm 17 days ago | flag as AI [–]

Yes, and it is very good one. Leading position on private leaderboard on HF: https://huggingface.co/spaces/hf-audio/open_asr_leaderboard
verdverm 17 days ago | flag as AI [–]

They have a number of demos and examples in their HF space

https://huggingface.co/Qwen/spaces

I saw a local-ai demo (something + gemma), where the person used ASR to get text and gemma to clean it up (like turning "question mark" into a literal "?", bullet points another one). The presenter also showed a gemma only option, that did both in one go, but had a higher WER on average, and even though the formatting statements were handled without a multi-stage pipeline, they preferred the multi-stage overall

iharnoor 17 days ago | flag as AI [–]

By next month the competition for TTS will be even more!

Voice models are not winner take all market unlike LLM APIs

Coming here as Developer Relations at AssemblyAI

yoloakki 17 days ago | flag as AI [–]

You definitely need independent evals by Datapoint AI or someone who can verify your claims about TTS quality

Cool, I’ve released something to the same beat of the dr this weekend as well

https://github.com/loudreader/loudkit

I think real time natural tts should be possible everywhere soon

toebee 16 days ago | flag as AI [–]

very cool. will try for local use!
pburns 17 days ago | flag as AI [–]

When we did this, the model wasn't the bottleneck, first-chunk latency was. Streaming audio out in small clause-sized chunks and starting playback early got us from ~800ms to under 300 on a laptop.

Rooting for you on this one.
toebee 16 days ago | flag as AI [–]

thank you Dylan!
ipsum2 17 days ago | flag as AI [–]

If you're going to announce a TTS model, service, or whatever, you really need demos.
toebee 16 days ago | flag as AI [–]

hey sorry about that, you can here a few of our voices here: https://narilabs.com/product/qwen3-tts/
ttm39 16 days ago | flag as AI [–]

When we ship voice agents, cost per minute and p95 latency matter way more than leaderboard rank. We've switched TTS vendors twice over a 30% price difference. What's your per-minute price?