-
Qwen3-TTS-Flash: Multi-timbre & Multi-lingual & Multi-dialect Speech Synthesis.
-
Autoregressive transformer models have proven highly effective for synthesizing speech from text. 🤗
-
🌟 MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder,
arXiv, 2505.07916, arxiv, pdf, cication: -1Bowen Zhang, Congchao Guo, Geng Yang, ..., Yuan Lu, Yucen He · (minimax-ai.github)
-
F5R-TTS: Improving Flow-Matching based Text-to-Speech with Group Relative Policy Optimization,
arXiv, 2504.02407, arxiv, pdf, cication: -1Xiaohui Sun, Ruitong Xiao, Jianye Mo, ..., Qun Yu, Baoxun Wang · (frontierlabs.github)
-
Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation 🤗
-
Hume AI dropped a text-to-speech LLM that understands emotional context 𝕏
-
Long-Form Speech Generation with Spoken Language Models,
arXiv, 2412.18603, arxiv, pdf, cication: -1Se Jin Park, Julian Salazar, Aren Jansen, ..., Yong Man Ro, RJ Skerry-Ryan · (google.github)
-
TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch,
arXiv, 2412.08237, arxiv, pdf, cication: -1Xingchen Song, Mengtao Xing, Changwei Ma, ..., Zhendong Peng, Zhiyong Wu
-
Debatts: Zero-Shot Debating Text-to-Speech Synthesis,
arXiv, 2411.06540, arxiv, pdf, cication: -1Yiqiao Huang, Yuancheng Wang, Jiaqi Li, ..., Shunsi Zhang, Zhizheng Wu
-
Very Attentive Tacotron: Robust and Unbounded Length Generalization in Autoregressive Transformer-Based Text-to-Speech,
arXiv, 2410.22179, arxiv, pdf, cication: -1Eric Battenberg, RJ Skerry-Ryan, Daisy Stanton, ..., Julian Salazar, David Kao · (sequence-layers - google)
· (x)
-
EmoSphere++: Emotion-Controllable Zero-Shot Text-to-Speech via Emotion-Adaptive Spherical Vector,
arXiv, 2411.02625, arxiv, pdf, cication: -1Deok-Hyeon Cho, Hyung-Seok Oh, Seung-Bin Kim, ..., Seong-Whan Lee
-
VibeVoice - microsoft
A Frontier Long Conversational Text-to-Speech Model
-
marvis-tts - Marvis-Labs
-
🌟 VibeVoice Technical Report,
arXiv, 2508.19205, arxiv, pdf, cication: -1Zhiliang Peng, Jianwei Yu, Wenhui Wang, ..., Yan Xia, Furu Wei
-
AudioStory: Generating Long-Form Narrative Audio with Large Language Models,
arXiv, 2508.20088, arxiv, pdf, cication: -1Yuxin Guo, Teng Wang, Yuying Ge, ..., Wei Zou, Ying Shan
-
tts - inworld-ai
-
🌟 higgs-audio - boson-ai
Redefining Expressiveness in Audio Generation
-
MOSS-TTSD - OpenMOSS
Text to Spoken Dialogue Generation
-
UniTTS - IDEA-Emdoor-Lab
An end-to-end TTS system without decoupling of acoustic and semantic information
-
🌟 dia - nari-labs
-
MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis,
arXiv, 2502.18924, arxiv, pdf, cication: -1Ziyue Jiang, Yi Ren, Ruiqi Li, ..., Xiang Yin, Zhou Zhao · (sditdemo.github)
-
csm - SesameAILabs
· (huggingface)
-
fastrtc - freddyaboulton
-
Overview of the Amphion Toolkit (v0.2),
arXiv, 2501.15442, arxiv, pdf, cication: -1Jiaqi Li, Xueyao Zhang, Yuancheng Wang, ..., Junan Zhang, Zhizheng Wu
-
LLaSA: Scaling Train-Time and Test-Time Compute for LLaMA-based Speech Synthesis 🤗
-
tts-generation-webui - rsxdalv
-
alltalk_tts - erew123