ChatTTS

ChatTTS

Conversational text-to-speech model for natural, expressive dialogue.

5.0
Rating
--
Visits/mo

Screenshots

ChatTTS screenshot

Overview

ChatTTS is a cutting-edge conversational text-to-speech (TTS) model designed for dialogue scenarios such as chatbots and virtual assistants. It transforms text into dynamic, natural-sounding speech, supporting both English and Chinese. The model is trained on extensive data (100,000+ hours for the full version, 40,000 hours for the open-source version) to deliver expressive speech with fine-grained control over prosodic features like laughter, pauses, and interjections.

How to Use

To use ChatTTS, users input text into the provided interface. They can then refine the text and adjust parameters such as audio temperature, top_P, top_K, audio seed, and text seed before generating the output audio.

Core Features

Optimized for dialogue scenarios (Conversational TTS) Fine-grained control over prosodic features (laughter, pauses, interjections) Superior prosody compared to most open-source TTS models Supports English and Chinese languages Trained on extensive data for natural, expressive speech

Use Cases

  1. 1 Enhancing chatbots with natural, expressive dialogue
  2. 2 Powering virtual assistants with lifelike speech
  3. 3 Research and development in text-to-speech technology

Frequently Asked Questions

ChatTTS is a conversational text-to-speech model that produces natural, expressive dialogue with controllable emotions and elements like laughter. It requires a GPU with sufficient VRAM for real-time inference, though model stability can vary.
For details, please visit the official website of ChatTTS.