Text to speech
Speech from text with Chatterbox, with adjustable delivery and share links.
This tool ran on a GPU at my place, and it's been offline since I moved to New York. Below is how it worked. The engineering behind all of them is in austn.net's GPU tools.
- model
- Chatterbox, on its own Flask server
Paste text, pick a voice, and tune how it's delivered: an exaggeration control for how expressive it sounds, and a guidance weight for how closely it sticks to the voice. Any clip could be shared with a link. I could also generate clips in bulk from a CSV (admin only).
How it worked
- Text to speech ran on its own small Python server next to ComfyUI, with a health route the site checked.
- Each request became a job like every other tool, so it waited its turn for the GPU.
- A share link points at a stored clip by token, with its own page and an embed view. Old links expire and get cleaned up on a schedule.
- There's also a small JSON API for generating clips from scripts.
Voice cloning
Chatterbox can clone a voice from a short sample. On the public site that's only open to me, the admin, and every generation is logged.