← Playground

Text to speech

Speech from text with Chatterbox, with adjustable delivery and share links.

This tool ran on a GPU at my place, and it's been offline since I moved to New York. Below is how it worked. The engineering behind all of them is in austn.net's GPU tools.

model
Chatterbox, on its own Flask server

Paste text, pick a voice, and tune how it's delivered: an exaggeration control for how expressive it sounds, and a guidance weight for how closely it sticks to the voice. Any clip could be shared with a link. I could also generate clips in bulk from a CSV (admin only).

How it worked

  • Text to speech ran on its own small Python server next to ComfyUI, with a health route the site checked.
  • Each request became a job like every other tool, so it waited its turn for the GPU.
  • A share link points at a stored clip by token, with its own page and an embed view. Old links expire and get cleaned up on a schedule.
  • There's also a small JSON API for generating clips from scripts.

Voice cloning

Chatterbox can clone a voice from a short sample. On the public site that's only open to me, the admin, and every generation is logged.