Models for multi-character voice experiences.
They are small and fast, built to run beside a live conversation rather than in a data centre. Each one does one job well: reading how a line was said, working out who it was said to, or speaking as a character. They answer in milliseconds on a CPU or a modest GPU, so when several AI characters share one conversation, in an interactive story, a game or a voice app, each can respond like a person in the room. They run locally, inside the application that uses them.
| Model | What it does | Latest release |
|---|---|---|
| SpellSpeak Tone | Reads one line of dialogue and says how the speaker sounds and what the line does to each person present: an emotion, a social act, how sure it is, and what the line does in the conversation | rc2.1 · 2026-10-09 |
| SpellSpeak Audience | Works out who a line is for when several AI characters are listening: one of them, the whole group, or unclear, so a character can ask "Who, me?" instead of taking offence | rc2.1 · 2026-10-09 |
Full precision is always there and is the reference. A release may also carry a smaller file with the usual _fp16 suffix, with its own measured numbers on the card. Every release has a MANIFEST.json listing its files with checksums.
Examples, demos and source: github.com/spellspeak/models. Everything is under the Apache-2.0 licence.