docs: specify Chatterbox voice studio
This commit is contained in:
parent
38a2a48f28
commit
44962e70c6
3 changed files with 65 additions and 1 deletions
2
.gitignore
vendored
2
.gitignore
vendored
|
|
@ -2,8 +2,8 @@
|
||||||
__pycache__/
|
__pycache__/
|
||||||
*.py[cod]
|
*.py[cod]
|
||||||
.pytest_cache/
|
.pytest_cache/
|
||||||
|
.superpowers/
|
||||||
|
|
||||||
# Reference clips are local operational data; keep the directory in Git only.
|
# Reference clips are local operational data; keep the directory in Git only.
|
||||||
voices/*
|
voices/*
|
||||||
!voices/.gitkeep
|
!voices/.gitkeep
|
||||||
|
|
||||||
|
|
|
||||||
BIN
docs/design/chatterbox-studio-concept.png
Normal file
BIN
docs/design/chatterbox-studio-concept.png
Normal file
Binary file not shown.
|
After Width: | Height: | Size: 1.1 MiB |
64
docs/superpowers/specs/2026-08-08-voice-studio-design.md
Normal file
64
docs/superpowers/specs/2026-08-08-voice-studio-design.md
Normal file
|
|
@ -0,0 +1,64 @@
|
||||||
|
# Chatterbox Voice Studio design
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
## Goal
|
||||||
|
|
||||||
|
Add a local-network web interface at `chatterbox.dimensionlab.net` where any
|
||||||
|
LAN visitor can upload, list, select, delete, and use local reference clips to
|
||||||
|
generate Chatterbox speech.
|
||||||
|
|
||||||
|
## Experience
|
||||||
|
|
||||||
|
The root route (`GET /`) serves a single-page studio. It uses a calm white,
|
||||||
|
two-column desktop layout with a responsive stacked mobile layout:
|
||||||
|
|
||||||
|
- The left voice-library panel lists stored clips with filename, duration,
|
||||||
|
format, and selection state. It provides an upload drop zone/button and an
|
||||||
|
explicit delete action with an in-page confirmation.
|
||||||
|
- The right generation panel contains the selected-voice summary, text field,
|
||||||
|
2,000-character counter, speed input, Generate button, and generated-audio
|
||||||
|
player with WAV download.
|
||||||
|
- Empty states tell visitors to upload a voice first. Generation, upload, and
|
||||||
|
delete show clear inline progress, success, and error states.
|
||||||
|
|
||||||
|
No authentication is required: this is intentionally available to everyone on
|
||||||
|
the local network. The service remains bound to loopback; DNS and Caddy define
|
||||||
|
its local-network exposure.
|
||||||
|
|
||||||
|
## Voice storage and API
|
||||||
|
|
||||||
|
Voice files continue to live exclusively under `VOICES_DIR`.
|
||||||
|
|
||||||
|
- `GET /v1/voices` returns safe metadata for each managed file: relative name,
|
||||||
|
byte size, duration, sample rate, and channel count.
|
||||||
|
- `POST /v1/voices` accepts a multipart `file` upload. The server sanitizes
|
||||||
|
the filename, rejects paths and collisions, accepts WAV only in this first
|
||||||
|
release, validates that the decoded audio is nonempty, and saves inside
|
||||||
|
`VOICES_DIR` only.
|
||||||
|
- `DELETE /v1/voices/{name}` deletes one validated, regular file below
|
||||||
|
`VOICES_DIR`; it rejects traversal, absolute paths, symlink escapes, and
|
||||||
|
missing files.
|
||||||
|
- `POST /v1/audio/speech` keeps its current OpenAI-compatible contract. The
|
||||||
|
studio submits the selected relative filename as `voice` and renders the
|
||||||
|
returned WAV locally in the browser.
|
||||||
|
|
||||||
|
The existing model lock continues to serialize synthesis. Voice-management
|
||||||
|
operations do not load the model. Deleting a clip does not interrupt an active
|
||||||
|
synthesis request because speech validates and resolves its reference path
|
||||||
|
before handing it to the model manager.
|
||||||
|
|
||||||
|
## Error handling and safety
|
||||||
|
|
||||||
|
All API failures are JSON with a clear 4xx/5xx detail. The UI renders these as
|
||||||
|
inline messages. Deletion requires a confirmation click but no login. The
|
||||||
|
upload limit is 20 MiB. Uploaded audio is never committed to Git; the existing
|
||||||
|
`voices/*` ignore rule remains in place.
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
|
||||||
|
Automated tests cover list metadata, WAV-only and 20 MiB upload validation,
|
||||||
|
filename/path containment, duplicate rejection, deletion behavior, and the
|
||||||
|
existing speech contract. Browser QA covers upload, selection, deletion
|
||||||
|
confirmation, speech generation/playback/download, desktop layout, and the
|
||||||
|
responsive stacked mobile layout.
|
||||||
Loading…
Add table
Add a link
Reference in a new issue