docs: plan Chatterbox voice studio

This commit is contained in:
vince 2026-08-08 23:54:12 +02:00
parent 96412f7629
commit 710139443c

View file

@ -0,0 +1,223 @@
# Chatterbox Voice Studio Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** Add LAN-open voice-file management and the approved two-column Chatterbox Studio without changing the public speech API shape.
**Architecture:** Extract safe local voice-file operations into `voice_store.py`, then expose a small JSON/multipart API from FastAPI. Serve a dependency-free HTML/CSS/JS studio from the same FastAPI service; the browser holds only selected relative names and sends them to the existing OpenAI-compatible speech endpoint.
**Tech Stack:** Python 3.12, FastAPI, Pydantic, python-multipart, standard-library `wave`, vanilla HTML/CSS/JavaScript, pytest.
## Global Constraints
- The UI is open to everyone on the local network; add no authentication.
- Store and serve reference audio only from `/home/vince/ai/apps/chatterbox-tts/voices`.
- Upload WAV only, with a maximum exact size of `20 * 1024 * 1024` bytes.
- Never accept absolute paths, traversal, symlink escapes, or non-regular files.
- Sanitize filenames for storage and render filenames via browser `textContent`, never `innerHTML`.
- Require an explicit UI confirmation before DELETE requests.
- Keep `/v1/audio/speech` OpenAI-compatible and WAV-only; a selected voice is its stored relative filename.
- Keep the existing 2,000-character input cap, one-request GPU lock, and 300-second unload.
- Do not change Kokoro, Caddy, local DNS, service binding, or systemd configuration.
---
### Task 1: Implement a safe, test-covered voice store and API
**Files:**
- Create: `voice_store.py`
- Modify: `app.py`
- Modify: `requirements.txt`
- Modify: `tests/test_app.py`
**Interfaces:**
- Produces: `VoiceStore.list() -> list[VoiceMetadata]`, `VoiceStore.save(filename: str, payload: bytes) -> VoiceMetadata`, `VoiceStore.delete(name: str) -> None`, and `VoiceStore.resolve(name: str) -> Path`.
- Produces: `GET /v1/voices`, `POST /v1/voices`, and `DELETE /v1/voices/{name}`.
- [ ] **Step 1: Write failing voice-management tests**
```python
def test_list_voices_returns_safe_wav_metadata(client, voices_dir):
write_wav(voices_dir / "warm voice.wav")
payload = client.get("/v1/voices").json()
assert payload["data"][0]["name"] == "warm-voice.wav"
assert payload["data"][0]["sample_rate"] == 24000
def test_upload_rejects_non_wav_and_oversized_files(client):
assert client.post("/v1/voices", files={"file": ("x.mp3", b"x")}).status_code == 400
assert client.post("/v1/voices", files={"file": ("x.wav", b"0" * (20 * 1024 * 1024 + 1))}).status_code == 413
def test_upload_and_delete_never_escape_voice_directory(client):
assert client.post("/v1/voices", files={"file": ("../../bad.wav", valid_wav_bytes())}).status_code == 400
assert client.delete("/v1/voices/../reference.wav").status_code == 400
```
- [ ] **Step 2: Run the new tests and verify RED**
Run: `PYTHONPATH=. .venv/bin/pytest -q tests/test_app.py -k 'voice or upload or delete'`
Expected: FAIL because the voice-store module and endpoints do not exist.
- [ ] **Step 3: Implement the safe store and endpoints**
Use `Path(filename).name` as a starting point, reject an empty/path-bearing name, normalize displayed/stored names to a safe ASCII `a-z`, `0-9`, `-`, `_`, `.` set, and preserve only the `.wav` suffix. In the endpoint, read the upload then reject payloads over the exact byte cap before calling `save()`. Validate nonempty WAV headers/frames with `wave.open(BytesIO(payload))`. Write first to a unique file inside `VOICES_DIR`, then atomically rename it to the final path after rechecking containment and collision. `list()` must skip symlinks/non-WAV files and return sorted metadata. `delete()` must resolve containment and refuse missing/non-regular files.
```python
@app.post("/v1/voices", status_code=201)
async def upload_voice(file: UploadFile) -> dict[str, VoiceMetadata]:
payload = await file.read()
return {"data": voice_store.save(file.filename or "", payload)}
@app.delete("/v1/voices/{name}", status_code=204)
def delete_voice(name: str) -> Response:
voice_store.delete(name)
return Response(status_code=204)
```
Add `python-multipart` to `requirements.txt` for multipart parsing.
- [ ] **Step 4: Run focused and full API tests**
Run: `PYTHONPATH=. .venv/bin/pytest -q tests/test_app.py`
Expected: PASS; no test loads CUDA or downloads a model.
- [ ] **Step 5: Commit the voice API**
```bash
git add voice_store.py app.py requirements.txt tests/test_app.py
git commit -m "feat: add safe voice management API"
```
### Task 2: Use the managed voice store for speech selection
**Files:**
- Modify: `app.py`
- Modify: `tests/test_app.py`
**Interfaces:**
- Consumes: `VoiceStore.resolve(name) -> Path` from Task 1.
- Produces: the existing `POST /v1/audio/speech` accepting only a stored relative voice name.
- [ ] **Step 1: Write failing speech-selection tests**
```python
def test_speech_uses_selected_stored_voice(client, voices_dir, monkeypatch):
write_wav(voices_dir / "selected.wav")
seen = {}
def generate(text, voice, speed):
seen["voice"] = voice
return b"wav"
monkeypatch.setattr(app.model_manager, "generate", generate)
response = client.post("/v1/audio/speech", json={"input": "Hello", "voice": "selected.wav"})
assert response.status_code == 200
assert seen["voice"] == voices_dir / "selected.wav"
def test_speech_rejects_deleted_or_escaped_voice(client):
assert client.post("/v1/audio/speech", json={"input": "Hello", "voice": "deleted.wav"}).status_code == 400
assert client.post("/v1/audio/speech", json={"input": "Hello", "voice": "../x.wav"}).status_code == 400
```
- [ ] **Step 2: Run the selection tests and verify RED**
Run: `PYTHONPATH=. .venv/bin/pytest -q tests/test_app.py -k 'selected_stored or deleted_or_escaped'`
Expected: FAIL until speech resolution delegates to `VoiceStore`.
- [ ] **Step 3: Refactor speech resolution to the store**
Replace the standalone `resolve_voice()` call with `voice_store.resolve(request.voice)`. Preserve error status/detail behavior, response content type `audio/wav`, 2,000-character Pydantic cap, and ModelManager locking/unload behavior. Do not introduce model loading to voice-management endpoints.
- [ ] **Step 4: Run all API tests**
Run: `PYTHONPATH=. .venv/bin/pytest -q tests/test_app.py`
Expected: PASS with speech selection and all previous safety cases covered.
- [ ] **Step 5: Commit the speech integration**
```bash
git add app.py tests/test_app.py
git commit -m "refactor: resolve speech voices through voice store"
```
### Task 3: Build the approved Studio UI and wire its flow
**Files:**
- Create: `static/index.html`
- Create: `static/studio.css`
- Create: `static/studio.js`
- Modify: `app.py`
- Modify: `tests/test_app.py`
**Interfaces:**
- Consumes: all Task 1 voice endpoints and `POST /v1/audio/speech` from Task 2.
- Produces: `GET /` serving the Studio and a responsive browser flow for upload, select, confirmed delete, generate, play, and download.
- [ ] **Step 1: Write failing static-route tests**
```python
def test_root_serves_studio(client):
response = client.get("/")
assert response.status_code == 200
assert "Chatterbox Studio" in response.text
assert 'src="/static/studio.js"' in response.text
```
- [ ] **Step 2: Run the route test and verify RED**
Run: `PYTHONPATH=. .venv/bin/pytest -q tests/test_app.py::test_root_serves_studio`
Expected: FAIL with 404 before static assets and the root route exist.
- [ ] **Step 3: Implement the studio**
Mount `/static` with FastAPI `StaticFiles` and return `static/index.html` for `/`. Implement the accepted true-white, left-library/right-generator layout with a mobile stacked media query. In `studio.js`, use `fetch` for `/v1/voices`; use `FormData` for WAV upload; create all list text using `element.textContent`; keep selected state as a relative filename; show an explicit in-page modal/panel with Cancel and Delete before calling DELETE; submit JSON to `/v1/audio/speech`; render the returned Blob using `URL.createObjectURL` in an `<audio controls>` element and set a download anchor with `download="chatterbox.wav"`. Render errors in an `aria-live` status region.
- [ ] **Step 4: Run tests and perform browser QA**
Run: `PYTHONPATH=. .venv/bin/pytest -q tests/test_app.py`
Expected: PASS.
Run the service locally and use the Browser plugin; if unavailable, use Playwright Chromium and record that fallback. Verify desktop and mobile: upload a valid WAV, select it, cancel then confirm deletion, generate a short result, play it, and download WAV. Capture a screenshot for comparison with `docs/design/chatterbox-studio-concept.png`.
- [ ] **Step 5: Commit the studio**
```bash
git add app.py static/index.html static/studio.css static/studio.js tests/test_app.py
git commit -m "feat: add Chatterbox voice studio"
```
### Task 4: Deploy and verify the updated live service
**Files:**
- Modify: none expected
**Interfaces:**
- Consumes: existing `chatterbox-tts.service`, Caddy route, and local DNS record.
- Produces: the live studio and voice APIs on `https://chatterbox.dimensionlab.net`.
- [ ] **Step 1: Restart only the Chatterbox user service**
Run: `systemctl --user restart chatterbox-tts.service && systemctl --user is-active chatterbox-tts.service`
Expected: `active`.
- [ ] **Step 2: Verify live API and root page**
Run:
```bash
curl -fsS https://chatterbox.dimensionlab.net/v1/voices
curl -fsS https://chatterbox.dimensionlab.net/ | rg 'Chatterbox Studio'
curl -fsS https://chatterbox.dimensionlab.net/health
```
Expected: voices JSON, studio HTML, and `{"status":"ok"}`.
- [ ] **Step 3: Push the reviewed commits**
Run: `git push origin main`
Expected: the clean, verified implementation is published to `vince/chatterbox-tts`.