diff --git a/docs/superpowers/plans/2026-08-08-voice-studio.md b/docs/superpowers/plans/2026-08-08-voice-studio.md new file mode 100644 index 0000000..2099ff5 --- /dev/null +++ b/docs/superpowers/plans/2026-08-08-voice-studio.md @@ -0,0 +1,223 @@ +# Chatterbox Voice Studio Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Add LAN-open voice-file management and the approved two-column Chatterbox Studio without changing the public speech API shape. + +**Architecture:** Extract safe local voice-file operations into `voice_store.py`, then expose a small JSON/multipart API from FastAPI. Serve a dependency-free HTML/CSS/JS studio from the same FastAPI service; the browser holds only selected relative names and sends them to the existing OpenAI-compatible speech endpoint. + +**Tech Stack:** Python 3.12, FastAPI, Pydantic, python-multipart, standard-library `wave`, vanilla HTML/CSS/JavaScript, pytest. + +## Global Constraints + +- The UI is open to everyone on the local network; add no authentication. +- Store and serve reference audio only from `/home/vince/ai/apps/chatterbox-tts/voices`. +- Upload WAV only, with a maximum exact size of `20 * 1024 * 1024` bytes. +- Never accept absolute paths, traversal, symlink escapes, or non-regular files. +- Sanitize filenames for storage and render filenames via browser `textContent`, never `innerHTML`. +- Require an explicit UI confirmation before DELETE requests. +- Keep `/v1/audio/speech` OpenAI-compatible and WAV-only; a selected voice is its stored relative filename. +- Keep the existing 2,000-character input cap, one-request GPU lock, and 300-second unload. +- Do not change Kokoro, Caddy, local DNS, service binding, or systemd configuration. + +--- + +### Task 1: Implement a safe, test-covered voice store and API + +**Files:** +- Create: `voice_store.py` +- Modify: `app.py` +- Modify: `requirements.txt` +- Modify: `tests/test_app.py` + +**Interfaces:** +- Produces: `VoiceStore.list() -> list[VoiceMetadata]`, `VoiceStore.save(filename: str, payload: bytes) -> VoiceMetadata`, `VoiceStore.delete(name: str) -> None`, and `VoiceStore.resolve(name: str) -> Path`. +- Produces: `GET /v1/voices`, `POST /v1/voices`, and `DELETE /v1/voices/{name}`. + +- [ ] **Step 1: Write failing voice-management tests** + +```python +def test_list_voices_returns_safe_wav_metadata(client, voices_dir): + write_wav(voices_dir / "warm voice.wav") + payload = client.get("/v1/voices").json() + assert payload["data"][0]["name"] == "warm-voice.wav" + assert payload["data"][0]["sample_rate"] == 24000 + +def test_upload_rejects_non_wav_and_oversized_files(client): + assert client.post("/v1/voices", files={"file": ("x.mp3", b"x")}).status_code == 400 + assert client.post("/v1/voices", files={"file": ("x.wav", b"0" * (20 * 1024 * 1024 + 1))}).status_code == 413 + +def test_upload_and_delete_never_escape_voice_directory(client): + assert client.post("/v1/voices", files={"file": ("../../bad.wav", valid_wav_bytes())}).status_code == 400 + assert client.delete("/v1/voices/../reference.wav").status_code == 400 +``` + +- [ ] **Step 2: Run the new tests and verify RED** + +Run: `PYTHONPATH=. .venv/bin/pytest -q tests/test_app.py -k 'voice or upload or delete'` + +Expected: FAIL because the voice-store module and endpoints do not exist. + +- [ ] **Step 3: Implement the safe store and endpoints** + +Use `Path(filename).name` as a starting point, reject an empty/path-bearing name, normalize displayed/stored names to a safe ASCII `a-z`, `0-9`, `-`, `_`, `.` set, and preserve only the `.wav` suffix. In the endpoint, read the upload then reject payloads over the exact byte cap before calling `save()`. Validate nonempty WAV headers/frames with `wave.open(BytesIO(payload))`. Write first to a unique file inside `VOICES_DIR`, then atomically rename it to the final path after rechecking containment and collision. `list()` must skip symlinks/non-WAV files and return sorted metadata. `delete()` must resolve containment and refuse missing/non-regular files. + +```python +@app.post("/v1/voices", status_code=201) +async def upload_voice(file: UploadFile) -> dict[str, VoiceMetadata]: + payload = await file.read() + return {"data": voice_store.save(file.filename or "", payload)} + +@app.delete("/v1/voices/{name}", status_code=204) +def delete_voice(name: str) -> Response: + voice_store.delete(name) + return Response(status_code=204) +``` + +Add `python-multipart` to `requirements.txt` for multipart parsing. + +- [ ] **Step 4: Run focused and full API tests** + +Run: `PYTHONPATH=. .venv/bin/pytest -q tests/test_app.py` + +Expected: PASS; no test loads CUDA or downloads a model. + +- [ ] **Step 5: Commit the voice API** + +```bash +git add voice_store.py app.py requirements.txt tests/test_app.py +git commit -m "feat: add safe voice management API" +``` + +### Task 2: Use the managed voice store for speech selection + +**Files:** +- Modify: `app.py` +- Modify: `tests/test_app.py` + +**Interfaces:** +- Consumes: `VoiceStore.resolve(name) -> Path` from Task 1. +- Produces: the existing `POST /v1/audio/speech` accepting only a stored relative voice name. + +- [ ] **Step 1: Write failing speech-selection tests** + +```python +def test_speech_uses_selected_stored_voice(client, voices_dir, monkeypatch): + write_wav(voices_dir / "selected.wav") + seen = {} + def generate(text, voice, speed): + seen["voice"] = voice + return b"wav" + monkeypatch.setattr(app.model_manager, "generate", generate) + response = client.post("/v1/audio/speech", json={"input": "Hello", "voice": "selected.wav"}) + assert response.status_code == 200 + assert seen["voice"] == voices_dir / "selected.wav" + +def test_speech_rejects_deleted_or_escaped_voice(client): + assert client.post("/v1/audio/speech", json={"input": "Hello", "voice": "deleted.wav"}).status_code == 400 + assert client.post("/v1/audio/speech", json={"input": "Hello", "voice": "../x.wav"}).status_code == 400 +``` + +- [ ] **Step 2: Run the selection tests and verify RED** + +Run: `PYTHONPATH=. .venv/bin/pytest -q tests/test_app.py -k 'selected_stored or deleted_or_escaped'` + +Expected: FAIL until speech resolution delegates to `VoiceStore`. + +- [ ] **Step 3: Refactor speech resolution to the store** + +Replace the standalone `resolve_voice()` call with `voice_store.resolve(request.voice)`. Preserve error status/detail behavior, response content type `audio/wav`, 2,000-character Pydantic cap, and ModelManager locking/unload behavior. Do not introduce model loading to voice-management endpoints. + +- [ ] **Step 4: Run all API tests** + +Run: `PYTHONPATH=. .venv/bin/pytest -q tests/test_app.py` + +Expected: PASS with speech selection and all previous safety cases covered. + +- [ ] **Step 5: Commit the speech integration** + +```bash +git add app.py tests/test_app.py +git commit -m "refactor: resolve speech voices through voice store" +``` + +### Task 3: Build the approved Studio UI and wire its flow + +**Files:** +- Create: `static/index.html` +- Create: `static/studio.css` +- Create: `static/studio.js` +- Modify: `app.py` +- Modify: `tests/test_app.py` + +**Interfaces:** +- Consumes: all Task 1 voice endpoints and `POST /v1/audio/speech` from Task 2. +- Produces: `GET /` serving the Studio and a responsive browser flow for upload, select, confirmed delete, generate, play, and download. + +- [ ] **Step 1: Write failing static-route tests** + +```python +def test_root_serves_studio(client): + response = client.get("/") + assert response.status_code == 200 + assert "Chatterbox Studio" in response.text + assert 'src="/static/studio.js"' in response.text +``` + +- [ ] **Step 2: Run the route test and verify RED** + +Run: `PYTHONPATH=. .venv/bin/pytest -q tests/test_app.py::test_root_serves_studio` + +Expected: FAIL with 404 before static assets and the root route exist. + +- [ ] **Step 3: Implement the studio** + +Mount `/static` with FastAPI `StaticFiles` and return `static/index.html` for `/`. Implement the accepted true-white, left-library/right-generator layout with a mobile stacked media query. In `studio.js`, use `fetch` for `/v1/voices`; use `FormData` for WAV upload; create all list text using `element.textContent`; keep selected state as a relative filename; show an explicit in-page modal/panel with Cancel and Delete before calling DELETE; submit JSON to `/v1/audio/speech`; render the returned Blob using `URL.createObjectURL` in an `