chatterbox-tts/docs/superpowers/plans/2026-08-08-voice-studio.md
2026-08-08 23:54:12 +02:00

9.8 KiB

Chatterbox Voice Studio Implementation Plan

For agentic workers: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (- [ ]) syntax for tracking.

Goal: Add LAN-open voice-file management and the approved two-column Chatterbox Studio without changing the public speech API shape.

Architecture: Extract safe local voice-file operations into voice_store.py, then expose a small JSON/multipart API from FastAPI. Serve a dependency-free HTML/CSS/JS studio from the same FastAPI service; the browser holds only selected relative names and sends them to the existing OpenAI-compatible speech endpoint.

Tech Stack: Python 3.12, FastAPI, Pydantic, python-multipart, standard-library wave, vanilla HTML/CSS/JavaScript, pytest.

Global Constraints

  • The UI is open to everyone on the local network; add no authentication.
  • Store and serve reference audio only from /home/vince/ai/apps/chatterbox-tts/voices.
  • Upload WAV only, with a maximum exact size of 20 * 1024 * 1024 bytes.
  • Never accept absolute paths, traversal, symlink escapes, or non-regular files.
  • Sanitize filenames for storage and render filenames via browser textContent, never innerHTML.
  • Require an explicit UI confirmation before DELETE requests.
  • Keep /v1/audio/speech OpenAI-compatible and WAV-only; a selected voice is its stored relative filename.
  • Keep the existing 2,000-character input cap, one-request GPU lock, and 300-second unload.
  • Do not change Kokoro, Caddy, local DNS, service binding, or systemd configuration.

Task 1: Implement a safe, test-covered voice store and API

Files:

  • Create: voice_store.py
  • Modify: app.py
  • Modify: requirements.txt
  • Modify: tests/test_app.py

Interfaces:

  • Produces: VoiceStore.list() -> list[VoiceMetadata], VoiceStore.save(filename: str, payload: bytes) -> VoiceMetadata, VoiceStore.delete(name: str) -> None, and VoiceStore.resolve(name: str) -> Path.

  • Produces: GET /v1/voices, POST /v1/voices, and DELETE /v1/voices/{name}.

  • Step 1: Write failing voice-management tests

def test_list_voices_returns_safe_wav_metadata(client, voices_dir):
    write_wav(voices_dir / "warm voice.wav")
    payload = client.get("/v1/voices").json()
    assert payload["data"][0]["name"] == "warm-voice.wav"
    assert payload["data"][0]["sample_rate"] == 24000

def test_upload_rejects_non_wav_and_oversized_files(client):
    assert client.post("/v1/voices", files={"file": ("x.mp3", b"x")}).status_code == 400
    assert client.post("/v1/voices", files={"file": ("x.wav", b"0" * (20 * 1024 * 1024 + 1))}).status_code == 413

def test_upload_and_delete_never_escape_voice_directory(client):
    assert client.post("/v1/voices", files={"file": ("../../bad.wav", valid_wav_bytes())}).status_code == 400
    assert client.delete("/v1/voices/../reference.wav").status_code == 400
  • Step 2: Run the new tests and verify RED

Run: PYTHONPATH=. .venv/bin/pytest -q tests/test_app.py -k 'voice or upload or delete'

Expected: FAIL because the voice-store module and endpoints do not exist.

  • Step 3: Implement the safe store and endpoints

Use Path(filename).name as a starting point, reject an empty/path-bearing name, normalize displayed/stored names to a safe ASCII a-z, 0-9, -, _, . set, and preserve only the .wav suffix. In the endpoint, read the upload then reject payloads over the exact byte cap before calling save(). Validate nonempty WAV headers/frames with wave.open(BytesIO(payload)). Write first to a unique file inside VOICES_DIR, then atomically rename it to the final path after rechecking containment and collision. list() must skip symlinks/non-WAV files and return sorted metadata. delete() must resolve containment and refuse missing/non-regular files.

@app.post("/v1/voices", status_code=201)
async def upload_voice(file: UploadFile) -> dict[str, VoiceMetadata]:
    payload = await file.read()
    return {"data": voice_store.save(file.filename or "", payload)}

@app.delete("/v1/voices/{name}", status_code=204)
def delete_voice(name: str) -> Response:
    voice_store.delete(name)
    return Response(status_code=204)

Add python-multipart to requirements.txt for multipart parsing.

  • Step 4: Run focused and full API tests

Run: PYTHONPATH=. .venv/bin/pytest -q tests/test_app.py

Expected: PASS; no test loads CUDA or downloads a model.

  • Step 5: Commit the voice API
git add voice_store.py app.py requirements.txt tests/test_app.py
git commit -m "feat: add safe voice management API"

Task 2: Use the managed voice store for speech selection

Files:

  • Modify: app.py
  • Modify: tests/test_app.py

Interfaces:

  • Consumes: VoiceStore.resolve(name) -> Path from Task 1.

  • Produces: the existing POST /v1/audio/speech accepting only a stored relative voice name.

  • Step 1: Write failing speech-selection tests

def test_speech_uses_selected_stored_voice(client, voices_dir, monkeypatch):
    write_wav(voices_dir / "selected.wav")
    seen = {}
    def generate(text, voice, speed):
        seen["voice"] = voice
        return b"wav"
    monkeypatch.setattr(app.model_manager, "generate", generate)
    response = client.post("/v1/audio/speech", json={"input": "Hello", "voice": "selected.wav"})
    assert response.status_code == 200
    assert seen["voice"] == voices_dir / "selected.wav"

def test_speech_rejects_deleted_or_escaped_voice(client):
    assert client.post("/v1/audio/speech", json={"input": "Hello", "voice": "deleted.wav"}).status_code == 400
    assert client.post("/v1/audio/speech", json={"input": "Hello", "voice": "../x.wav"}).status_code == 400
  • Step 2: Run the selection tests and verify RED

Run: PYTHONPATH=. .venv/bin/pytest -q tests/test_app.py -k 'selected_stored or deleted_or_escaped'

Expected: FAIL until speech resolution delegates to VoiceStore.

  • Step 3: Refactor speech resolution to the store

Replace the standalone resolve_voice() call with voice_store.resolve(request.voice). Preserve error status/detail behavior, response content type audio/wav, 2,000-character Pydantic cap, and ModelManager locking/unload behavior. Do not introduce model loading to voice-management endpoints.

  • Step 4: Run all API tests

Run: PYTHONPATH=. .venv/bin/pytest -q tests/test_app.py

Expected: PASS with speech selection and all previous safety cases covered.

  • Step 5: Commit the speech integration
git add app.py tests/test_app.py
git commit -m "refactor: resolve speech voices through voice store"

Task 3: Build the approved Studio UI and wire its flow

Files:

  • Create: static/index.html
  • Create: static/studio.css
  • Create: static/studio.js
  • Modify: app.py
  • Modify: tests/test_app.py

Interfaces:

  • Consumes: all Task 1 voice endpoints and POST /v1/audio/speech from Task 2.

  • Produces: GET / serving the Studio and a responsive browser flow for upload, select, confirmed delete, generate, play, and download.

  • Step 1: Write failing static-route tests

def test_root_serves_studio(client):
    response = client.get("/")
    assert response.status_code == 200
    assert "Chatterbox Studio" in response.text
    assert 'src="/static/studio.js"' in response.text
  • Step 2: Run the route test and verify RED

Run: PYTHONPATH=. .venv/bin/pytest -q tests/test_app.py::test_root_serves_studio

Expected: FAIL with 404 before static assets and the root route exist.

  • Step 3: Implement the studio

Mount /static with FastAPI StaticFiles and return static/index.html for /. Implement the accepted true-white, left-library/right-generator layout with a mobile stacked media query. In studio.js, use fetch for /v1/voices; use FormData for WAV upload; create all list text using element.textContent; keep selected state as a relative filename; show an explicit in-page modal/panel with Cancel and Delete before calling DELETE; submit JSON to /v1/audio/speech; render the returned Blob using URL.createObjectURL in an <audio controls> element and set a download anchor with download="chatterbox.wav". Render errors in an aria-live status region.

  • Step 4: Run tests and perform browser QA

Run: PYTHONPATH=. .venv/bin/pytest -q tests/test_app.py

Expected: PASS.

Run the service locally and use the Browser plugin; if unavailable, use Playwright Chromium and record that fallback. Verify desktop and mobile: upload a valid WAV, select it, cancel then confirm deletion, generate a short result, play it, and download WAV. Capture a screenshot for comparison with docs/design/chatterbox-studio-concept.png.

  • Step 5: Commit the studio
git add app.py static/index.html static/studio.css static/studio.js tests/test_app.py
git commit -m "feat: add Chatterbox voice studio"

Task 4: Deploy and verify the updated live service

Files:

  • Modify: none expected

Interfaces:

  • Consumes: existing chatterbox-tts.service, Caddy route, and local DNS record.

  • Produces: the live studio and voice APIs on https://chatterbox.dimensionlab.net.

  • Step 1: Restart only the Chatterbox user service

Run: systemctl --user restart chatterbox-tts.service && systemctl --user is-active chatterbox-tts.service

Expected: active.

  • Step 2: Verify live API and root page

Run:

curl -fsS https://chatterbox.dimensionlab.net/v1/voices
curl -fsS https://chatterbox.dimensionlab.net/ | rg 'Chatterbox Studio'
curl -fsS https://chatterbox.dimensionlab.net/health

Expected: voices JSON, studio HTML, and {"status":"ok"}.

  • Step 3: Push the reviewed commits

Run: git push origin main

Expected: the clean, verified implementation is published to vince/chatterbox-tts.