Launch guardrails, honest live-preview states, drop the dead script toggle
Browse filesCost caps for the public demo: scripts up to 7,000 words (duration menu
tops out at 45 min to match), plus a global budget of 60 generations per
rolling day on top of the per-IP limit, with a friendly quota message.
/api/status reports the remaining daily budget and the word cap.
The generation stage's preview row used to appear only once audio was
already scheduled and always said "playing", even when muted. It now shows
from the start with a note explaining that audio streams in once the
first stretch of dialogue renders and passes the gate, then switches to
playing/muted with the chunk count. The preview AudioContext is created
inside the Generate click so browsers don't start it suspended.
"View full script" on the player stage is removed along with its hidden
transcript pane; the dock transcript is the one that works.
The Modal runner (tracked) gains deploy-time scaling profiles
(VIBEVOICE_PROFILE=launch|tail) and cold/warm markers in container logs.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MARbUWrg2FB73kxtqenL2x
- README.md +6 -6
- app.py +32 -12
- backend_modal/modal_runner.py +30 -3
- static/app.js +25 -18
- static/index.html +2 -4
- static/styles.css +4 -2
|
@@ -30,7 +30,7 @@ _A 3-speaker example — Wizard, Orc, and Mom — generated from a single senten
|
|
| 30 |
|
| 31 |
**Writing**
|
| 32 |
- **Prompt-to-script** — describe the scenario ("a 4-person product meeting about pricing") and Qwen2.5-Coder-32B writes the full conversation, with a title
|
| 33 |
-
- **Target length** — pick 1 to
|
| 34 |
- **Bring your own script** — paste or upload text using `Speaker N:` tags or named characters
|
| 35 |
- **Turn editor** — reassign speakers or rewrite any line before rendering
|
| 36 |
- **Gender-aware casting** — characters get a matching voice automatically, with one-click override
|
|
@@ -86,9 +86,9 @@ The lightweight FastAPI frontend (this repo, hosted as a Docker Space) is separa
|
|
| 86 |
```
|
| 87 |
|
| 88 |
- **Frontend** (`app.py` + `static/`): script generation via the HF Inference API, script parsing, an SSE endpoint that relays progress and streamed chunk audio from Modal, post-processing (tone shelves, loudness normalization, spectral denoise, time-stretch), MP3 encoding, and byte-range serving of finished takes. Takes live in memory for 15 minutes.
|
| 89 |
-
- **Backend** (`backend_modal/modal_runner.py`
|
| 90 |
- **Scaling profiles**: the backend deploys with `VIBEVOICE_PROFILE=launch` (one container always warm, one buffer under load) or `tail` (scale to zero). Both cap at 4 concurrent GPUs; extra requests queue. The first generation after idle in `tail` mode takes ~3 minutes to load models, and the UI says so when the GPU is cold.
|
| 91 |
-
- **Limits**: 3 generations per hour per IP, 4 in flight globally.
|
| 92 |
|
| 93 |
---
|
| 94 |
|
|
@@ -127,7 +127,7 @@ VibeVoice is Microsoft's open-source long-form, multi-speaker TTS model. It uses
|
|
| 127 |
| Janus | M | Bright, Conversational | Original preset |
|
| 128 |
| Starchild | F | Airy, Dreamy | Original preset |
|
| 129 |
|
| 130 |
-
Preview clips live in `public/voices/`; the 60-second reference WAVs the backend conditions on
|
| 131 |
|
| 132 |
---
|
| 133 |
|
|
@@ -141,7 +141,7 @@ pip install -r requirements.txt
|
|
| 141 |
# Hugging Face token for the script-writing LLM (Inference API access)
|
| 142 |
export HF_TOKEN=your_hf_token_here
|
| 143 |
|
| 144 |
-
# Deploy the GPU backend separately (needs the gitignored backend_modal/
|
| 145 |
# VIBEVOICE_PROFILE=tail modal deploy backend_modal/modal_runner.py
|
| 146 |
|
| 147 |
python app.py # http://localhost:7860
|
|
@@ -172,7 +172,7 @@ Required env:
|
|
| 172 |
│ └── voices/ # Voice preview clips
|
| 173 |
├── text_examples/ # Example scripts (1–4 speakers)
|
| 174 |
├── tests/ # Script-parser tests + example prompts
|
| 175 |
-
└── backend_modal/ #
|
| 176 |
```
|
| 177 |
|
| 178 |
---
|
|
|
|
| 30 |
|
| 31 |
**Writing**
|
| 32 |
- **Prompt-to-script** — describe the scenario ("a 4-person product meeting about pricing") and Qwen2.5-Coder-32B writes the full conversation, with a title
|
| 33 |
+
- **Target length** — pick 1 to 45 minutes; the script is extended in continuation rounds until it reaches the word budget
|
| 34 |
- **Bring your own script** — paste or upload text using `Speaker N:` tags or named characters
|
| 35 |
- **Turn editor** — reassign speakers or rewrite any line before rendering
|
| 36 |
- **Gender-aware casting** — characters get a matching voice automatically, with one-click override
|
|
|
|
| 86 |
```
|
| 87 |
|
| 88 |
- **Frontend** (`app.py` + `static/`): script generation via the HF Inference API, script parsing, an SSE endpoint that relays progress and streamed chunk audio from Modal, post-processing (tone shelves, loudness normalization, spectral denoise, time-stretch), MP3 encoding, and byte-range serving of finished takes. Takes live in memory for 15 minutes.
|
| 89 |
+
- **Backend** (`backend_modal/modal_runner.py`): a Modal class that loads both models at container start and exposes `generate_podcast` as a streaming generator. Deployed separately. The VibeVoice model code and reference voice WAVs alongside it are gitignored.
|
| 90 |
- **Scaling profiles**: the backend deploys with `VIBEVOICE_PROFILE=launch` (one container always warm, one buffer under load) or `tail` (scale to zero). Both cap at 4 concurrent GPUs; extra requests queue. The first generation after idle in `tail` mode takes ~3 minutes to load models, and the UI says so when the GPU is cold.
|
| 91 |
+
- **Limits** (public demo): scripts up to 7,000 words (~45 min), 3 generations per hour per IP, 60 per day across all visitors, 4 in flight globally. The backend has no length ceiling; the multi-hour records were rendered by calling it directly.
|
| 92 |
|
| 93 |
---
|
| 94 |
|
|
|
|
| 127 |
| Janus | M | Bright, Conversational | Original preset |
|
| 128 |
| Starchild | F | Airy, Dreamy | Original preset |
|
| 129 |
|
| 130 |
+
Preview clips live in `public/voices/`; the 60-second reference WAVs the backend conditions on live under the gitignored `backend_modal/voices/`.
|
| 131 |
|
| 132 |
---
|
| 133 |
|
|
|
|
| 141 |
# Hugging Face token for the script-writing LLM (Inference API access)
|
| 142 |
export HF_TOKEN=your_hf_token_here
|
| 143 |
|
| 144 |
+
# Deploy the GPU backend separately (needs the gitignored model code + voices under backend_modal/)
|
| 145 |
# VIBEVOICE_PROFILE=tail modal deploy backend_modal/modal_runner.py
|
| 146 |
|
| 147 |
python app.py # http://localhost:7860
|
|
|
|
| 172 |
│ └── voices/ # Voice preview clips
|
| 173 |
├── text_examples/ # Example scripts (1–4 speakers)
|
| 174 |
├── tests/ # Script-parser tests + example prompts
|
| 175 |
+
└── backend_modal/ # Modal runner (tracked); VibeVoice model code + reference voices (gitignored)
|
| 176 |
```
|
| 177 |
|
| 178 |
---
|
|
@@ -58,13 +58,13 @@ DEFAULT_SPEAKERS = ["Cylinder", "Statesman", "Novella", "Eyre"]
|
|
| 58 |
|
| 59 |
SCRIPT_GEN_MODEL = "Qwen/Qwen2.5-Coder-32B-Instruct"
|
| 60 |
WORDS_PER_MINUTE = 150 # Matches the pace assumed by the client's duration estimate
|
| 61 |
-
DURATION_OPTIONS_MINUTES = [1, 2, 5, 10, 15, 20, 30, 45
|
| 62 |
MAX_COMPLETION_TOKENS = 8192 # Good-faith ceiling for a single chat_completion call; the
|
| 63 |
# underlying provider may cap lower, in which case the longest
|
| 64 |
# duration options may come back shorter than requested
|
| 65 |
-
MAX_SCRIPT_WORDS =
|
| 66 |
-
#
|
| 67 |
-
#
|
| 68 |
MAX_TURNS = 250 # Hard ceiling regardless of target length (safety valve)
|
| 69 |
MAX_GEN_ROUNDS = 24 # Hard cap on LLM calls per script regardless of target length
|
| 70 |
MIN_TARGET_FRACTION = 0.9 # Stop extending once the script reaches 90% of the word target
|
|
@@ -601,6 +601,9 @@ def _prune_audio_store() -> None:
|
|
| 601 |
_RATE_LOG: dict[str, deque] = defaultdict(deque)
|
| 602 |
SCRIPT_RATE_LIMIT = (5, 600) # 5 script generations per 10 min per IP (hits paid HF inference)
|
| 603 |
AUDIO_RATE_LIMIT = (3, 3600) # 3 audio generations per hour per IP (hits paid Modal GPU time)
|
|
|
|
|
|
|
|
|
|
| 604 |
GENERATION_CONCURRENCY = asyncio.Semaphore(4) # at most 4 Modal generations in flight at once, globally —
|
| 605 |
# matches the backend's max_containers cap (2026-09-06)
|
| 606 |
|
|
@@ -612,22 +615,31 @@ def _client_ip(request: Request) -> str:
|
|
| 612 |
return request.client.host if request.client else "unknown"
|
| 613 |
|
| 614 |
|
| 615 |
-
def _enforce_rate_limit(bucket: str, request: Request, limit: int, window_seconds: int
|
| 616 |
-
|
|
|
|
|
|
|
|
|
|
| 617 |
now = time.time()
|
| 618 |
log = _RATE_LOG[key]
|
| 619 |
while log and now - log[0] > window_seconds:
|
| 620 |
log.popleft()
|
| 621 |
if len(log) >= limit:
|
| 622 |
retry_after = max(1, int(window_seconds - (now - log[0])))
|
| 623 |
-
|
| 624 |
-
|
| 625 |
-
|
| 626 |
-
headers={"Retry-After": str(retry_after)},
|
| 627 |
-
)
|
| 628 |
log.append(now)
|
| 629 |
|
| 630 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 631 |
# ========================================================
|
| 632 |
# FASTAPI APP
|
| 633 |
# ========================================================
|
|
@@ -716,7 +728,11 @@ async def _backend_stats() -> dict | None:
|
|
| 716 |
async def api_status() -> dict:
|
| 717 |
"""Backend reachability plus, when Modal reports it, whether a GPU container
|
| 718 |
is already hot — the UI uses that to warn about the ~3 min cold path."""
|
| 719 |
-
payload = {
|
|
|
|
|
|
|
|
|
|
|
|
|
| 720 |
stats = await _backend_stats()
|
| 721 |
if stats is not None:
|
| 722 |
payload.update(stats)
|
|
@@ -856,6 +872,10 @@ class GenerateRequest(BaseModel):
|
|
| 856 |
@app.post("/api/generate")
|
| 857 |
async def api_generate(payload: GenerateRequest, request: Request) -> StreamingResponse:
|
| 858 |
_enforce_rate_limit("audio", request, *AUDIO_RATE_LIMIT)
|
|
|
|
|
|
|
|
|
|
|
|
|
| 859 |
if remote_generate_function is None:
|
| 860 |
raise HTTPException(status_code=503, detail="Modal backend is offline.")
|
| 861 |
|
|
|
|
| 58 |
|
| 59 |
SCRIPT_GEN_MODEL = "Qwen/Qwen2.5-Coder-32B-Instruct"
|
| 60 |
WORDS_PER_MINUTE = 150 # Matches the pace assumed by the client's duration estimate
|
| 61 |
+
DURATION_OPTIONS_MINUTES = [1, 2, 5, 10, 15, 20, 30, 45] # 45 × 150 wpm fits MAX_SCRIPT_WORDS
|
| 62 |
MAX_COMPLETION_TOKENS = 8192 # Good-faith ceiling for a single chat_completion call; the
|
| 63 |
# underlying provider may cap lower, in which case the longest
|
| 64 |
# duration options may come back shorter than requested
|
| 65 |
+
MAX_SCRIPT_WORDS = 7000 # Public cap (2026-09-07, launch): ~45 min of audio, ~8 min of A100
|
| 66 |
+
# time per job. The backend itself has no ceiling — record-length
|
| 67 |
+
# renders call Modal directly and bypass this guard.
|
| 68 |
MAX_TURNS = 250 # Hard ceiling regardless of target length (safety valve)
|
| 69 |
MAX_GEN_ROUNDS = 24 # Hard cap on LLM calls per script regardless of target length
|
| 70 |
MIN_TARGET_FRACTION = 0.9 # Stop extending once the script reaches 90% of the word target
|
|
|
|
| 601 |
_RATE_LOG: dict[str, deque] = defaultdict(deque)
|
| 602 |
SCRIPT_RATE_LIMIT = (5, 600) # 5 script generations per 10 min per IP (hits paid HF inference)
|
| 603 |
AUDIO_RATE_LIMIT = (3, 3600) # 3 audio generations per hour per IP (hits paid Modal GPU time)
|
| 604 |
+
AUDIO_DAILY_BUDGET = (60, 86400) # 60 audio generations per rolling day across ALL visitors — bounds the
|
| 605 |
+
# worst-case GPU bill regardless of how many IPs show up (2026-09-07)
|
| 606 |
+
GLOBAL_KEY = "__global__"
|
| 607 |
GENERATION_CONCURRENCY = asyncio.Semaphore(4) # at most 4 Modal generations in flight at once, globally —
|
| 608 |
# matches the backend's max_containers cap (2026-09-06)
|
| 609 |
|
|
|
|
| 615 |
return request.client.host if request.client else "unknown"
|
| 616 |
|
| 617 |
|
| 618 |
+
def _enforce_rate_limit(bucket: str, request: Request | None, limit: int, window_seconds: int,
|
| 619 |
+
message: str | None = None) -> None:
|
| 620 |
+
"""Sliding-window counter. request=None counts globally (every visitor shares the bucket)."""
|
| 621 |
+
who = _client_ip(request) if request is not None else GLOBAL_KEY
|
| 622 |
+
key = f"{bucket}:{who}"
|
| 623 |
now = time.time()
|
| 624 |
log = _RATE_LOG[key]
|
| 625 |
while log and now - log[0] > window_seconds:
|
| 626 |
log.popleft()
|
| 627 |
if len(log) >= limit:
|
| 628 |
retry_after = max(1, int(window_seconds - (now - log[0])))
|
| 629 |
+
if message is None:
|
| 630 |
+
message = f"Rate limit reached ({limit} per {window_seconds // 60} min). Try again in about {retry_after}s."
|
| 631 |
+
raise HTTPException(status_code=429, detail=message, headers={"Retry-After": str(retry_after)})
|
|
|
|
|
|
|
| 632 |
log.append(now)
|
| 633 |
|
| 634 |
|
| 635 |
+
def _audio_budget_remaining() -> int:
|
| 636 |
+
now = time.time()
|
| 637 |
+
log = _RATE_LOG[f"audio-day:{GLOBAL_KEY}"]
|
| 638 |
+
while log and now - log[0] > AUDIO_DAILY_BUDGET[1]:
|
| 639 |
+
log.popleft()
|
| 640 |
+
return max(0, AUDIO_DAILY_BUDGET[0] - len(log))
|
| 641 |
+
|
| 642 |
+
|
| 643 |
# ========================================================
|
| 644 |
# FASTAPI APP
|
| 645 |
# ========================================================
|
|
|
|
| 728 |
async def api_status() -> dict:
|
| 729 |
"""Backend reachability plus, when Modal reports it, whether a GPU container
|
| 730 |
is already hot — the UI uses that to warn about the ~3 min cold path."""
|
| 731 |
+
payload = {
|
| 732 |
+
"backend": "ready" if remote_generate_function is not None else "offline",
|
| 733 |
+
"daily_remaining": _audio_budget_remaining(),
|
| 734 |
+
"max_script_words": MAX_SCRIPT_WORDS,
|
| 735 |
+
}
|
| 736 |
stats = await _backend_stats()
|
| 737 |
if stats is not None:
|
| 738 |
payload.update(stats)
|
|
|
|
| 872 |
@app.post("/api/generate")
|
| 873 |
async def api_generate(payload: GenerateRequest, request: Request) -> StreamingResponse:
|
| 874 |
_enforce_rate_limit("audio", request, *AUDIO_RATE_LIMIT)
|
| 875 |
+
_enforce_rate_limit(
|
| 876 |
+
"audio-day", None, *AUDIO_DAILY_BUDGET,
|
| 877 |
+
message="Today's free demo quota is used up. It refills over the next 24 hours — please come back later.",
|
| 878 |
+
)
|
| 879 |
if remote_generate_function is None:
|
| 880 |
raise HTTPException(status_code=503, detail="Modal backend is offline.")
|
| 881 |
|
|
@@ -14,6 +14,23 @@ import pickle
|
|
| 14 |
# Modal-specific imports
|
| 15 |
import modal
|
| 16 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 17 |
# Define the Modal Stub
|
| 18 |
image = (
|
| 19 |
modal.Image.debian_slim(python_version="3.10")
|
|
@@ -40,6 +57,7 @@ image = (
|
|
| 40 |
"ln -s /root/schedule /root/vibevoice/schedule"
|
| 41 |
)
|
| 42 |
.env({"PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True"}) # fights fragmentation across waves
|
|
|
|
| 43 |
.add_local_dir("backend_modal/modular", remote_path="/root/modular")
|
| 44 |
.add_local_dir("backend_modal/processor", remote_path="/root/processor")
|
| 45 |
.add_local_dir("backend_modal/voices", remote_path="/root/voices")
|
|
@@ -65,12 +83,12 @@ cache_volume = modal.Volume.from_name("vibevoice-cache", create_if_missing=True)
|
|
| 65 |
|
| 66 |
@app.cls(
|
| 67 |
gpu="A100-40GB",
|
| 68 |
-
scaledown_window=300,
|
| 69 |
timeout=7200, # was 3600: a 120-min+ record attempt hit the 1h ceiling at the
|
| 70 |
# finish line (2026-08-19), losing the whole render. Long-form
|
| 71 |
# with cloned voices runs slower than the preset-voice record
|
| 72 |
# pace (bigger reference prefill per chunk), so give 2h.
|
| 73 |
-
volumes={"/cache": cache_volume}
|
|
|
|
| 74 |
)
|
| 75 |
class VibeVoiceModel:
|
| 76 |
@modal.enter()
|
|
@@ -89,7 +107,9 @@ class VibeVoiceModel:
|
|
| 89 |
from modular.modeling_vibevoice_inference import VibeVoiceForConditionalGenerationInference
|
| 90 |
from processor.vibevoice_processor import VibeVoiceProcessor
|
| 91 |
|
| 92 |
-
|
|
|
|
|
|
|
| 93 |
|
| 94 |
# Set compiler flags for better performance
|
| 95 |
if torch.cuda.is_available() and hasattr(torch, '_inductor'):
|
|
@@ -122,6 +142,8 @@ class VibeVoiceModel:
|
|
| 122 |
|
| 123 |
self.setup_voice_presets()
|
| 124 |
self.ready_at = time.time() # for cold-start detection in timing reports
|
|
|
|
|
|
|
| 125 |
# VRAM baseline right after model load: requests arriving to a GPU far
|
| 126 |
# above this are hitting a poisoned container (e.g. a cancelled run's
|
| 127 |
# zombie generation thread) and must recycle, not proceed (2026-08-14)
|
|
@@ -816,6 +838,11 @@ class VibeVoiceModel:
|
|
| 816 |
falls back to the named preset in speaker_N.
|
| 817 |
"""
|
| 818 |
try:
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 819 |
if model_name not in self.models:
|
| 820 |
raise ValueError(f"Unknown model: {model_name}")
|
| 821 |
|
|
|
|
| 14 |
# Modal-specific imports
|
| 15 |
import modal
|
| 16 |
|
| 17 |
+
# --- Scaling profiles (2026-09-06) ---------------------------------------
|
| 18 |
+
# Chosen at deploy time: VIBEVOICE_PROFILE=launch modal deploy backend_modal/modal_runner.py
|
| 19 |
+
# "launch": one container always hot + one pre-warmed buffer under load, so a
|
| 20 |
+
# public-post burst lands warm instead of paying the ~3 min model load.
|
| 21 |
+
# "tail": scale to zero when idle (default). Same cap and idle window, so the
|
| 22 |
+
# only thing that changes is the standing spend.
|
| 23 |
+
# max_containers is the spending cap in both: a 5th concurrent request queues
|
| 24 |
+
# rather than spawning a 5th A100.
|
| 25 |
+
SCALING_PROFILES = {
|
| 26 |
+
"launch": dict(min_containers=1, buffer_containers=1, max_containers=4, scaledown_window=1200),
|
| 27 |
+
"tail": dict(min_containers=0, buffer_containers=0, max_containers=4, scaledown_window=1200),
|
| 28 |
+
}
|
| 29 |
+
SCALING_PROFILE = os.environ.get("VIBEVOICE_PROFILE", "tail").strip().lower()
|
| 30 |
+
if SCALING_PROFILE not in SCALING_PROFILES:
|
| 31 |
+
raise SystemExit(f"VIBEVOICE_PROFILE must be one of {sorted(SCALING_PROFILES)}, got {SCALING_PROFILE!r}")
|
| 32 |
+
print(f"Deploying with scaling profile '{SCALING_PROFILE}': {SCALING_PROFILES[SCALING_PROFILE]}")
|
| 33 |
+
|
| 34 |
# Define the Modal Stub
|
| 35 |
image = (
|
| 36 |
modal.Image.debian_slim(python_version="3.10")
|
|
|
|
| 57 |
"ln -s /root/schedule /root/vibevoice/schedule"
|
| 58 |
)
|
| 59 |
.env({"PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True"}) # fights fragmentation across waves
|
| 60 |
+
.env({"VIBEVOICE_PROFILE": SCALING_PROFILE}) # so container logs name the profile they run under
|
| 61 |
.add_local_dir("backend_modal/modular", remote_path="/root/modular")
|
| 62 |
.add_local_dir("backend_modal/processor", remote_path="/root/processor")
|
| 63 |
.add_local_dir("backend_modal/voices", remote_path="/root/voices")
|
|
|
|
| 83 |
|
| 84 |
@app.cls(
|
| 85 |
gpu="A100-40GB",
|
|
|
|
| 86 |
timeout=7200, # was 3600: a 120-min+ record attempt hit the 1h ceiling at the
|
| 87 |
# finish line (2026-08-19), losing the whole render. Long-form
|
| 88 |
# with cloned voices runs slower than the preset-voice record
|
| 89 |
# pace (bigger reference prefill per chunk), so give 2h.
|
| 90 |
+
volumes={"/cache": cache_volume},
|
| 91 |
+
**SCALING_PROFILES[SCALING_PROFILE],
|
| 92 |
)
|
| 93 |
class VibeVoiceModel:
|
| 94 |
@modal.enter()
|
|
|
|
| 107 |
from modular.modeling_vibevoice_inference import VibeVoiceForConditionalGenerationInference
|
| 108 |
from processor.vibevoice_processor import VibeVoiceProcessor
|
| 109 |
|
| 110 |
+
self.boot_started = time.time()
|
| 111 |
+
print(f"CONTAINER_START profile={SCALING_PROFILE} at={datetime.utcnow().isoformat()}Z "
|
| 112 |
+
"— loading models to GPU...")
|
| 113 |
|
| 114 |
# Set compiler flags for better performance
|
| 115 |
if torch.cuda.is_available() and hasattr(torch, '_inductor'):
|
|
|
|
| 142 |
|
| 143 |
self.setup_voice_presets()
|
| 144 |
self.ready_at = time.time() # for cold-start detection in timing reports
|
| 145 |
+
self.jobs_served = 0
|
| 146 |
+
print(f"CONTAINER_READY load_seconds={self.ready_at - self.boot_started:.0f}")
|
| 147 |
# VRAM baseline right after model load: requests arriving to a GPU far
|
| 148 |
# above this are hitting a poisoned container (e.g. a cancelled run's
|
| 149 |
# zombie generation thread) and must recycle, not proceed (2026-08-14)
|
|
|
|
| 838 |
falls back to the named preset in speaker_N.
|
| 839 |
"""
|
| 840 |
try:
|
| 841 |
+
self.jobs_served = getattr(self, "jobs_served", 0) + 1
|
| 842 |
+
container_age = time.time() - getattr(self, "ready_at", time.time())
|
| 843 |
+
print(f"JOB_START cold={container_age < 120} container_age_s={container_age:.0f} "
|
| 844 |
+
f"job_n={self.jobs_served} model={model_name} words={len(script.split())} "
|
| 845 |
+
f"at={datetime.utcnow().isoformat()}Z")
|
| 846 |
if model_name not in self.models:
|
| 847 |
raise ValueError(f"Unknown model: {model_name}")
|
| 848 |
|
|
@@ -83,7 +83,7 @@ const el = {};
|
|
| 83 |
"composerCollapsedStrip", "collapsedSummary", "composerBody",
|
| 84 |
"playerStage", "stageTitle", "stagePlayBtn", "stageWaveform", "stageTime",
|
| 85 |
"stageDot", "stageLine", "stageSpeaker", "stageCloseBtn", "stageDownloadBtn", "stageOrbs",
|
| 86 |
-
"
|
| 87 |
"generationTime", "audioDuration", "resultModel", "downloadBtn",
|
| 88 |
"realtimeRow", "realtimeFactor", "warmupRow", "warmupTime",
|
| 89 |
"downloadMp3Btn", "stageDownloadMp3Btn",
|
|
@@ -1341,9 +1341,8 @@ function buildSyncedTranscript(snapshot) {
|
|
| 1341 |
lastCaptionKey = "";
|
| 1342 |
|
| 1343 |
el.syncedTranscript.innerHTML = "";
|
| 1344 |
-
el.stageTranscript.innerHTML = "";
|
| 1345 |
state.resultTurns.forEach((turn, i) => {
|
| 1346 |
-
state.resultTurns[i].rows = [el.syncedTranscript
|
| 1347 |
const row = document.createElement("div");
|
| 1348 |
row.className = "sync-line";
|
| 1349 |
const dot = document.createElement("span");
|
|
@@ -1866,15 +1865,6 @@ el.openPlayerBtn.addEventListener("click", openPlayerStage);
|
|
| 1866 |
el.stageCloseBtn.addEventListener("click", () => el.playerStage.close());
|
| 1867 |
el.playerStage.addEventListener("click", (e) => { if (e.target === el.playerStage) el.playerStage.close(); });
|
| 1868 |
|
| 1869 |
-
el.stageScriptToggle.addEventListener("click", () => {
|
| 1870 |
-
el.stageTranscript.hidden = !el.stageTranscript.hidden;
|
| 1871 |
-
el.stageScriptToggle.textContent = el.stageTranscript.hidden ? "View full script" : "Hide script";
|
| 1872 |
-
if (!el.stageTranscript.hidden) {
|
| 1873 |
-
const active = state.resultTurns[state.activeSyncIndex];
|
| 1874 |
-
const target = (active && active.rows && active.rows[1]) || el.stageTranscript;
|
| 1875 |
-
target.scrollIntoView({ block: "nearest", behavior: "smooth" });
|
| 1876 |
-
}
|
| 1877 |
-
});
|
| 1878 |
|
| 1879 |
document.addEventListener("keydown", (e) => {
|
| 1880 |
if (!el.playerStage.open || e.code !== "Space") return;
|
|
@@ -2303,12 +2293,13 @@ function resetPreview() {
|
|
| 2303 |
preview.nextTakeTime = 0;
|
| 2304 |
preview.normGain = 1;
|
| 2305 |
preview.done = false;
|
| 2306 |
-
|
| 2307 |
}
|
| 2308 |
|
| 2309 |
function stopPreview(fadeSecs = 0.3) {
|
| 2310 |
preview.done = true;
|
| 2311 |
preview.active = false;
|
|
|
|
| 2312 |
if (!preview.ctx) return;
|
| 2313 |
const ctx = preview.ctx;
|
| 2314 |
preview.ctx = null;
|
|
@@ -2439,25 +2430,40 @@ function previewPositionSeconds() {
|
|
| 2439 |
return pos;
|
| 2440 |
}
|
| 2441 |
|
|
|
|
|
|
|
|
|
|
| 2442 |
function updatePreviewUI() {
|
| 2443 |
-
if (preview.done
|
| 2444 |
el.genPreviewRow.hidden = false;
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 2445 |
const n = Math.min(preview.nextIndex, preview.total || preview.nextIndex);
|
|
|
|
| 2446 |
if (preview.ctx && preview.ctx.state === "suspended") {
|
| 2447 |
// Browser blocked audio start without a fresh gesture.
|
| 2448 |
el.genPreviewLabel.textContent = "Preview ready — click Unmute to listen while it renders";
|
|
|
|
| 2449 |
return;
|
| 2450 |
}
|
| 2451 |
-
el.genPreviewLabel.textContent = preview.
|
| 2452 |
-
? `Live preview playing · ${n} of ${preview.total} chunks rendered`
|
| 2453 |
-
: "Live preview playing";
|
| 2454 |
}
|
| 2455 |
|
| 2456 |
el.genPreviewMute.addEventListener("click", () => {
|
| 2457 |
preview.muted = !preview.muted;
|
| 2458 |
el.genPreviewMute.textContent = preview.muted ? "Unmute" : "Mute";
|
| 2459 |
if (preview.master) preview.master.gain.value = preview.muted ? 0 : preview.normGain;
|
| 2460 |
-
if (preview.ctx && preview.ctx.state === "suspended")
|
|
|
|
|
|
|
|
|
|
| 2461 |
});
|
| 2462 |
|
| 2463 |
/* ---------------- Generate ---------------- */
|
|
@@ -2514,6 +2520,7 @@ el.generateBtn.addEventListener("click", async () => {
|
|
| 2514 |
el.logToggleBtn.hidden = true;
|
| 2515 |
el.logToggleBtn.textContent = "View generation log";
|
| 2516 |
startProgress();
|
|
|
|
| 2517 |
// The render is already submitted — edits here can't reach it, so lock the
|
| 2518 |
// workspace rather than let controls silently no-op, and put progress
|
| 2519 |
// front and centre.
|
|
|
|
| 83 |
"composerCollapsedStrip", "collapsedSummary", "composerBody",
|
| 84 |
"playerStage", "stageTitle", "stagePlayBtn", "stageWaveform", "stageTime",
|
| 85 |
"stageDot", "stageLine", "stageSpeaker", "stageCloseBtn", "stageDownloadBtn", "stageOrbs",
|
| 86 |
+
"soundSeg", "soundHint", "polishStatus", "cleanNoiseCheckbox",
|
| 87 |
"generationTime", "audioDuration", "resultModel", "downloadBtn",
|
| 88 |
"realtimeRow", "realtimeFactor", "warmupRow", "warmupTime",
|
| 89 |
"downloadMp3Btn", "stageDownloadMp3Btn",
|
|
|
|
| 1341 |
lastCaptionKey = "";
|
| 1342 |
|
| 1343 |
el.syncedTranscript.innerHTML = "";
|
|
|
|
| 1344 |
state.resultTurns.forEach((turn, i) => {
|
| 1345 |
+
state.resultTurns[i].rows = [el.syncedTranscript].map((container) => {
|
| 1346 |
const row = document.createElement("div");
|
| 1347 |
row.className = "sync-line";
|
| 1348 |
const dot = document.createElement("span");
|
|
|
|
| 1865 |
el.stageCloseBtn.addEventListener("click", () => el.playerStage.close());
|
| 1866 |
el.playerStage.addEventListener("click", (e) => { if (e.target === el.playerStage) el.playerStage.close(); });
|
| 1867 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1868 |
|
| 1869 |
document.addEventListener("keydown", (e) => {
|
| 1870 |
if (!el.playerStage.open || e.code !== "Space") return;
|
|
|
|
| 2293 |
preview.nextTakeTime = 0;
|
| 2294 |
preview.normGain = 1;
|
| 2295 |
preview.done = false;
|
| 2296 |
+
updatePreviewUI();
|
| 2297 |
}
|
| 2298 |
|
| 2299 |
function stopPreview(fadeSecs = 0.3) {
|
| 2300 |
preview.done = true;
|
| 2301 |
preview.active = false;
|
| 2302 |
+
el.genPreviewRow.hidden = true;
|
| 2303 |
if (!preview.ctx) return;
|
| 2304 |
const ctx = preview.ctx;
|
| 2305 |
preview.ctx = null;
|
|
|
|
| 2430 |
return pos;
|
| 2431 |
}
|
| 2432 |
|
| 2433 |
+
const PREVIEW_WAITING_NOTE =
|
| 2434 |
+
"Audio starts streaming here as soon as the first stretch of dialogue is rendered and passes the quality check — usually a minute or two in.";
|
| 2435 |
+
|
| 2436 |
function updatePreviewUI() {
|
| 2437 |
+
if (preview.done) return;
|
| 2438 |
el.genPreviewRow.hidden = false;
|
| 2439 |
+
if (!preview.active) {
|
| 2440 |
+
// Nothing is audible yet: say what will happen rather than claim playback.
|
| 2441 |
+
el.genPreviewRow.classList.add("waiting");
|
| 2442 |
+
el.genPreviewMute.hidden = true;
|
| 2443 |
+
el.genPreviewLabel.textContent = PREVIEW_WAITING_NOTE;
|
| 2444 |
+
return;
|
| 2445 |
+
}
|
| 2446 |
+
el.genPreviewRow.classList.remove("waiting");
|
| 2447 |
+
el.genPreviewMute.hidden = false;
|
| 2448 |
const n = Math.min(preview.nextIndex, preview.total || preview.nextIndex);
|
| 2449 |
+
const progressNote = preview.total ? ` · ${n} of ${preview.total} chunks rendered` : "";
|
| 2450 |
if (preview.ctx && preview.ctx.state === "suspended") {
|
| 2451 |
// Browser blocked audio start without a fresh gesture.
|
| 2452 |
el.genPreviewLabel.textContent = "Preview ready — click Unmute to listen while it renders";
|
| 2453 |
+
el.genPreviewMute.textContent = "Unmute";
|
| 2454 |
return;
|
| 2455 |
}
|
| 2456 |
+
el.genPreviewLabel.textContent = (preview.muted ? "Live preview muted" : "Live preview playing") + progressNote;
|
|
|
|
|
|
|
| 2457 |
}
|
| 2458 |
|
| 2459 |
el.genPreviewMute.addEventListener("click", () => {
|
| 2460 |
preview.muted = !preview.muted;
|
| 2461 |
el.genPreviewMute.textContent = preview.muted ? "Unmute" : "Mute";
|
| 2462 |
if (preview.master) preview.master.gain.value = preview.muted ? 0 : preview.normGain;
|
| 2463 |
+
if (preview.ctx && preview.ctx.state === "suspended") {
|
| 2464 |
+
preview.ctx.resume().then(updatePreviewUI).catch(() => {});
|
| 2465 |
+
}
|
| 2466 |
+
updatePreviewUI();
|
| 2467 |
});
|
| 2468 |
|
| 2469 |
/* ---------------- Generate ---------------- */
|
|
|
|
| 2520 |
el.logToggleBtn.hidden = true;
|
| 2521 |
el.logToggleBtn.textContent = "View generation log";
|
| 2522 |
startProgress();
|
| 2523 |
+
ensurePreviewCtx(); // inside the click gesture: an AudioContext made here starts unblocked
|
| 2524 |
// The render is already submitted — edits here can't reach it, so lock the
|
| 2525 |
// workspace rather than let controls silently no-op, and put progress
|
| 2526 |
// front and centre.
|
|
@@ -203,10 +203,10 @@
|
|
| 203 |
<div class="progress-fill" id="genStageFill"></div>
|
| 204 |
</div>
|
| 205 |
<div class="gen-stage-meta" id="genStageMeta"></div>
|
| 206 |
-
<div class="gen-preview" id="genPreviewRow" hidden>
|
| 207 |
<span class="gen-preview-dot"></span>
|
| 208 |
<span id="genPreviewLabel">Live preview</span>
|
| 209 |
-
<button class="gen-preview-mute" id="genPreviewMute" type="button">Mute</button>
|
| 210 |
</div>
|
| 211 |
<div class="gen-stage-desc" id="genStageDesc"></div>
|
| 212 |
<div class="gen-stage-log" id="genStageLog"></div>
|
|
@@ -248,9 +248,7 @@
|
|
| 248 |
<a class="btn btn-ink" id="stageDownloadBtn" download="chorus-audio.wav">Download WAV</a>
|
| 249 |
<a class="btn btn-pill-outline" id="stageDownloadMp3Btn" download="chorus-audio.mp3"
|
| 250 |
title="Much smaller file — encoded on the server on first click">MP3</a>
|
| 251 |
-
<button class="btn btn-pill-outline" id="stageScriptToggle" type="button">View full script</button>
|
| 252 |
</div>
|
| 253 |
-
<div class="synced-transcript stage-transcript" id="stageTranscript" hidden></div>
|
| 254 |
</div>
|
| 255 |
</div>
|
| 256 |
</dialog>
|
|
|
|
| 203 |
<div class="progress-fill" id="genStageFill"></div>
|
| 204 |
</div>
|
| 205 |
<div class="gen-stage-meta" id="genStageMeta"></div>
|
| 206 |
+
<div class="gen-preview waiting" id="genPreviewRow" hidden>
|
| 207 |
<span class="gen-preview-dot"></span>
|
| 208 |
<span id="genPreviewLabel">Live preview</span>
|
| 209 |
+
<button class="gen-preview-mute" id="genPreviewMute" type="button" hidden>Mute</button>
|
| 210 |
</div>
|
| 211 |
<div class="gen-stage-desc" id="genStageDesc"></div>
|
| 212 |
<div class="gen-stage-log" id="genStageLog"></div>
|
|
|
|
| 248 |
<a class="btn btn-ink" id="stageDownloadBtn" download="chorus-audio.wav">Download WAV</a>
|
| 249 |
<a class="btn btn-pill-outline" id="stageDownloadMp3Btn" download="chorus-audio.mp3"
|
| 250 |
title="Much smaller file — encoded on the server on first click">MP3</a>
|
|
|
|
| 251 |
</div>
|
|
|
|
| 252 |
</div>
|
| 253 |
</div>
|
| 254 |
</dialog>
|
|
@@ -980,7 +980,11 @@ body.is-generating .canvas { opacity: 0.6; transition: opacity 0.2s; }
|
|
| 980 |
border-radius: 50%;
|
| 981 |
background: var(--accent);
|
| 982 |
animation: pulse 1.2s ease-in-out infinite;
|
|
|
|
| 983 |
}
|
|
|
|
|
|
|
|
|
|
| 984 |
.gen-preview-mute {
|
| 985 |
border: 1px solid var(--border);
|
| 986 |
background: transparent;
|
|
@@ -1077,8 +1081,6 @@ body.is-generating .canvas { opacity: 0.6; transition: opacity 0.2s; }
|
|
| 1077 |
text-align: center;
|
| 1078 |
}
|
| 1079 |
|
| 1080 |
-
.stage-transcript { margin: 20px 0 0; max-height: 260px; }
|
| 1081 |
-
.stage-transcript .sync-text { font-size: 0.88rem; }
|
| 1082 |
|
| 1083 |
.icon-btn {
|
| 1084 |
width: 30px; height: 30px;
|
|
|
|
| 980 |
border-radius: 50%;
|
| 981 |
background: var(--accent);
|
| 982 |
animation: pulse 1.2s ease-in-out infinite;
|
| 983 |
+
flex: none;
|
| 984 |
}
|
| 985 |
+
/* Before the first chunk lands: a quiet note, not a "playing" claim. */
|
| 986 |
+
.gen-preview.waiting { font-weight: 500; flex-wrap: wrap; text-align: center; max-width: 440px; margin-left: auto; margin-right: auto; }
|
| 987 |
+
.gen-preview.waiting .gen-preview-dot { animation: none; background: var(--border-strong); }
|
| 988 |
.gen-preview-mute {
|
| 989 |
border: 1px solid var(--border);
|
| 990 |
background: transparent;
|
|
|
|
| 1081 |
text-align: center;
|
| 1082 |
}
|
| 1083 |
|
|
|
|
|
|
|
| 1084 |
|
| 1085 |
.icon-btn {
|
| 1086 |
width: 30px; height: 30px;
|