ACloudCenter Claude Fable 5.1 commited on
Commit
797f6c7
·
1 Parent(s): e3b6645

Launch guardrails, honest live-preview states, drop the dead script toggle

Browse files

Cost caps for the public demo: scripts up to 7,000 words (duration menu
tops out at 45 min to match), plus a global budget of 60 generations per
rolling day on top of the per-IP limit, with a friendly quota message.
/api/status reports the remaining daily budget and the word cap.

The generation stage's preview row used to appear only once audio was
already scheduled and always said "playing", even when muted. It now shows
from the start with a note explaining that audio streams in once the
first stretch of dialogue renders and passes the gate, then switches to
playing/muted with the chunk count. The preview AudioContext is created
inside the Generate click so browsers don't start it suspended.

"View full script" on the player stage is removed along with its hidden
transcript pane; the dock transcript is the one that works.

The Modal runner (tracked) gains deploy-time scaling profiles
(VIBEVOICE_PROFILE=launch|tail) and cold/warm markers in container logs.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MARbUWrg2FB73kxtqenL2x

Files changed (6) hide show
  1. README.md +6 -6
  2. app.py +32 -12
  3. backend_modal/modal_runner.py +30 -3
  4. static/app.js +25 -18
  5. static/index.html +2 -4
  6. static/styles.css +4 -2
README.md CHANGED
@@ -30,7 +30,7 @@ _A 3-speaker example — Wizard, Orc, and Mom — generated from a single senten
30
 
31
  **Writing**
32
  - **Prompt-to-script** — describe the scenario ("a 4-person product meeting about pricing") and Qwen2.5-Coder-32B writes the full conversation, with a title
33
- - **Target length** — pick 1 to 60 minutes; the script is extended in continuation rounds until it reaches the word budget
34
  - **Bring your own script** — paste or upload text using `Speaker N:` tags or named characters
35
  - **Turn editor** — reassign speakers or rewrite any line before rendering
36
  - **Gender-aware casting** — characters get a matching voice automatically, with one-click override
@@ -86,9 +86,9 @@ The lightweight FastAPI frontend (this repo, hosted as a Docker Space) is separa
86
  ```
87
 
88
  - **Frontend** (`app.py` + `static/`): script generation via the HF Inference API, script parsing, an SSE endpoint that relays progress and streamed chunk audio from Modal, post-processing (tone shelves, loudness normalization, spectral denoise, time-stretch), MP3 encoding, and byte-range serving of finished takes. Takes live in memory for 15 minutes.
89
- - **Backend** (`backend_modal/modal_runner.py`, gitignored): a Modal class that loads both models at container start and exposes `generate_podcast` as a streaming generator. Deployed separately.
90
  - **Scaling profiles**: the backend deploys with `VIBEVOICE_PROFILE=launch` (one container always warm, one buffer under load) or `tail` (scale to zero). Both cap at 4 concurrent GPUs; extra requests queue. The first generation after idle in `tail` mode takes ~3 minutes to load models, and the UI says so when the GPU is cold.
91
- - **Limits**: 3 generations per hour per IP, 4 in flight globally.
92
 
93
  ---
94
 
@@ -127,7 +127,7 @@ VibeVoice is Microsoft's open-source long-form, multi-speaker TTS model. It uses
127
  | Janus | M | Bright, Conversational | Original preset |
128
  | Starchild | F | Airy, Dreamy | Original preset |
129
 
130
- Preview clips live in `public/voices/`; the 60-second reference WAVs the backend conditions on are under the gitignored `backend_modal/voices/`.
131
 
132
  ---
133
 
@@ -141,7 +141,7 @@ pip install -r requirements.txt
141
  # Hugging Face token for the script-writing LLM (Inference API access)
142
  export HF_TOKEN=your_hf_token_here
143
 
144
- # Deploy the GPU backend separately (needs the gitignored backend_modal/ tree)
145
  # VIBEVOICE_PROFILE=tail modal deploy backend_modal/modal_runner.py
146
 
147
  python app.py # http://localhost:7860
@@ -172,7 +172,7 @@ Required env:
172
  │ └── voices/ # Voice preview clips
173
  ├── text_examples/ # Example scripts (1–4 speakers)
174
  ├── tests/ # Script-parser tests + example prompts
175
- └── backend_modal/ # (gitignored) Modal runner, VibeVoice model code, reference voices
176
  ```
177
 
178
  ---
 
30
 
31
  **Writing**
32
  - **Prompt-to-script** — describe the scenario ("a 4-person product meeting about pricing") and Qwen2.5-Coder-32B writes the full conversation, with a title
33
+ - **Target length** — pick 1 to 45 minutes; the script is extended in continuation rounds until it reaches the word budget
34
  - **Bring your own script** — paste or upload text using `Speaker N:` tags or named characters
35
  - **Turn editor** — reassign speakers or rewrite any line before rendering
36
  - **Gender-aware casting** — characters get a matching voice automatically, with one-click override
 
86
  ```
87
 
88
  - **Frontend** (`app.py` + `static/`): script generation via the HF Inference API, script parsing, an SSE endpoint that relays progress and streamed chunk audio from Modal, post-processing (tone shelves, loudness normalization, spectral denoise, time-stretch), MP3 encoding, and byte-range serving of finished takes. Takes live in memory for 15 minutes.
89
+ - **Backend** (`backend_modal/modal_runner.py`): a Modal class that loads both models at container start and exposes `generate_podcast` as a streaming generator. Deployed separately. The VibeVoice model code and reference voice WAVs alongside it are gitignored.
90
  - **Scaling profiles**: the backend deploys with `VIBEVOICE_PROFILE=launch` (one container always warm, one buffer under load) or `tail` (scale to zero). Both cap at 4 concurrent GPUs; extra requests queue. The first generation after idle in `tail` mode takes ~3 minutes to load models, and the UI says so when the GPU is cold.
91
+ - **Limits** (public demo): scripts up to 7,000 words (~45 min), 3 generations per hour per IP, 60 per day across all visitors, 4 in flight globally. The backend has no length ceiling; the multi-hour records were rendered by calling it directly.
92
 
93
  ---
94
 
 
127
  | Janus | M | Bright, Conversational | Original preset |
128
  | Starchild | F | Airy, Dreamy | Original preset |
129
 
130
+ Preview clips live in `public/voices/`; the 60-second reference WAVs the backend conditions on live under the gitignored `backend_modal/voices/`.
131
 
132
  ---
133
 
 
141
  # Hugging Face token for the script-writing LLM (Inference API access)
142
  export HF_TOKEN=your_hf_token_here
143
 
144
+ # Deploy the GPU backend separately (needs the gitignored model code + voices under backend_modal/)
145
  # VIBEVOICE_PROFILE=tail modal deploy backend_modal/modal_runner.py
146
 
147
  python app.py # http://localhost:7860
 
172
  │ └── voices/ # Voice preview clips
173
  ├── text_examples/ # Example scripts (1–4 speakers)
174
  ├── tests/ # Script-parser tests + example prompts
175
+ └── backend_modal/ # Modal runner (tracked); VibeVoice model code + reference voices (gitignored)
176
  ```
177
 
178
  ---
app.py CHANGED
@@ -58,13 +58,13 @@ DEFAULT_SPEAKERS = ["Cylinder", "Statesman", "Novella", "Eyre"]
58
 
59
  SCRIPT_GEN_MODEL = "Qwen/Qwen2.5-Coder-32B-Instruct"
60
  WORDS_PER_MINUTE = 150 # Matches the pace assumed by the client's duration estimate
61
- DURATION_OPTIONS_MINUTES = [1, 2, 5, 10, 15, 20, 30, 45, 60]
62
  MAX_COMPLETION_TOKENS = 8192 # Good-faith ceiling for a single chat_completion call; the
63
  # underlying provider may cap lower, in which case the longest
64
  # duration options may come back shorter than requested
65
- MAX_SCRIPT_WORDS = 100000 # Effectively uncapped (2026-08-14, Josh) — backend chunking has no
66
- # real ceiling; the old 20,000 cap was a leftover UI guess that
67
- # blocked genuine long-form renders below the backend's actual limit
68
  MAX_TURNS = 250 # Hard ceiling regardless of target length (safety valve)
69
  MAX_GEN_ROUNDS = 24 # Hard cap on LLM calls per script regardless of target length
70
  MIN_TARGET_FRACTION = 0.9 # Stop extending once the script reaches 90% of the word target
@@ -601,6 +601,9 @@ def _prune_audio_store() -> None:
601
  _RATE_LOG: dict[str, deque] = defaultdict(deque)
602
  SCRIPT_RATE_LIMIT = (5, 600) # 5 script generations per 10 min per IP (hits paid HF inference)
603
  AUDIO_RATE_LIMIT = (3, 3600) # 3 audio generations per hour per IP (hits paid Modal GPU time)
 
 
 
604
  GENERATION_CONCURRENCY = asyncio.Semaphore(4) # at most 4 Modal generations in flight at once, globally —
605
  # matches the backend's max_containers cap (2026-09-06)
606
 
@@ -612,22 +615,31 @@ def _client_ip(request: Request) -> str:
612
  return request.client.host if request.client else "unknown"
613
 
614
 
615
- def _enforce_rate_limit(bucket: str, request: Request, limit: int, window_seconds: int) -> None:
616
- key = f"{bucket}:{_client_ip(request)}"
 
 
 
617
  now = time.time()
618
  log = _RATE_LOG[key]
619
  while log and now - log[0] > window_seconds:
620
  log.popleft()
621
  if len(log) >= limit:
622
  retry_after = max(1, int(window_seconds - (now - log[0])))
623
- raise HTTPException(
624
- status_code=429,
625
- detail=f"Rate limit reached ({limit} per {window_seconds // 60} min). Try again in about {retry_after}s.",
626
- headers={"Retry-After": str(retry_after)},
627
- )
628
  log.append(now)
629
 
630
 
 
 
 
 
 
 
 
 
631
  # ========================================================
632
  # FASTAPI APP
633
  # ========================================================
@@ -716,7 +728,11 @@ async def _backend_stats() -> dict | None:
716
  async def api_status() -> dict:
717
  """Backend reachability plus, when Modal reports it, whether a GPU container
718
  is already hot — the UI uses that to warn about the ~3 min cold path."""
719
- payload = {"backend": "ready" if remote_generate_function is not None else "offline"}
 
 
 
 
720
  stats = await _backend_stats()
721
  if stats is not None:
722
  payload.update(stats)
@@ -856,6 +872,10 @@ class GenerateRequest(BaseModel):
856
  @app.post("/api/generate")
857
  async def api_generate(payload: GenerateRequest, request: Request) -> StreamingResponse:
858
  _enforce_rate_limit("audio", request, *AUDIO_RATE_LIMIT)
 
 
 
 
859
  if remote_generate_function is None:
860
  raise HTTPException(status_code=503, detail="Modal backend is offline.")
861
 
 
58
 
59
  SCRIPT_GEN_MODEL = "Qwen/Qwen2.5-Coder-32B-Instruct"
60
  WORDS_PER_MINUTE = 150 # Matches the pace assumed by the client's duration estimate
61
+ DURATION_OPTIONS_MINUTES = [1, 2, 5, 10, 15, 20, 30, 45] # 45 × 150 wpm fits MAX_SCRIPT_WORDS
62
  MAX_COMPLETION_TOKENS = 8192 # Good-faith ceiling for a single chat_completion call; the
63
  # underlying provider may cap lower, in which case the longest
64
  # duration options may come back shorter than requested
65
+ MAX_SCRIPT_WORDS = 7000 # Public cap (2026-09-07, launch): ~45 min of audio, ~8 min of A100
66
+ # time per job. The backend itself has no ceiling — record-length
67
+ # renders call Modal directly and bypass this guard.
68
  MAX_TURNS = 250 # Hard ceiling regardless of target length (safety valve)
69
  MAX_GEN_ROUNDS = 24 # Hard cap on LLM calls per script regardless of target length
70
  MIN_TARGET_FRACTION = 0.9 # Stop extending once the script reaches 90% of the word target
 
601
  _RATE_LOG: dict[str, deque] = defaultdict(deque)
602
  SCRIPT_RATE_LIMIT = (5, 600) # 5 script generations per 10 min per IP (hits paid HF inference)
603
  AUDIO_RATE_LIMIT = (3, 3600) # 3 audio generations per hour per IP (hits paid Modal GPU time)
604
+ AUDIO_DAILY_BUDGET = (60, 86400) # 60 audio generations per rolling day across ALL visitors — bounds the
605
+ # worst-case GPU bill regardless of how many IPs show up (2026-09-07)
606
+ GLOBAL_KEY = "__global__"
607
  GENERATION_CONCURRENCY = asyncio.Semaphore(4) # at most 4 Modal generations in flight at once, globally —
608
  # matches the backend's max_containers cap (2026-09-06)
609
 
 
615
  return request.client.host if request.client else "unknown"
616
 
617
 
618
+ def _enforce_rate_limit(bucket: str, request: Request | None, limit: int, window_seconds: int,
619
+ message: str | None = None) -> None:
620
+ """Sliding-window counter. request=None counts globally (every visitor shares the bucket)."""
621
+ who = _client_ip(request) if request is not None else GLOBAL_KEY
622
+ key = f"{bucket}:{who}"
623
  now = time.time()
624
  log = _RATE_LOG[key]
625
  while log and now - log[0] > window_seconds:
626
  log.popleft()
627
  if len(log) >= limit:
628
  retry_after = max(1, int(window_seconds - (now - log[0])))
629
+ if message is None:
630
+ message = f"Rate limit reached ({limit} per {window_seconds // 60} min). Try again in about {retry_after}s."
631
+ raise HTTPException(status_code=429, detail=message, headers={"Retry-After": str(retry_after)})
 
 
632
  log.append(now)
633
 
634
 
635
+ def _audio_budget_remaining() -> int:
636
+ now = time.time()
637
+ log = _RATE_LOG[f"audio-day:{GLOBAL_KEY}"]
638
+ while log and now - log[0] > AUDIO_DAILY_BUDGET[1]:
639
+ log.popleft()
640
+ return max(0, AUDIO_DAILY_BUDGET[0] - len(log))
641
+
642
+
643
  # ========================================================
644
  # FASTAPI APP
645
  # ========================================================
 
728
  async def api_status() -> dict:
729
  """Backend reachability plus, when Modal reports it, whether a GPU container
730
  is already hot — the UI uses that to warn about the ~3 min cold path."""
731
+ payload = {
732
+ "backend": "ready" if remote_generate_function is not None else "offline",
733
+ "daily_remaining": _audio_budget_remaining(),
734
+ "max_script_words": MAX_SCRIPT_WORDS,
735
+ }
736
  stats = await _backend_stats()
737
  if stats is not None:
738
  payload.update(stats)
 
872
  @app.post("/api/generate")
873
  async def api_generate(payload: GenerateRequest, request: Request) -> StreamingResponse:
874
  _enforce_rate_limit("audio", request, *AUDIO_RATE_LIMIT)
875
+ _enforce_rate_limit(
876
+ "audio-day", None, *AUDIO_DAILY_BUDGET,
877
+ message="Today's free demo quota is used up. It refills over the next 24 hours — please come back later.",
878
+ )
879
  if remote_generate_function is None:
880
  raise HTTPException(status_code=503, detail="Modal backend is offline.")
881
 
backend_modal/modal_runner.py CHANGED
@@ -14,6 +14,23 @@ import pickle
14
  # Modal-specific imports
15
  import modal
16
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
17
  # Define the Modal Stub
18
  image = (
19
  modal.Image.debian_slim(python_version="3.10")
@@ -40,6 +57,7 @@ image = (
40
  "ln -s /root/schedule /root/vibevoice/schedule"
41
  )
42
  .env({"PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True"}) # fights fragmentation across waves
 
43
  .add_local_dir("backend_modal/modular", remote_path="/root/modular")
44
  .add_local_dir("backend_modal/processor", remote_path="/root/processor")
45
  .add_local_dir("backend_modal/voices", remote_path="/root/voices")
@@ -65,12 +83,12 @@ cache_volume = modal.Volume.from_name("vibevoice-cache", create_if_missing=True)
65
 
66
  @app.cls(
67
  gpu="A100-40GB",
68
- scaledown_window=300,
69
  timeout=7200, # was 3600: a 120-min+ record attempt hit the 1h ceiling at the
70
  # finish line (2026-08-19), losing the whole render. Long-form
71
  # with cloned voices runs slower than the preset-voice record
72
  # pace (bigger reference prefill per chunk), so give 2h.
73
- volumes={"/cache": cache_volume}
 
74
  )
75
  class VibeVoiceModel:
76
  @modal.enter()
@@ -89,7 +107,9 @@ class VibeVoiceModel:
89
  from modular.modeling_vibevoice_inference import VibeVoiceForConditionalGenerationInference
90
  from processor.vibevoice_processor import VibeVoiceProcessor
91
 
92
- print("Entering container and loading models to GPU...")
 
 
93
 
94
  # Set compiler flags for better performance
95
  if torch.cuda.is_available() and hasattr(torch, '_inductor'):
@@ -122,6 +142,8 @@ class VibeVoiceModel:
122
 
123
  self.setup_voice_presets()
124
  self.ready_at = time.time() # for cold-start detection in timing reports
 
 
125
  # VRAM baseline right after model load: requests arriving to a GPU far
126
  # above this are hitting a poisoned container (e.g. a cancelled run's
127
  # zombie generation thread) and must recycle, not proceed (2026-08-14)
@@ -816,6 +838,11 @@ class VibeVoiceModel:
816
  falls back to the named preset in speaker_N.
817
  """
818
  try:
 
 
 
 
 
819
  if model_name not in self.models:
820
  raise ValueError(f"Unknown model: {model_name}")
821
 
 
14
  # Modal-specific imports
15
  import modal
16
 
17
+ # --- Scaling profiles (2026-09-06) ---------------------------------------
18
+ # Chosen at deploy time: VIBEVOICE_PROFILE=launch modal deploy backend_modal/modal_runner.py
19
+ # "launch": one container always hot + one pre-warmed buffer under load, so a
20
+ # public-post burst lands warm instead of paying the ~3 min model load.
21
+ # "tail": scale to zero when idle (default). Same cap and idle window, so the
22
+ # only thing that changes is the standing spend.
23
+ # max_containers is the spending cap in both: a 5th concurrent request queues
24
+ # rather than spawning a 5th A100.
25
+ SCALING_PROFILES = {
26
+ "launch": dict(min_containers=1, buffer_containers=1, max_containers=4, scaledown_window=1200),
27
+ "tail": dict(min_containers=0, buffer_containers=0, max_containers=4, scaledown_window=1200),
28
+ }
29
+ SCALING_PROFILE = os.environ.get("VIBEVOICE_PROFILE", "tail").strip().lower()
30
+ if SCALING_PROFILE not in SCALING_PROFILES:
31
+ raise SystemExit(f"VIBEVOICE_PROFILE must be one of {sorted(SCALING_PROFILES)}, got {SCALING_PROFILE!r}")
32
+ print(f"Deploying with scaling profile '{SCALING_PROFILE}': {SCALING_PROFILES[SCALING_PROFILE]}")
33
+
34
  # Define the Modal Stub
35
  image = (
36
  modal.Image.debian_slim(python_version="3.10")
 
57
  "ln -s /root/schedule /root/vibevoice/schedule"
58
  )
59
  .env({"PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True"}) # fights fragmentation across waves
60
+ .env({"VIBEVOICE_PROFILE": SCALING_PROFILE}) # so container logs name the profile they run under
61
  .add_local_dir("backend_modal/modular", remote_path="/root/modular")
62
  .add_local_dir("backend_modal/processor", remote_path="/root/processor")
63
  .add_local_dir("backend_modal/voices", remote_path="/root/voices")
 
83
 
84
  @app.cls(
85
  gpu="A100-40GB",
 
86
  timeout=7200, # was 3600: a 120-min+ record attempt hit the 1h ceiling at the
87
  # finish line (2026-08-19), losing the whole render. Long-form
88
  # with cloned voices runs slower than the preset-voice record
89
  # pace (bigger reference prefill per chunk), so give 2h.
90
+ volumes={"/cache": cache_volume},
91
+ **SCALING_PROFILES[SCALING_PROFILE],
92
  )
93
  class VibeVoiceModel:
94
  @modal.enter()
 
107
  from modular.modeling_vibevoice_inference import VibeVoiceForConditionalGenerationInference
108
  from processor.vibevoice_processor import VibeVoiceProcessor
109
 
110
+ self.boot_started = time.time()
111
+ print(f"CONTAINER_START profile={SCALING_PROFILE} at={datetime.utcnow().isoformat()}Z "
112
+ "— loading models to GPU...")
113
 
114
  # Set compiler flags for better performance
115
  if torch.cuda.is_available() and hasattr(torch, '_inductor'):
 
142
 
143
  self.setup_voice_presets()
144
  self.ready_at = time.time() # for cold-start detection in timing reports
145
+ self.jobs_served = 0
146
+ print(f"CONTAINER_READY load_seconds={self.ready_at - self.boot_started:.0f}")
147
  # VRAM baseline right after model load: requests arriving to a GPU far
148
  # above this are hitting a poisoned container (e.g. a cancelled run's
149
  # zombie generation thread) and must recycle, not proceed (2026-08-14)
 
838
  falls back to the named preset in speaker_N.
839
  """
840
  try:
841
+ self.jobs_served = getattr(self, "jobs_served", 0) + 1
842
+ container_age = time.time() - getattr(self, "ready_at", time.time())
843
+ print(f"JOB_START cold={container_age < 120} container_age_s={container_age:.0f} "
844
+ f"job_n={self.jobs_served} model={model_name} words={len(script.split())} "
845
+ f"at={datetime.utcnow().isoformat()}Z")
846
  if model_name not in self.models:
847
  raise ValueError(f"Unknown model: {model_name}")
848
 
static/app.js CHANGED
@@ -83,7 +83,7 @@ const el = {};
83
  "composerCollapsedStrip", "collapsedSummary", "composerBody",
84
  "playerStage", "stageTitle", "stagePlayBtn", "stageWaveform", "stageTime",
85
  "stageDot", "stageLine", "stageSpeaker", "stageCloseBtn", "stageDownloadBtn", "stageOrbs",
86
- "stageScriptToggle", "stageTranscript", "soundSeg", "soundHint", "polishStatus", "cleanNoiseCheckbox",
87
  "generationTime", "audioDuration", "resultModel", "downloadBtn",
88
  "realtimeRow", "realtimeFactor", "warmupRow", "warmupTime",
89
  "downloadMp3Btn", "stageDownloadMp3Btn",
@@ -1341,9 +1341,8 @@ function buildSyncedTranscript(snapshot) {
1341
  lastCaptionKey = "";
1342
 
1343
  el.syncedTranscript.innerHTML = "";
1344
- el.stageTranscript.innerHTML = "";
1345
  state.resultTurns.forEach((turn, i) => {
1346
- state.resultTurns[i].rows = [el.syncedTranscript, el.stageTranscript].map((container) => {
1347
  const row = document.createElement("div");
1348
  row.className = "sync-line";
1349
  const dot = document.createElement("span");
@@ -1866,15 +1865,6 @@ el.openPlayerBtn.addEventListener("click", openPlayerStage);
1866
  el.stageCloseBtn.addEventListener("click", () => el.playerStage.close());
1867
  el.playerStage.addEventListener("click", (e) => { if (e.target === el.playerStage) el.playerStage.close(); });
1868
 
1869
- el.stageScriptToggle.addEventListener("click", () => {
1870
- el.stageTranscript.hidden = !el.stageTranscript.hidden;
1871
- el.stageScriptToggle.textContent = el.stageTranscript.hidden ? "View full script" : "Hide script";
1872
- if (!el.stageTranscript.hidden) {
1873
- const active = state.resultTurns[state.activeSyncIndex];
1874
- const target = (active && active.rows && active.rows[1]) || el.stageTranscript;
1875
- target.scrollIntoView({ block: "nearest", behavior: "smooth" });
1876
- }
1877
- });
1878
 
1879
  document.addEventListener("keydown", (e) => {
1880
  if (!el.playerStage.open || e.code !== "Space") return;
@@ -2303,12 +2293,13 @@ function resetPreview() {
2303
  preview.nextTakeTime = 0;
2304
  preview.normGain = 1;
2305
  preview.done = false;
2306
- el.genPreviewRow.hidden = true;
2307
  }
2308
 
2309
  function stopPreview(fadeSecs = 0.3) {
2310
  preview.done = true;
2311
  preview.active = false;
 
2312
  if (!preview.ctx) return;
2313
  const ctx = preview.ctx;
2314
  preview.ctx = null;
@@ -2439,25 +2430,40 @@ function previewPositionSeconds() {
2439
  return pos;
2440
  }
2441
 
 
 
 
2442
  function updatePreviewUI() {
2443
- if (preview.done || !preview.active) return;
2444
  el.genPreviewRow.hidden = false;
 
 
 
 
 
 
 
 
 
2445
  const n = Math.min(preview.nextIndex, preview.total || preview.nextIndex);
 
2446
  if (preview.ctx && preview.ctx.state === "suspended") {
2447
  // Browser blocked audio start without a fresh gesture.
2448
  el.genPreviewLabel.textContent = "Preview ready — click Unmute to listen while it renders";
 
2449
  return;
2450
  }
2451
- el.genPreviewLabel.textContent = preview.total
2452
- ? `Live preview playing · ${n} of ${preview.total} chunks rendered`
2453
- : "Live preview playing";
2454
  }
2455
 
2456
  el.genPreviewMute.addEventListener("click", () => {
2457
  preview.muted = !preview.muted;
2458
  el.genPreviewMute.textContent = preview.muted ? "Unmute" : "Mute";
2459
  if (preview.master) preview.master.gain.value = preview.muted ? 0 : preview.normGain;
2460
- if (preview.ctx && preview.ctx.state === "suspended") preview.ctx.resume().catch(() => {});
 
 
 
2461
  });
2462
 
2463
  /* ---------------- Generate ---------------- */
@@ -2514,6 +2520,7 @@ el.generateBtn.addEventListener("click", async () => {
2514
  el.logToggleBtn.hidden = true;
2515
  el.logToggleBtn.textContent = "View generation log";
2516
  startProgress();
 
2517
  // The render is already submitted — edits here can't reach it, so lock the
2518
  // workspace rather than let controls silently no-op, and put progress
2519
  // front and centre.
 
83
  "composerCollapsedStrip", "collapsedSummary", "composerBody",
84
  "playerStage", "stageTitle", "stagePlayBtn", "stageWaveform", "stageTime",
85
  "stageDot", "stageLine", "stageSpeaker", "stageCloseBtn", "stageDownloadBtn", "stageOrbs",
86
+ "soundSeg", "soundHint", "polishStatus", "cleanNoiseCheckbox",
87
  "generationTime", "audioDuration", "resultModel", "downloadBtn",
88
  "realtimeRow", "realtimeFactor", "warmupRow", "warmupTime",
89
  "downloadMp3Btn", "stageDownloadMp3Btn",
 
1341
  lastCaptionKey = "";
1342
 
1343
  el.syncedTranscript.innerHTML = "";
 
1344
  state.resultTurns.forEach((turn, i) => {
1345
+ state.resultTurns[i].rows = [el.syncedTranscript].map((container) => {
1346
  const row = document.createElement("div");
1347
  row.className = "sync-line";
1348
  const dot = document.createElement("span");
 
1865
  el.stageCloseBtn.addEventListener("click", () => el.playerStage.close());
1866
  el.playerStage.addEventListener("click", (e) => { if (e.target === el.playerStage) el.playerStage.close(); });
1867
 
 
 
 
 
 
 
 
 
 
1868
 
1869
  document.addEventListener("keydown", (e) => {
1870
  if (!el.playerStage.open || e.code !== "Space") return;
 
2293
  preview.nextTakeTime = 0;
2294
  preview.normGain = 1;
2295
  preview.done = false;
2296
+ updatePreviewUI();
2297
  }
2298
 
2299
  function stopPreview(fadeSecs = 0.3) {
2300
  preview.done = true;
2301
  preview.active = false;
2302
+ el.genPreviewRow.hidden = true;
2303
  if (!preview.ctx) return;
2304
  const ctx = preview.ctx;
2305
  preview.ctx = null;
 
2430
  return pos;
2431
  }
2432
 
2433
+ const PREVIEW_WAITING_NOTE =
2434
+ "Audio starts streaming here as soon as the first stretch of dialogue is rendered and passes the quality check — usually a minute or two in.";
2435
+
2436
  function updatePreviewUI() {
2437
+ if (preview.done) return;
2438
  el.genPreviewRow.hidden = false;
2439
+ if (!preview.active) {
2440
+ // Nothing is audible yet: say what will happen rather than claim playback.
2441
+ el.genPreviewRow.classList.add("waiting");
2442
+ el.genPreviewMute.hidden = true;
2443
+ el.genPreviewLabel.textContent = PREVIEW_WAITING_NOTE;
2444
+ return;
2445
+ }
2446
+ el.genPreviewRow.classList.remove("waiting");
2447
+ el.genPreviewMute.hidden = false;
2448
  const n = Math.min(preview.nextIndex, preview.total || preview.nextIndex);
2449
+ const progressNote = preview.total ? ` · ${n} of ${preview.total} chunks rendered` : "";
2450
  if (preview.ctx && preview.ctx.state === "suspended") {
2451
  // Browser blocked audio start without a fresh gesture.
2452
  el.genPreviewLabel.textContent = "Preview ready — click Unmute to listen while it renders";
2453
+ el.genPreviewMute.textContent = "Unmute";
2454
  return;
2455
  }
2456
+ el.genPreviewLabel.textContent = (preview.muted ? "Live preview muted" : "Live preview playing") + progressNote;
 
 
2457
  }
2458
 
2459
  el.genPreviewMute.addEventListener("click", () => {
2460
  preview.muted = !preview.muted;
2461
  el.genPreviewMute.textContent = preview.muted ? "Unmute" : "Mute";
2462
  if (preview.master) preview.master.gain.value = preview.muted ? 0 : preview.normGain;
2463
+ if (preview.ctx && preview.ctx.state === "suspended") {
2464
+ preview.ctx.resume().then(updatePreviewUI).catch(() => {});
2465
+ }
2466
+ updatePreviewUI();
2467
  });
2468
 
2469
  /* ---------------- Generate ---------------- */
 
2520
  el.logToggleBtn.hidden = true;
2521
  el.logToggleBtn.textContent = "View generation log";
2522
  startProgress();
2523
+ ensurePreviewCtx(); // inside the click gesture: an AudioContext made here starts unblocked
2524
  // The render is already submitted — edits here can't reach it, so lock the
2525
  // workspace rather than let controls silently no-op, and put progress
2526
  // front and centre.
static/index.html CHANGED
@@ -203,10 +203,10 @@
203
  <div class="progress-fill" id="genStageFill"></div>
204
  </div>
205
  <div class="gen-stage-meta" id="genStageMeta"></div>
206
- <div class="gen-preview" id="genPreviewRow" hidden>
207
  <span class="gen-preview-dot"></span>
208
  <span id="genPreviewLabel">Live preview</span>
209
- <button class="gen-preview-mute" id="genPreviewMute" type="button">Mute</button>
210
  </div>
211
  <div class="gen-stage-desc" id="genStageDesc"></div>
212
  <div class="gen-stage-log" id="genStageLog"></div>
@@ -248,9 +248,7 @@
248
  <a class="btn btn-ink" id="stageDownloadBtn" download="chorus-audio.wav">Download WAV</a>
249
  <a class="btn btn-pill-outline" id="stageDownloadMp3Btn" download="chorus-audio.mp3"
250
  title="Much smaller file — encoded on the server on first click">MP3</a>
251
- <button class="btn btn-pill-outline" id="stageScriptToggle" type="button">View full script</button>
252
  </div>
253
- <div class="synced-transcript stage-transcript" id="stageTranscript" hidden></div>
254
  </div>
255
  </div>
256
  </dialog>
 
203
  <div class="progress-fill" id="genStageFill"></div>
204
  </div>
205
  <div class="gen-stage-meta" id="genStageMeta"></div>
206
+ <div class="gen-preview waiting" id="genPreviewRow" hidden>
207
  <span class="gen-preview-dot"></span>
208
  <span id="genPreviewLabel">Live preview</span>
209
+ <button class="gen-preview-mute" id="genPreviewMute" type="button" hidden>Mute</button>
210
  </div>
211
  <div class="gen-stage-desc" id="genStageDesc"></div>
212
  <div class="gen-stage-log" id="genStageLog"></div>
 
248
  <a class="btn btn-ink" id="stageDownloadBtn" download="chorus-audio.wav">Download WAV</a>
249
  <a class="btn btn-pill-outline" id="stageDownloadMp3Btn" download="chorus-audio.mp3"
250
  title="Much smaller file — encoded on the server on first click">MP3</a>
 
251
  </div>
 
252
  </div>
253
  </div>
254
  </dialog>
static/styles.css CHANGED
@@ -980,7 +980,11 @@ body.is-generating .canvas { opacity: 0.6; transition: opacity 0.2s; }
980
  border-radius: 50%;
981
  background: var(--accent);
982
  animation: pulse 1.2s ease-in-out infinite;
 
983
  }
 
 
 
984
  .gen-preview-mute {
985
  border: 1px solid var(--border);
986
  background: transparent;
@@ -1077,8 +1081,6 @@ body.is-generating .canvas { opacity: 0.6; transition: opacity 0.2s; }
1077
  text-align: center;
1078
  }
1079
 
1080
- .stage-transcript { margin: 20px 0 0; max-height: 260px; }
1081
- .stage-transcript .sync-text { font-size: 0.88rem; }
1082
 
1083
  .icon-btn {
1084
  width: 30px; height: 30px;
 
980
  border-radius: 50%;
981
  background: var(--accent);
982
  animation: pulse 1.2s ease-in-out infinite;
983
+ flex: none;
984
  }
985
+ /* Before the first chunk lands: a quiet note, not a "playing" claim. */
986
+ .gen-preview.waiting { font-weight: 500; flex-wrap: wrap; text-align: center; max-width: 440px; margin-left: auto; margin-right: auto; }
987
+ .gen-preview.waiting .gen-preview-dot { animation: none; background: var(--border-strong); }
988
  .gen-preview-mute {
989
  border: 1px solid var(--border);
990
  background: transparent;
 
1081
  text-align: center;
1082
  }
1083
 
 
 
1084
 
1085
  .icon-btn {
1086
  width: 30px; height: 30px;