Skip to content

feat: show AssemblyAI speaker diarization in transcripts - #2416

Open
pavzagor wants to merge 3 commits into
CapSoftware:mainfrom
pavzagor:codex/assemblyai-diarization
Open

pavzagor wants to merge 3 commits into
CapSoftware:mainfrom
pavzagor:codex/assemblyai-diarization

Conversation

@pavzagor

@pavzagor pavzagor commented Oct 5, 2026 •

Copy link
Copy Markdown

Cap already stores AssemblyAI word speakers, but transcription did not request them and captions/UI discarded them. This enables diarization for full recordings and editable-transcript backfills, shows Speaker A/B labels in the transcript, editor and player captions, and preserves labels through transcript edits, video cuts, copying, VTT/text downloads and agent API round-trips. Existing transcripts without labels continue to render normally.

Live chunks remain provisional and use no speaker labels: AssemblyAI identities are scoped to a transcription request. On recording completion, Cap queues a full-recording transcription instead of promoting independent chunks into a misleading final transcript. This adds a full transcription pass for recordings previously eligible for live promotion; the final labels appear when that pass completes. Queue failures propagate for workflow retry.

Validation:

  • 2,749 web tests passed; 27 existing opt-in tests skipped. Added speaker-boundary, storage/edit/export round-trip, unknown-speaker, escaping, agent API, live handoff and translation validation tests.
  • pnpm typecheck and pnpm exec biome ci . --linter-enabled=false passed; scoped Biome checks passed.
  • Real AssemblyAI run through the local app, MySQL, MinIO and media server: 33.5-second synthetic two-speaker recording, four alternating turns, correctly identified A/B; transcript ID b7ffdd4a-2d21-4425-b8bd-bffa3f28159f.
  • Browser verified transcript seeking, speaker captions, caption edit → save → reload, editor labels, word deletion → render → reload with remapped timestamps, and VTT download preserving voice tags. Mobile checked at 360px and 393px without horizontal overflow.
  • Independent review completed; agent API metadata/entity findings resolved. Greptile translation finding fixed: both provider responses and cached translations must preserve cue IDs, timings and voice tags.

Also corrected the existing Slack-manifest test's stale expected brand color to match the current manifest, so the full web suite passes. No database migration or new environment variable is required; uses the existing ASSEMBLY_API_KEY.

Upstream validation on 2766dc0: CI and Recording Reliability passed. Greptile re-reviewed 24 files and added no new comments; security checks passed. Vercel preview remains blocked on Cap Software team authorization.

October 5 review fixes

Legacy cues containing literal text such as 2 < 3 no longer lose text. The parser strips recognized WebVTT tags and timestamps, then decodes entities; escaped VTT writes and React text rendering remain intact. Regression coverage includes caption display, copying, downloads, and the agent API.

Validation on ce473ccc: 49 focused transcript tests and 27 scheduling/live-handoff tests passed; scoped Biome passed. The actual-component transcript fixture used by #2417 also verifies complete literal text.

Cost-control review: recording-wide speaker identities require one canonical full-recording pass after provisional live chunks. This deliberately adds a pass to the previous successful-live-promotion path. The pre-feature workflow already used chunk transcription plus a full pass when promotion failed (321ae61b9, queueFullPassFallback). The canonical database claim prevents competing triggers from scheduling multiple canonical passes. Account-wide usage budgets and trusted media-duration enforcement remain broader pre-existing cost-control gaps; this change does not establish a new quota policy.

RetriggerConfidence Score: 4/5

The PR is not yet safe to merge because existing transcripts containing literal angle brackets can display and export truncated text.

Findings

  1. P1 Legacy cue text gets truncated ▶
Fix with agent prompt
### Issue 1
apps/web/lib/transcript-vtt.ts:36
When an existing transcript contains literal text such as `2 < 3`, the previous caption writer stored it without escaping the `<`. This expression treats the unmatched `<` as the start of a tag and removes the rest of the cue. The transcript view, copied text, and downloads then omit part of the spoken text. Strip recognized VTT tags without consuming literal angle brackets.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Summary

The PR requests speaker diarization for full-recording transcription, keeps live chunks provisional, and carries speaker labels through captions, editing, exports, translations, and the agent API.

  • The new VTT parser can truncate literal angle-bracket text in existing transcripts.

Reviews (1) · Last reviewed commit: "fix: validate translated transcript spea..."

Comment thread apps/web/lib/transcript-vtt.ts Outdated
@richiemcilroy

Copy link
Copy Markdown
Member

Thanks for the thorough testing here. We require Greptile 5/5 on the latest commit and green CI before review, so please rebase on current main (the branch is far behind) and trigger a Greptile re-review on the angle-bracket fix.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants