Skip to content

fix: group transcript captions into readable sentences - #2417

Open
pavzagor wants to merge 7 commits into
CapSoftware:mainfrom
pavzagor:codex/fix-shitty-transcription
Open

pavzagor wants to merge 7 commits into
CapSoftware:mainfrom
pavzagor:codex/fix-shitty-transcription

Conversation

@pavzagor

@pavzagor pavzagor commented Oct 5, 2026 •

Copy link
Copy Markdown

The Transcript tab currently renders every subtitle cue as a separate row, splitting speech at commas, short pauses, and every eight words. This groups adjacent cues into sentence rows while preserving speaker boundaries and exact outer timestamps. Existing transcripts, translated captions, and provisional live transcripts benefit without retranscription.

Subtitle files retain their original timing. Owners can select Edit transcript to expose the original cues for precise corrections; edits never save a merged sentence under a single cue ID. Timestamped copying uses the readable sentence rows. Grouping also stops at long silence, overlaps, and bounded length/duration when punctuation is missing.

Depends on #2416 (AssemblyAI diarization). This branch is based on that PR's head, 2766dc0; merge that PR first. Until then, GitHub's cumulative diff includes the prerequisite. Review only this follow-up's four files.

English before/after proof

A fresh AssemblyAI transcription of NASA's public JFK archival clip produces 12 caption fragments before → 3 sentence rows after, preserving all 72 words. Selecting the matching row on either side seeks the embedded source video to 12.850 seconds.

Watch/download the before/after MP4 · GitHub video page · Source clip · All proof files and methodology

English transcript before and after

The proof uses the actual old/new React components and real ASR output, with storage/auth hooks mocked. It is not production footage or proof of a persisted backend edit. Full local share-page verification was blocked by unavailable MySQL at 127.0.0.1:3306. NASA's clip splices two speeches; A/B are the unmodified ASR labels.

Validation

  • 91 focused tests passed: sentence grouping/UI, diarization, VTT, text formatting, translation, caption generation, and transcript editing.
  • Scoped Biome check and git diff --check passed.
  • Web tsc --noEmit passed in the existing development checkout; the isolated checkout reused dependencies and was unsuitable for the workspace reference type check.
  • Browser checked sentence seeking, fragment editing controls, and 360px/393px layouts without horizontal overflow. Timestamped copying is covered by the UI tests.
  • No migration, new environment variable, or paid reprocessing is required by this change.

Review fix: multilingual sentence endings

Addressed the Arabic question-mark finding in decfd40 with Unicode Sentence_Terminal. Regression tests cover Arabic, quoted Arabic and Hindi, plus original/translated Arabic rendering and exact seeking. All 91 focused tests, TypeScript, scoped Biome and whitespace checks pass. The approved English output remains byte-for-byte equivalent as parsed sentence entries (12 fragments → 3 sentences).

Review-fix walkthrough · Validation log · English parity check

Arabic and Hindi sentence boundaries before and after

This additional proof uses the actual component with a synthetic multilingual fixture and mocked storage/auth.

October 5 review fixes

Single-letter labels such as Option A. now end their sentence. Standalone and consecutive initials, titles, and ellipses still remain grouped. A lone embedded initial remains ambiguous with a label and is treated as terminal. Switching between cue editing and sentence reading clears cue selection explicitly, so a selected non-leading cue cannot leave an inconsistent sentence highlight. The legacy literal-angle-bracket fix from #2416 is included.

Validation on e261911a: 81 focused transcript/parser/edit/export tests passed; scoped Biome passed. Actual-component fixtures verify sentence boundaries, complete 2 < 3, precise seeking to 2 seconds, and selection clearing at desktop and 390px mobile widths. Auth/storage/network are mocked in these fixtures; this is not authenticated full-page proof. The attempted full web typecheck did not pass: the isolated checkout lacks built shared-project declarations and reuses dependency links from the primary checkout. Fresh CI remains subject to upstream workflow approval.

Mobile reading fixture · After Done editing · Fixture methodology

Security review: the extra canonical full-recording pass belongs to prerequisite #2416 and is required for recording-wide speaker identity. The pre-feature workflow already fell back to a full pass after chunks when live promotion failed. The existing canonical database claim remains covered by scheduling tests. Per-owner usage budgets are not introduced by this sentence-grouping follow-up; the broader cost-control concern needs a separately defined quota policy.

RetriggerConfidence Score: 4/5

The follow-up appears safe to merge, with two non-blocking sentence-display issues worth correcting.

Findings

  1. P2 Single-letter endings merge sentences ▶
  2. P2 Selection disappears after editing ▶
Fix with agent prompt
### Issue 1
apps/web/lib/transcript-sentences.ts:7
A cue ending with a single-letter label such as “Option A.” is treated as an abbreviation. If the next cue starts “Choose another.”, the transcript displays both as one sentence row. This makes some sentence boundaries harder to read.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

### Issue 2
apps/web/app/s/[videoId]/_components/tabs/Transcript.tsx:443
If an owner selects the second cue of a merged sentence in edit mode and then clicks “Done editing,” no sentence row is highlighted. The merged row keeps only the first cue’s ID, so it cannot match the selected ID. Map the selection to the sentence row or clear it when switching modes.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Summary

The follow-up groups adjacent transcript cues into readable rows for canonical and provisional transcripts, while retaining original cues for editing and VTT download.

  • Adds sentence-boundary, speaker, silence, overlap, and size checks.
  • Adds tests for grouping, seeking, copying, editing, and multilingual endings.

Reviews (1) · Last reviewed commit: "fix: recognize multilingual sentence ter..."

@superagent-security superagent-security Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Superagent found 1 security concern(s).

Comment thread apps/web/workflows/live-transcribe.ts
Comment thread apps/web/lib/transcript-sentences.ts Outdated
Comment thread apps/web/app/s/[videoId]/_components/tabs/Transcript.tsx

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant