Purplelink
← All tools

Free web tool

Transcript Cleaner

Turn a Zoom or Microsoft Teams transcript (.vtt), a subtitle file (.srt) or a text file with speaker labels into plain text you can read and quote. The WEBVTT header, cue numbers and timestamps are removed, and each speaker's lines are joined into one paragraph. It does not summarise, correct or reword anything. Everything runs in your browser. Your file is never uploaded.

A Zoom-style .vtt file

WEBVTT

1
00:00:03.120 --> 00:00:06.480
Jane Doe: Thanks for making time.

2
00:00:07.000 --> 00:00:09.400
Jane Doe: Let's start with the survey.

3
00:00:09.900 --> 00:00:12.000
Alex Sample: Sure.

After cleaning

Jane Doe: Thanks for making time. Let's start with the survey.

Alex Sample: Sure.
A made-up example, with speaker names kept and timestamps off.

Drop .vtt, .srt or .txt files here

You can add several files at once, up to 50 MB each.

Options
Paste text instead

Bytes of your text sent anywhere: 0

Files are read and cleaned in this tab, in a background worker. The only request the page makes is an anonymous usage count that names the file type. Your files and text are never sent anywhere.

What it changes and what it leaves alone

How to clean a Zoom or Teams transcript

  1. Get the transcript file. Download the transcript as a .vtt file. Zoom saves the transcript of a cloud recording as .vtt, and Microsoft Teams lets you download a meeting transcript as .vtt. An .srt file or a .txt file with speaker labels works too.
  2. Add the file. Drop it on the page or choose it. You can add several files at once.
  3. Pick the options. Keep speaker names, put a time before each turn, join each turn into one paragraph, and remove bracketed sound tags such as [inaudible]. Changing an option updates every file at once.
  4. Check the preview. Read a few turns and compare them with your original file. The cleaner does not correct the words, so any error in the transcript is still there.
  5. Copy or download. Copy the text, or download it as a .txt or .md file.

Questions

Is my file uploaded?
No. The page reads the file with your browser and cleans it in a background worker on your device. The meter on the page counts the requests the page makes while it works. The only one is an anonymous usage count that names the file type (vtt, srt or text), never the file name or any of its text.
Does it change the words?
No. It does not summarise, correct, translate or reword anything. Spelling, grammar, false starts and filler words stay as they are. The changes to the text itself are small: line breaks inside a cue become spaces, repeated spaces become one, the HTML codes for ampersand, less-than, greater-than and non-breaking space that VTT files use are turned back into characters, and caption tags such as italics are removed. If you tick the sound tag option, a fixed list of bracketed tags is removed too.
Which files and layouts can it read?
It reads .vtt (WebVTT), .srt (SubRip) and .txt files. In a VTT or SRT file, speaker names are taken from voice tags such as <v Jane Doe> (the layout Microsoft Teams uses) or from a name and colon at the start of each cue, such as Jane Doe: (the layout Zoom uses). A .txt file can use Name: lines, or a line such as [Jane Doe] 00:00:05 above each turn. It was built and tested on sample files written from the published formats, not on real exports. If a real Zoom or Teams file comes out wrong, send the first few lines with the names changed to ben@purplelink.llc.
How does it decide who is speaking?
Voice tags are taken as written. For Name: labels, the start of a line counts as a speaker only when at least half of the cues start with a label, or when at least two different capitalised names each appear three times or more. Ordinary text such as Note: see page 4 is not mistaken for a speaker. A cue with no name stays with the speaker above it.
What happens when people talk over each other?
Cues stay in the order they appear in the file. If one cue holds two voice tags, or two lines with different names, each speaker gets their own turn. If a cue starts before the previous speaker's cue has ended, the result says how many times that happened.
What does it do when there are no speaker names?
A pause of more than three seconds between cues starts a new paragraph. In a text file with no times, the blank lines you already have are kept as paragraph breaks.
Which sound tags are removed?
Only when you tick the option, and only a fixed list in square or round brackets: inaudible, unintelligible, crosstalk, laughter, applause, music, silence, pause, noise and a few similar words, with or without a time inside the brackets. Any other bracketed text, such as [NAME] or [sic], stays.
Does it remove names or other personal details?
No. It only removes the structure of the file. Names, places and other details in what people said are left as they are, so check the text before you share it.
How large a file can it handle?
Up to 50 MB per file. A two-hour interview is usually under 1 MB. Files are cleaned in a background worker, so the page stays usable. The preview shows the first 200,000 characters. Copy and the downloads include everything.

If this saves you time, you can leave a tip. It helps keep these tools free and online.