Home Pricing Documentation News About Guides 中文
Sign up
Speaker Labels in Transcription · Meeting or interview to labelled text

Speaker Labels in Transcription

Transcribe a meeting, interview or podcast with speaker labels: every line shows who said it, with time codes. Speaker diarization online, free to start.

No credit card · No watermark · Works in browser
Free tier to start 67 languages Privacy first
speakvid.com/tools/transcribe/workbench
Tasks
Speaker Labels in Transcription
Subtitle
Voiceover
TTS
S
Speaker Labels in Transcription
Progress visible in browser
Recording in · speaker-labelled transcript out Speaker-labelled transcript AI
Recording in · speaker-labelled transcript out
Output
SRT subtitle✓
Dubbed video✓
Transcript✓
Progress72%
67+Translation languages
300+AI voices
90Recognition languages
100%No watermark on export
Sound Familiar?

The old way is slow and expensive

Manual translation, hired voice actors, days of waiting. Creators and teams lose time and money on every single video.

Skip the old workflow

The transcript is one wall of text and I cannot tell which person said which line.

I have to keep rewinding the audio to work out who is speaking.

The dubbed version came out with one voice reading both sides of a conversation.

How It Works

Done in four simple steps

Upload, pick languages, and let the AI handle the rest. No software to install.

01

Upload the meeting, interview or podcast file — one recording, however many people are on it

A few clicks, that is all.

02

Choose the spoken language and run it; speaker separation is an option in the transcription step

A few clicks, that is all.

03

Read the result: each line is tagged with a speaker and a time code

A few clicks, that is all.

04

Rename the speakers in your editor if you want real names — or continue to dubbing and give each one a voice

A few clicks, that is all.

The hard part of an interview or a meeting transcript is not the words, it is the attribution. Without speaker labels a two-person conversation reads as one continuous voice, and you end up re-listening to work out who said what — which defeats the point of having a transcript. Turning on speaker separation during transcription splits the recording by voice and marks each spoken segment with a speaker label, so the result reads as a script: one speaker per line, with time codes, ready to be read top to bottom without going back to the audio. It changes the job more than it sounds: meeting minutes become possible without the recording, a quote can be attributed with confidence, and a podcast episode can be cut into a written interview. The same speaker labels carry into dubbing as well — when you translate and voice a multi-person recording, you can give each speaker a different voice, so the dubbed version does not collapse into one narrator reading a dialogue. It works best on recordings with a few clearly separated voices; heavy crosstalk on a single room microphone is the case to check. Costs 1 credit per minute of audio, with 20 free credits at sign-up and a daily free allowance; failed tasks are not charged.

A SpeakVid user story
Why SpeakVid

Everything you need, nothing you do not

<path d="M17 21v-2a4 4 0 0 0-4-4H5a4 4 0 0 0-4 4v2"/><circle cx="9" cy="7" r="4"/><path d="M23 21v-2a4 4 0 0 0-3-3.87"/><path d="M16 3.13a4 4 0 0 1 0 7.75"/>

Attribution, not just words

Each spoken segment is labelled by speaker, so the transcript reads as a script instead of a stream. You can follow a single person's thread, or paste a quote knowing exactly whose it is.

<circle cx="12" cy="12" r="10"/><path d="M12 6v6l4 2"/>

Still timestamped

Speaker labels come on top of the time codes, not instead of them — so you can see both who spoke and when, and export the whole thing as SRT or VTT.

<path d="M12 1a3 3 0 0 0-3 3v8a3 3 0 0 0 6 0V4a3 3 0 0 0-3-3z"/><path d="M19 10v2a7 7 0 0 1-14 0v-2"/><path d="M12 19v4"/>

A voice per speaker when dubbing

In the dubbing step you can assign a different voice to each speaker, so a translated multi-person recording still sounds like a conversation rather than a monologue.

What You Get

Inputs and limits

Tell who said what — every line labelled by speaker Upload a recording, get SPEAKER 1 / SPEAKER 2 on each line
SRT · VTT · TXT (with speaker labels) Give each speaker their own voice when dubbing
No watermark Clean exports, always
Fast processing Minutes, not days
FAQ

Common questions

What is speaker diarization?+

It is the process of splitting a recording by voice and labelling each spoken segment with its speaker, so the transcript shows who said what instead of one continuous block of text. The output here is a speaker label on each line, alongside the time codes.

How do I get speaker labels in a transcript?+

Upload the recording, and turn on speaker separation in the transcription step. The transcript that comes back tags every line with a speaker label and a time code, and can be exported as SRT, VTT or TXT. No separate tool is needed — it happens inside the same transcription run.

How many speakers can it tell apart?+

It is built for the usual case: two to four people in a meeting, an interview or a podcast, with clearly distinguishable voices. More than that, or several people sharing one room microphone with frequent interruptions, makes the boundaries less reliable — worth reading through once and fixing the odd line.

Can I put real names on the labels?+

The labels come out as neutral speaker tags (SPEAKER 1, SPEAKER 2 and so on), because nothing in the audio says a name. Rename them to real names in your editor or notes app — it is a find-and-replace on a handful of tags. In the dubbing step, the same labels are what let you give each speaker a different voice.

Does it work for two people on one microphone?+

Yes, that is a common setup and it usually works: two voices take turns and each turn gets labelled. What it cannot fully solve is talking over each other — where two voices overlap at once, the boundary is a best guess, and that is the part to check if a line seems attributed to the wrong person.

Is it free?+

Transcription costs 1 credit per minute of audio, sign-up includes 20 free credits plus a daily free allowance, and failed tasks are never charged. Speaker separation is part of the transcription step, not a separate paid add-on.

Still have questions? Contact us
Explore More

Related tasks you can also do

Ready to reach a global audience?

Try it now — free, fast and private.

Open the Workspace