What the labels mean
Speaker diarization groups speech that appears to come from similar voices. Person 1 and Person 2 are temporary labels created for the current session. They do not identify a person, authenticate a voice or prove that the same real-world person spoke every segment carrying the same label.
Diarization is separate from understanding the words. A transcript or translation can be broadly correct while its speaker label is wrong, and a correct speaker label does not make the words or translation correct.
Why results vary between models and methods
Übersetz supports streaming cloud providers, local Whisper-based workflows and Apple translation. Those workflows do not divide or return speech at identical boundaries or at identical times. Some results arrive progressively; others arrive after a completed audio chunk. Übersetz correlates audio intervals, transcript parts, translations and speaker estimates, but very short, delayed or corrected results can still be attributed imperfectly.
Changing the translation provider can therefore change the visible separation even when the voices are the same. A result that looks good with one conversation, language or provider is not evidence that the next session will be equally accurate.
Conditions that commonly cause errors
- Two people speaking at once, rapid interruptions or one-word responses.
- Similar voices, changing speaking style, whispering, shouting or speech that is too quiet.
- Background noise, music, room echo, poor microphones or compressed call audio.
- A voice reaching both microphone and system audio, especially when speakers are used instead of headphones.
- An incorrect attending-speaker count, too little speech from a participant or a new participant joining late.
- Long pauses, clipped beginnings or endings, provider corrections and translation chunks that do not match natural speaker turns.
Use it deliberately
- Set the known number of attending speakers when you know it; use Auto-detect only when the count is genuinely unknown.
- At the start, let each person speak a clear sentence in turn and check whether the labels separate correctly before relying on the display.
- Use headphones for calls, reduce background noise and avoid overlapping speech where practical.
- Verify labels again after changing the input source, translation provider, room, microphone or participant group.
- Rename Person labels only after you have independently confirmed the mapping. A name you assign does not make later attribution reliable.
- If separation degrades, select Clear to reset the session, numbering and learned speaker state, then start again with the correct speaker count—or turn diarization off.
When not to rely on it
Do not use speaker labels as the sole record for medical, legal, employment, financial, compliance, safety-critical or disciplinary decisions. Do not use them for voice identification, attendance verification, consent or accusations about who made a statement.
Übersetz does not archive raw audio. Save exports the text and labels, not evidence that can establish the speaker. If attribution matters, compare the result with an authorized original recording or another authoritative source and follow the consent and recording rules that apply to you.
Report a reproducible problem
If the same failure repeats, send the app version, macOS version, Mac model, microphone or system-audio choice, translation provider, attending-speaker setting and a short description of the turn pattern. Never send API keys. Share audio or a transcript only when you have the right to do so and have removed confidential material.