# The [Inaudible] Problem: What Transcripts Hide

Every transcript has them: [inaudible], [unclear], [crosstalk]. Brackets that mark a gap where the transcriber could not resolve the audio.

Those brackets are honest — but they hide a question: Did a human ear try, or did the machine quit?

On difficult evidence audio, the most important lines are often the faintest. The threat muttered under breath. The admission spoken while turning away. The name half-obscured by wind.

A single-pass AI transcript marks those sections [inaudible] and moves on. The gap enters the record as if it were empty — when it may not be.

What "inaudible" really means

In a human-transcribed record, [inaudible] means the transcriber listened and could not make out the words.

In an AI transcript, [inaudible] can mean:

  • The model's confidence was below threshold
  • The audio was silent or pure noise
  • The model hallucinated something implausible and the post-processor flagged it
  • The speech was overlapping or distorted

The bracket does not tell you which. It does not tell you whether a human ever listened.

The problem with single-pass transcripts

A single AI pass produces one opinion. Where the audio is unclear, that opinion may be:

  • A hallucination (invented words)
  • [inaudible] (a guess that failed post-processing)
  • Nothing (the model skipped the section entirely)

You do not know which unless you listen to the audio yourself — at which point the transcript is redundant.

Multi-pass scatter as signal

Five AI passes listening to the same unclear audio will produce five different guesses. That scatter is not noise — it is data.

This is the practical defense against Whisper hallucinations and similar single-model failures — disagreement surfaces what one pass would confidently hide.

High agreement (5/5 or 4/5) — the audio is clear enough that independent models converge. Trust is warranted.

Moderate agreement (3/5) — the audio is marginal. Some models resolved it; some did not. Flag for review.

Low agreement (2/5 or 1/5) — the models are guessing. Mark for human verification.

No agreement (0/5) — the models could not converge on anything. This line needs an ear, urgently.

The scatter tells you where the transcript is solid and where it is not.

The ear-queue, sorted by importance

Not every [inaudible] matters equally. Background chatter, irrelevant exchanges, sections already corroborated by other evidence — those can stay bracketed.

But faint speech at a critical moment — a contested fact, an admission, a threat — cannot.

VeriVox queues low-agreement sections for human review and sorts them by importance:

  • Speaker transitions (who started speaking?)
  • Elevated volume (raised voices, arguments)
  • Long pauses followed by faint speech (hesitation, then admission)
  • Segments flagged by the user as critical

The ear-queue is a to-do list: start at the top, work down, stop when you have verified the lines that matter.

What a verified line looks like

A human listens to the queued section and makes a call:

  • Verified — the ear resolved it; the line is certified
  • Corrected — the machine guessed wrong; the correction is logged, original preserved
  • Truly inaudible — even a human ear cannot make it out; mark it honestly

The export labels which:


[03:12.4–03:16.1] ✓ SPEAKER-A: You were never supposed to be here today.
  agree 5/5 · verified by ear

[07:41.8–07:44.0] ✎ SPEAKER-A: The paperwork was already gone by then.
  agree 3/5 · corrected
  (ASR original: "the paper it was already gone been")

[22:03.1–22:04.6] ⚑ SPEAKER-A: [grave excerpt — operator-verified]
  agree 0/5 — machine could not resolve

[24:48.2–24:50.0] · UNATTRIBUTED: [inaudible — background media]
  excluded from excerpts · truly inaudible

Every line declares how it earned its place. The brackets mean something.

Disagreement as provenance

When multiple passes disagree, the disagreement itself is part of the record. It proves the section was difficult — not just for one model, but for all of them.

If opposing counsel challenges a verified line, you can show:

  • Five models produced five different outputs
  • None agreed
  • A human listened and certified the result
  • The original machine outputs are preserved

That is defensible. A single-pass [inaudible] is not.

Real pattern from real audio

In a contested recording, a single-pass Whisper transcript produced:


[22:03] [inaudible]

The audio was faint — wind, distance, overlap. Whisper's confidence was too low to guess.

When the same audio was run through VeriVox's multi-pass pipeline:

  • Pass 1 (Whisper): [inaudible]
  • Pass 2 (Wav2Vec2): "can do"
  • Pass 3 (NeMo): "thank you"
  • Pass 4 (Vosk): [silence]
  • Pass 5 (Faster-Whisper): "think through"

Agreement: 0/5. The line went to the ear-queue, flagged as grave (speaker conflict + volume spike).

A human verified it. The actual words were neither "can do" nor "thank you" — and they mattered to the case.

The single-pass bracket would have hidden that.

When to accept [inaudible]

Some audio is truly inaudible. Heavy static, complete speaker overlap, recording failure, physical obstruction of the microphone.

When five passes scatter and a human ear cannot resolve it, the honest answer is [inaudible — verified unresolvable].

That bracket is defensible. It says you tried.

What commercial transcription services do

Services like Rev and GMR use human transcribers. When they cannot make out a word, they mark it [inaudible] and move on.

Their quality depends on the transcriber's skill, the playback equipment, and how much time they spend on difficult sections. There is no multi-pass vote. There is no provenance for which sections were reviewed multiple times.

For clear audio, human transcription is excellent. For faint, contested evidence audio, one human ear is still one opinion.

The VeriVox workflow

VeriVox's Consensus and Verify stages handle unclear audio:

  1. Five transcription passes, aligned word-by-word
  2. Agreement scoring per line
  3. Low-agreement lines queued for review, sorted by importance
  4. User listens, verifies, corrects, or confirms inaudibility
  5. Export labels every line: verified, corrected, or machine-only
  6. Original machine outputs preserved

The [inaudible] brackets mean "a human tried and could not resolve it" — not "the machine quit."

See the full pipeline →

For court reporters and legal transcriptionists, learn how VeriVox supports professional workflows →

Why disagreement belongs in the record

When you present a transcript with verified corrections, opposing counsel may ask: "How do you know your correction is right?"

Your answer:

  • Five independent models disagreed
  • Here are the five outputs (preserved)
  • A qualified person listened multiple times
  • The verified line is certified

The disagreement is not a weakness. It is proof that the section was scrutinized.

Checklist: Handling unclear audio

  • [ ] Run multiple transcription passes (or use a multi-pass tool)
  • [ ] Align passes and score agreement per line
  • [ ] Queue low-agreement sections for human review
  • [ ] Sort the queue by importance (contested facts first)
  • [ ] Verify, correct, or confirm inaudibility — with ears
  • [ ] Label the result: verified, corrected, or truly inaudible
  • [ ] Preserve the original machine outputs
  • [ ] Include provenance in the export

FAQ

What if I only have a single-pass transcript?

Listen to every [inaudible] section yourself. If you can resolve it, note the correction and label it as human-verified. If you cannot, confirm it is truly inaudible.

How many passes do I need?

Three is the minimum for meaningful disagreement detection. Five is better. More than five has diminishing returns.

What if all five passes agree on something that sounds wrong?

Possible — rare, but possible. If the audio is clear and all models converge on implausible words, listen yourself. Trust your ear over consensus when the result is absurd.

Can I use different models for each pass?

Yes — different models are better. They make independent errors. Running the same model five times will produce the same output five times.

What if the disputed section is only two words?

Two words can change a case. If those two words are contested, they deserve the same scrutiny as a full paragraph. Queue them, verify them, certify them.

How does this relate to speaker attribution?

Disagreement in transcription and disagreement in speaker identification are related problems. Voice identification can flag segments where the diarization conflicts with enrolled voice scores — both forms of disagreement require human resolution.