Voice biometrics

Identifying a person by the characteristics of their voice. How a voiceprint is built, what degrades it, what attacks it, and what controls it needs around it.

In short

Voice biometrics is the set of techniques that identify or verify a person from the measurable characteristics of their voice. A model processes an audio clip, extracts traits that depend on the anatomy of the vocal tract and on that person's way of speaking, and turns them into a biometric template.

What gets compared is not the content of what was said. It is how the speaker sounds.

How a voiceprint is built

The process has two moments separated in time, and the quality of the first conditions everything else.

  • Enrollment — the person provides one or several audio samples, from which the system derives the reference template. An enrollment done with noise, with a low-quality channel, or with too little audio produces a poor reference that drags errors along afterward.
  • Verification — on every later attempt, the system captures a new sample, generates its template, and compares it against the reference one. The result is a similarity measure checked against a threshold, the same as in facial biometrics.

The traits that get measured are not pitch or perceived volume, but properties of the sound spectrum derived from the speaker's anatomy and articulatory habits. That is what makes two people who "sound alike" to a human ear not necessarily look alike to the model.

Text-dependent and text-independent mode solve different cases

It is the design decision that most changes the user experience and the system's level of demand.

  • Text-dependent — the person says a specific phrase, the same one from enrollment or one the system indicates. The comparison is more constrained and usually needs less audio, in exchange for an explicit step.
  • Text-independent — the system verifies over natural speech, with no fixed phrase. It works for validating during an ongoing conversation, and generally requires more audio to reach the same confidence.

Text-dependent mode shows up in access and confirmations. Text-independent mode shows up in phone channels, where verification happens while the person explains what they need.

What degrades a voice and what attacks it

It is worth separating the two problems, because one is about quality and the other about intent, and they are not solved with the same thing.

What degrades it with no one attacking: background noise, channel compression, a microphone different from the one used at enrollment, a poor-quality call. The person themselves too: a cold, fatigue, the passing of the years. A voice changes more than a face, and quite a bit more than a fingerprint, which is why voice systems tend to need re-enrollment policies that other modalities do not.

What attacks it, with intent:

  • Replay — the attacker records the person and plays the audio back in front of the microphone. It is the equivalent of showing a photograph to the camera.
  • Synthesis and cloning — a model generates a voice that mimics the person from samples of theirs, which today can be obtained from a call, a public video, or a voice note. It is the voice variant of the same phenomenon that produces a facial deepfake.
  • Injection — the audio does not pass through the microphone: it is fed into the stream before it reaches the application. It is the same layer problem as a video injection attack.

The first two are presentation attacks and are covered with PAD applied to the voice modality. The third is a different matter, and requires verifying the origin of the stream and the integrity of the execution environment.

What controls it needs around it

A voice comparison alone answers a partial question, the same as a face comparison alone.

  • Presentation attack detection for voice — distinguishes a voice produced by a present person from one that is replayed or generated. It is the control equivalent to facial liveness detection, tried on another modality.
  • Explicit thresholds — error is configurable, and its two forms are FMR and FNMR. On a phone channel with variable audio, a threshold set for laboratory conditions rejects legitimate users all day long.
  • A second factor when the operation warrants it — voice works well as one of several authentication factors, and worse as the only one, especially on channels where the audio arrives compressed.

A facial liveness certification says nothing about a voice system. Each modality is tested separately, and a test report states which modality it was done on.

Frequently asked questions

It is the set of techniques that identify or verify a person from the measurable characteristics of their voice. A model processes an audio clip, extracts traits derived from the anatomy of the vocal tract and the way of speaking, and turns them into a biometric template that is compared against a reference one. What gets compared is how the speaker sounds, not what they say.

In text-dependent mode, the person says a specific phrase: the comparison is more constrained and usually requires less audio, in exchange for an explicit step for the user. In text-independent mode, the system verifies over natural speech, with no fixed phrase, which works for validating during an ongoing conversation and generally requires more audio to reach the same confidence.

It is the attack the modality has to solve, which is why a voice comparison alone is not enough. A model can generate a voice that mimics a person from public samples, and the biometric comparison answers for the resemblance, not for the origin. The right control is presentation attack detection on the voice modality, which distinguishes a voice produced by someone present from one that is replayed or generated. Audio injection, where the sound never passes through the microphone, is a different problem and is covered separately.

Almost always because of conditions, not an attack. Background noise, channel compression, a microphone different from the one used at enrollment, or a poor-quality call move the sample enough to fall below the threshold. The person also changes: a cold, fatigue, or the passing of time alter the voice more than they alter a face. That is why voice systems need re-enrollment policies that other modalities do not.

One identity, one SDK

VU ONE brings identity verification, authentication and fraud protection together on a single identity graph.

The verification you run at signup stays available to authentication and to your fraud rules, with no repeated processes and no duplicated data.

Verify, Authenticate and Protect, consolidated in one place.

Request a demo