Overview of calibration models

Nuance provides a variety of factory calibration models and algorithms for tuning the biometric engine’s score based on your deployment’s data. Contact Nuance representative for advice about choosing algorithms.

Only the latest version of NTSL (v3), NVSL (v10), NCSL (v2), and NRSL (v1) are mentioned in this document. For more information on earlier releases of NTSL, NVSL, NCSL, and NRSL, contact your Nuance representative.

The performance of each model depends on the characteristics of the audio segments. It’s possible to have a general idea how a given algorithm affects accuracy, CPU load, and memory usage. But without a test, it isn’t possible to know the exact impact. In general, use the following factors to choose an algorithm:

  • The type of verification needed: text-dependent, text-independent, or text-prompted
  • For text-dependent verification, the type of passphrases needed: common or unique
  • For text-independent verification, the expected duration of the spoken audio

To choose a model, perform the following actions with the help of Nuance Professional Services:

  1. Use the following tables to identify the calibration models that best match the audio segments in your assessment. (Match the audio attributes with the table columns.)
  2. Test one model. Choose one of the models, and perform the assessment. When choosing a model, CPU, memory usage, voiceprint size as well as the size of the model itself, should all be taken into consideration as they might vary between different algorithms or calibration models.
  3. When processing completes, review the accuracy of the results.
  4. To test another calibration model, repeat the process.

The following table lists factory calibration models and the features they support. When creating a custom calibration model, select the baseline algorithm that the system should use.

Factory calibration models
Calibration model Text-dependent Text-independent Gender Detection Speaker Segmentation Adaptation from Print Identification from Print Multi-speaker identification Text verification & Audio consistency check Extended calibration
FACTORY_TD_COMMON_10 Yes No Yes No Yes Yes No Yes Yes
FACTORY_TD_CUSTOM_10 Yes No Yes No Yes Yes No Yes Yes
FACTORY_TD_CUSTOM_MVIMP_10 Yes No Yes No Yes Yes No Yes Deprecated
FACTORY_TD_UNIQUE_10 Yes No Yes No No No No Yes No
FACTORY_TI_10 No Yes Yes Yes Yes Yes Yes No Yes
FACTORY_TI_LIVENESS_10 No Yes Yes No Yes Yes No No Yes

When the system loads calibration models into memory, the amount of required space is primarily determined by the algorithms used to build the models. Other elements such as probability distributions and per-instance memory occupy small amounts of memory. In the following table, the memory estimates are for the algorithms alone (and the actual memory occupied by the calibration model is larger.) The memory for prints is highly variable, and largely depends on the number and length of the audio or text enrollment segments.

The following table lists hardware specifications and use cases for factory calibration models:

Specifications and use cases for factory calibration models
Algorithm Global Memory Print size Use case
FACTORY_TD_COMMON_10 31 MB 131 KB This is the seed model for speaker verification, text-dependent, common passphrase applications. This engine models passphrases with the format “At CompanyName, my Voice Is My Password”, but isn’t intended as a standalone model (use FACTORY_TD_CUSTOM_MVIMP_10 instead). Use this engine when the available training data is somewhat limited. The engine is based on HMM SVM technology. Use at least 250 utterances of the common passphrase spoken by different people. If multiple utterances of a big speakers set (1000 or more) are available, use FACTORY_TD_CUSTOM_10 instead (it likely provides better results). Recommended: a minimum of 250 audio elements. It’s okay to include a single recording of each speaker. To limit the model size and the enrollment time, avoid using more than 5000 extended-calibration utterances. The recordings should be homogeneous in terms of lexical content with the passwords used by the application. For optimal calibration, have one utterance per speaker, and maximize the speaker variability. If using a single person to speak in more than one recording, you must associate a numerical speaker identifier with the audio elements to denote utterances pronounced by the same person.
FACTORY_TD_CUSTOM_10 117 MB 14 KB This configuration is a seed model for speaker verification, text-dependent, common passphrase applications, when large quantities of training material are available. This engine is based on HMM i-vector technology. Use this model when a large number of audio files are available (1000 or more speakers, each one having more than 4 associated utterances). Recommended: For an extended calibration, a minimum of 5,000 audio elements that contain the common passphrase from 800 to 1,000 speakers. The recordings should be homogeneous in terms of lexical content with the passwords used by the application. For optimal calibration, have one utterance per speaker, and maximize the speaker variability. If using a single person to speak in more than one recording, you must associate a numerical speaker identifier with the audio elements to denote utterances pronounced by the same person.
FACTORY_TD_CUSTOM_MVIMP_10 113 MB 13 KB This configuration is a ready-to-use engine for text-dependent, common passphrase applications. It uses a special voice activity detection mechanism, optimized for passphrases with the format “At myCompanyName, my voice is my Password”. This engine configuration is based on HMM i-vector as well as DNN Embedding technologies Because there is no need for data collection for calibration and tuning purposes, the engine speeds application deployment. For more details see Managing calibration models. This engine doesn’t work well with different passwords. In general, it should be used as a standalone model and not as a seed model for different passphrases.
FACTORY_TD_UNIQUE_10 104 MB ~17 KB This engine is tailored for text-dependent unique passphrase scenarios. You can tune the engine by performing a basic calibration. Extended calibration isn’t supported by this engine. The tuning would improve the accuracy of this engine, especially for non-English applications. This is a double technology configuration, combining i-vector and HMGM models. Recommended: a minimum of 500 utterances to calibrate unique passphrases. Ideally data representing the possible passphrases should be used for the calibration.
FACTORY_TI_10 111 MB 5.3 KB This model is one of the text-independent frameworks. It includes speaker segmentation and multi-speaker identification capabilities. It’s the first of a new generation of speaker recognition models that are based on DNN embedding technology. Voiceprints created with previous text-independent engine configurations (NVSLv5, NVSLv6, NVSLv7, NVSLv8, and NVSLv9) aren’t compatible with the new text-independent framework. To train Gender Detection, provide audio files labeled with a minimum of 45 male and 45 female speakers (100 of each gender is recommended). If you don’t provide gender labels, the extended calibration inherits the Gender Detection accuracy of the base model. Recommended: a minimum of 400 audio elements for creating a custom calibration. The recordings should be homogeneous in terms of lexical content with the passwords used by the application. For optimal calibration, have one utterance per speaker, and maximize the speaker variability. If using a single person to speak in more than one recording, you must associate a numerical speaker identifier with the audio elements to denote utterances pronounced by the same person.
FACTORY_TI_LIVENESS_10 73 MB 7.8 KB This configuration is a text-independent engine particularly suited for Liveness-Detection tasks. It’s optimized for EN-US, telephonic applications but is easily tuned to other languages through extended PLDA calibration. It’s based on DNN embedding technology. To improve the accuracy of this engine towards specific deployments, use the training toolkit software (see the Speaker ID Training Toolkit Guide) to perform extended calibration (PLDA tuning), or create custom calibration sets. Voiceprints created with previous liveness frameworks from previous NVSL versions aren’t compatible with the new framework: you must retrain the voiceprints from the enrollment audio files. Recommended: a minimum of 400 audio elements for creating a custom calibration set. The recordings should be homogeneous in terms of lexical content with the passwords used by the application. For optimal calibration, have one utterance per speaker, and maximize the speaker variability. If using a single person to speak in more than one recording, you must associate a numerical speaker identifier with the audio elements to denote utterances pronounced by the same person.

An additional model, FACTORY_ANTISPOOFING_10, is used to evaluate non-voiceprint factors that might indicate call spoofing.

Following factory demonstration models are provided for testing purposes and aren’t to be used in a production environment:

  • FACTORY_TI_RISK_ENGINE_1_5_MIN_FALSE_ALARM
  • FACTORY_TI_RISK_ENGINE_1_5_BALANCED
  • FACTORY_TI_RISK_ENGINE_1_5_MAX_SPOOF_DETECTION
  • FACTORY CP_2 (NSTL2 demo engine configuration)
  • FACTORY_CP_3 (NSTL3 demo engine configuration)