Manage calibration models
A calibration model is a codified representation of a typical speaker and acoustic environment for enrolling and verifying speakers. It normalizes and removes the variations in the input data so that computations can focus on the relevant data attributes.
To manage calibration models, click the calibration icon on the left. The page lists all factory models in the system and all basic and extended models in the scope you selected. You can filter by calibration mode (including factory, basic, and extended), calibration type (including Voiceprints, Playback, and Risk Engine), or calibration state (including Active, Validated, Inactive, Initializing, and Invalid).
Click the three dot menu to see additional information including description, date created, date modified, and calibration model revision ID. You can also use the menu to download or activate calibration models.
The limitations are as follows:
- You cannot download a factory model.
- You can only activate a factory model if you are using the system scope.
- In all non system scopes you can only download or activate custom models (custom models are models of type Basic or Extended).
- You can only activate validated models or inactive models.
- You can only deactivate active models.
Upload a new calibration model
You can only upload a factory calibration model to the system scope. To upload a new model:
- Click Upload.
- Enter a calibration model ID, calibration type, a description, and select a configuration set. Engine framework and calibration mode are defined inside the calibration file.
- Click Browse and select the file to upload.
- Click Upload. By default, the model’s state is Validated.
- To start using the model, click the three-dot menu and select Activate.
Collect audio for calibration
When collecting and selecting audio for system calibration, it is important to use audio that represents the typical speakers and audio environments of the production application. If the system is not tuned on the type of audio used in production, the accuracy is sub-optimal.
To create a calibration model, collect at least 250 recordings (depending on the algorithm used) from a variety of speakers. Doing this provides a good representation of the local environment. Using more calls assures a good representation of the target speakers and environment, but it is not necessary to collect extreme quantities of audio (a few hundred files are sufficient).
Collect data from live calls
Silent Mode Data Collection (SMDC) collects data from live calls in contact centers without any agent interaction. The agent will not open a web agent console and is not aware that the call is recorded and used for data collection. You can define which extensions to use and for how long as well as enable automatic consent through the API. Transactions resulting from SMDC are billable by default, but often subject to special terms and conditions extended through a Professional Services engagement.
Collect audio according to the task and algorithm
The recommended number of total audio files and the recommended number of audio files for a speaker depend on the task type and the algorithm used.
The following table lists the suggested sizes depending on the algorithm and engine type:
| Model | Recommendation |
|---|---|
| FACTORY_TD_COMMON_9 | Seed configuration for text-dependent common passphrase scenarios when reduced calibration data is available. Extended calibration is strongly recommended and mandatory for phrases other than “At #siteName, my voice is my password”. |
| FACTORY_TD_CUSTOM_9 | Seed configuration for text-dependent common passphrase scenarios where substantial calibration data is available. Extended calibration is strongly recommended and mandatory for phrases other than “At #siteName, my voice is my password”. |
| FACTORY_TD_CUSTOM_MVIMP_9 | Ready-to-use text-dependent common password engine configuration that models the American English passphrase, “My Voice is My Password”, with an optional “At XYZ” variable preamble. This built-in framework provides peak accuracy without requiring any additional calibration data. |
| FACTORY_TD_UNIQUE_9 | Text-dependent speaker recognition for unique passphrase scenarios as well as scenarios with little to no calibration data. 250-5000 utterances (ideally a single audio for each speaker). |
| FACTORY_TI_9 | Text-independent speaker recognition for call center applications. 250-5000 utterances (ideally a single audio for each speaker). Also supports extended calibration. |
| FACTORY_TI_LIVENESS_9 | Text-independent speaker recognition for text-prompted liveness detection. 250-5000 utterances (ideally a single audio for each speaker). Also supports extended calibration. |
| FACTORY_ANTISPOOFING_9 | Special engine configuration including the Anti-Spoofing features only: playback detection, synthetic speech detection and music detection. |
Recommended: Prior to operating the system, perform a small-scale experiment using audio from several speakers to set the decision threshold at the optimum point. In this experiment, measure the distribution of scores. This allows setting the decision threshold at the optimal working point. This process is called tuning.
Collect audio for fraud pattern detection
Corpus for creating a calibration model - Recommended: use a minimum of 1000 texts for fraud and 1000 texts for non-fraud, sampled to represent your general population. Ideally, each text should come from a different speaker. At a minimum, there should be a variety of 1000 different speakers (or writers for the digital world) in this text collection. Each text should have 80 to 100 unique words.
Corpus for the TUIT experiment - Recommended: approximately 3000 texts for fraud and 3000 texts for non-fraud, Ideally, each text should originate from different authors and have 80 to 100 unique words. Use NTE fast mode for the experiment (because it delivers a good trade-off between accuracy and computational time).
Create an extended calibration
Extended calibration is a technique to maximize the ability of the system to discriminate between speakers and to ignore variations for individual speakers. This is particularly powerful for audio data that originates from different sources (telephone, microphone, mobile lines, and so on) because the system can learn to ignore differences between the sources.
For text-dependent engines, extended calibration learns the acoustic features of the selected common passphrase, increasing the precision and the accuracy of the model. This is why extended calibration is mandatory for text-dependent engines where the common passphrase is not “At myCompanyName, my voice is my password”.
Extended calibration can be performed with these models:
- Text-independent enrollment, verification, and identification operations performed without regard for the spoken words in the audio segments. These operations rely primarily on biometric features and not the content or meaning of the speech. (FACTORY_TI_9)
- Text-dependent common passphrase (FACTORY_TD_COMMON_9, FACTORY_TD_CUSTOM_9)
- Liveness (FACTORY_TI_LIVENESS_9)
- Anti-spoofing algorithm (FACTORY_ANTISPOOFING_9)
Note:
Extended calibration can only be applied to common passphrase text-dependent or text-independent. Extended calibration is not available with FACTORY_TD_UNIQUE_9.To generate a good extended calibration, you need many utterances per speaker under different acoustic conditions, and you must know the speakers’ identity (unless you are using TD_COMMON where speaker label information is optional, but still recommended in order to obtain the best accuracy). Recommended: label the gender of the audio used for extended calibration.
To create an extended calibration:
- Create a dedicated folder for extended calibration audio.
- Create one sub-folder for each speaker.
- Add audio samples to each speaker folder. (Using more samples improves intra-speaker discrimination.) For extended calibration with certain algorithms, the total number of audio samples must be at least 400 more than the number of speakers.
- Calibrate the audio and chose Extended calibration in the calibration type.
- Give the new calibration a meaningful name.
Calibrate the risk engine
Each seed model (text-dependent and text-independent) must be calibrated independently. Given the seed models, the risk engine produced adequate results using a minimum of 18,000 trials. The risk engine can work with unbalanced training data. Although, Nuance recommends using around 6,000 trials per class: authentic, mismatch, and fraud. The collected calibration data must not include trials of factors that are different from the ones supported in NRSL 1.0.
If the deployment uses a subset of factors, contact Nuance for instructions on retraining the calibration model. Although the risk engine supports training the calibration model with missing factors, Nuance recommends including all representative trials for the factor that you plan to use. After calibrating the risk engine, run a small-scale experiment to determine the working thresholds for the three classes: authentic, mismatch, and fraud.