When you start Cantaris for the first time, you will see the license activation screen. A valid license key is required before you can use Cantaris.
Steps
You should have received a license key when you purchased Cantaris (from cantaris-software.com). It is a text string, for example XXXX-XXXX-XXXX-XXXX.
On the activation screen, enter your license key in the text field (it shows a placeholder of the form XXXX-XXXX-XXXX-XXXX).
Click Activate License.
Cantaris will contact the license server to verify your key. If successful, Cantaris moves on to the login screen.
If activation fails, an error message will appear. Double-check that you typed the key correctly and that you have an active internet connection.
Checking or Changing Your License
Once you are on the login screen, you can check or change your license at any time:
Right-click (or middle-click on macOS) the "Licensed until..." text at the bottom of the login screen.
A menu appears with two options:
check — contacts the license server to verify your current license and updates the displayed expiry date if the server has a newer one.
change license — opens a dialog where you can enter a different license key. The new key is activated and the old one is deactivated automatically.
This is useful when you renew your subscription and receive a new key, or if you want to confirm that the server still recognizes your license.
Important Notes
Each license key is tied to the device it is activated on. The activation registers your computer's identity with the license server.
The license has an expiration date. Cantaris displays "Licensed until [date]" on the login screen so you can see when your license expires.
Cantaris periodically re-validates your license online. If you are within 7 days of expiration, Cantaris checks with the server each time it starts (at most once every 24 hours). Outside that window it starts directly without an online check.
If your license has expired, Cantaris will attempt to renew it online at startup. If the server confirms the license is no longer valid, you will be returned to the activation screen where you can enter a new key.
If you are offline and your license has not yet expired, Cantaris will continue to work normally.
Your existing measurement data is never lost when a license expires. When you re-activate (with the same or a new key), all your previous users and data will still be available.
Praat Installation
On first startup, if Cantaris detects that Praat is not installed on your computer, it will offer to download and install it for you automatically. Praat is a free acoustic analysis tool used by Cantaris for advanced voice analysis features such as the Open in Praat function. You can accept the automatic installation or choose to install Praat manually later from praat.org.
After activating your license, you arrive at the login screen. If you are a new user, you need to create an account first.
Steps
On the login screen, type your desired username in the top field (the dropdown/text field).
Type a password in the password field below it.
Click New user.
If registration is successful, a dialog will appear showing your backup code. This is a long text string.
Write down or securely save this backup code. It is the only way to recover your account if you forget your password. It cannot be retrieved later.
Click OK. Your username is now selected in the login field and you can log in immediately.
Important Notes
Usernames must be unique and are case-insensitive — "Alice" and "alice" are treated as the same account.
If a username already exists (either on this computer or in the shared workspace), registration will fail.
Your password protects all your measurement data with strong encryption. There is no way to recover the data without either the password or the backup code.
The backup code is the only spare key to your data. Treat it like a bank PIN — keep it private and safe.
The following functions are all accessed from the Options button on the login screen. Click Options, then select the relevant item in the sidebar on the left. The sidebar lists, in order: Language, Workspace, Import Vox, Functions, Batch export, Change password, Reset password, and Delete account. (Language is selected by default when the Options panel opens.)
Functions is covered in Section 14 (default mode and hiding dashboard items). Batch export opens a wizard with two choices — Export data (CSV) and Export WAVs — for exporting measurement data or audio across all people in one pass.
4. Changing Your Password
Steps
On the login screen, click Options.
Click Change password in the sidebar.
Fill in:
Username — pre-filled with the name you had selected on the login screen.
Current password — your existing password.
New password — the password you want to switch to.
Repeat new password — type the new password again to confirm.
Click Change password.
If successful, a confirmation dialog appears. Your password is now updated.
Important Notes
Changing your password keeps all your existing data accessible — only the password that unlocks it changes.
If you have the workspace on a shared network drive, the password change is synced to the workspace so you can use the new password on any computer that accesses the same workspace.
You can change this to any folder you like — for example, a shared network drive so that multiple computers can access the same data, or a specific local folder of your choice.
Steps
On the login screen, click Options.
Click Workspace in the sidebar.
You will see the current workspace path displayed in a text field.
To change it:
Click Browse... to open a folder picker dialog, then navigate to and select the folder you want.
Or type/paste the path directly into the text field.
Click Apply and close.
Cantaris will save the new workspace location and then close. The next time you start Cantaris, all data operations (person records, measurements, recordings) will use the new location.
Additional Buttons
Open — opens the current workspace folder in your file manager so you can inspect its contents.
Reset — reverts the path back to the default workspace location.
Important Notes
Changing the workspace does not move existing data. If you want to keep your old data, manually copy the contents of the old workspace folder to the new location before switching.
The workspace stores all user data in a data/ subdirectory, organized by user ID. It also contains an accounts/ directory that enables account roaming between computers.
Configuration files, license information, and authentication data stay in the application data folder (separate from the workspace). Only measurement data and person records live in the workspace.
A shared network workspace allows multiple computers to see the same clients and recordings. When you log in on a different machine that points to the same workspace, your account credentials are synced automatically.
If you previously used the Vox software (the predecessor to Cantaris), you can import your existing Vox data into Cantaris.
Steps
On the login screen, click Options.
Click Import Vox in the sidebar, then click Start Import Wizard....
The import wizard guides you through four steps:
Step 1: Log in — select which Cantaris user account will receive the imported data.
Step 2: Select Vox directory — browse to the Vox source directory on your computer. Cantaris validates that it is a genuine Vox data folder.
Step 3: Select measurements — a checklist appears showing all persons and measurements found in the Vox data. Select which ones you want to import.
Step 4: Importing... — Cantaris copies and converts the data. A progress bar and log show what is happening. VRP recordings are split by attempt, SRP recordings are matched to their condition (Normal, Loud, Soft), and all audio is re-analyzed using Cantaris's analysis engine with the original Vox calibration data preserved.
When the import is complete, close the wizard and log in. The imported persons and measurements will be available on the dashboard.
Important Notes
The import does not modify or delete your original Vox data.
Imported data is re-analyzed by Cantaris, so the results may differ slightly from what Vox showed, due to differences in analysis parameters.
Person records, voice classifications, and measurement metadata are carried over.
The following settings are accessed after logging in, via the Settings screen (gear icon) on the dashboard.
10. Setting Up Recording Equipment
Cantaris uses equipment profiles to manage different recording setups. Each profile stores the audio input device, microphone-to-mouth distance, and calibration data. This allows you to switch between setups (for example, a desktop microphone at the office and a head-mounted microphone for fieldwork) without losing calibration.
Accessing Equipment Settings
Log in and open the dashboard.
Click the ⋯ (Options) button and choose Audio and calibration settings.
The left panel shows Equipment & calibration.
Creating a New Equipment Profile
In Settings, click + Add (in the actions row below the equipment dropdown).
A setup dialog appears where you:
Select the audio input device from a dropdown. A live level meter is shown (with RMS and Peak readouts) so you can verify the microphone is working — speak into it and confirm you see the meter respond.
If you just plugged in a new device, click ↻ Refresh to rescan.
Name the profile (the Setup name field, e.g., "Studio headset", "Classroom mic").
Click OK to create the profile. It becomes the active equipment immediately.
A reminder appears to calibrate the new profile before using it for measurements. See Section 12.
Selecting and Switching Equipment
Use the Equipment dropdown in Settings to switch between profiles.
When you select a different equipment profile, Cantaris loads the calibration data associated with that profile. This means each setup maintains its own independent calibration.
Renaming an Equipment Profile
Select the profile you want to rename from the equipment dropdown.
Click Rename….
Type the new name in the dialog and confirm.
Deleting an Equipment Profile
Select the profile from the dropdown.
Click Delete.
Confirm in the dialog that appears. The profile and its calibration data will be removed.
You cannot delete the last remaining profile.
Audio Output Device
In the right panel of Settings (Audio & analysis), you can select the Audio output device. This controls which speakers or headphones Cantaris uses for playback.
If your audio device is not listed or was just connected, click the ↻ Refresh button next to the dropdown to rescan.
Additional Audio Settings
The Audio & analysis panel on the right also offers:
Measure CPPS — a toggle switch to enable or disable CPPS (Cepstral Peak Prominence Smoothed) analysis for voice quality measurement.
Speaking voice CPPS — when CPPS is enabled, a dropdown lets you choose how CPPS is computed for speech: Overall (one value per recording) or Per-frame.
Live CPPS biofeedback — when enabled, the live biofeedback crosshair changes color based on real-time voice quality (see Section 16).
Show spectrogram with live feedback — toggles a live spectrogram strip alongside the biofeedback display.
Show crosshair when recording — toggles the live pitch/loudness crosshair during a recording.
Saving Settings
Click Save to persist all changes (equipment profile, audio output device, and voice quality settings). Click Cancel to discard changes.
The microphone you use has a direct impact on the accuracy and reliability of your voice measurements. While Cantaris works with any microphone, choosing the right one will give you significantly better results — especially for measuring soft phonation and voice quality.
Headset Microphone (Recommended)
A head-mounted (headset) microphone is strongly recommended for two reasons:
Consistent distance. The microphone stays at the same distance from the mouth at all times, even when the person moves their head. With a stand microphone, small head movements change the distance and distort intensity measurements.
Better signal-to-noise ratio. A headset mic is close to the mouth (typically 8 cm), so it picks up much more voice signal relative to background noise. This is critical for measuring the softest sounds accurately. Research has shown that a headset setup with proper calibration produces significantly more accurate phonetographic measurements than stand-mounted alternatives (Nijkamp & Maryn, 2018).
When using a headset at close range (e.g. 8 cm), Cantaris can automatically recalculate intensity values to the standard 30 cm distance for comparison purposes. See the distance recalculation option in Section 12: Calibrating dB Levels.
Microphone Specifications
If you select a different microphone, look for these specifications:
Type: Condenser (more sensitive than dynamic microphones)
Polar pattern: Cardioid (rejects sound from the sides and back)
Frequency range: At least 50 Hz – 20,000 Hz
Frequency response: As flat as possible (minimal colouration of the signal)
Self-noise: As low as possible
Sensitivity: High (around −60 dB or better)
Maximum SPL: 126 dB or higher (to handle loud phonation without distortion)
Room Environment
The recording environment matters as much as the microphone. Background noise reduces the signal-to-noise ratio (SNR), which directly affects measurement accuracy — particularly for soft phonation. Cantaris can actually measure your room's quietness for you; see Section 13.
SNR Level
Quality
Suitable For
≥30 dB
Excellent
All voice measurements including voice quality (CPPS)
15–30 dB
Adequate
Basic pitch and loudness analysis
<15 dB
Insufficient
Not suitable for acoustic voice analysis
Troubleshooting noisy recordings
If your signal-to-noise ratio is too low, check these factors:
Are windows and doors closed?
Is the computer producing audible fan noise?
Is noise coming from adjacent rooms or corridors?
Are the walls, floor, and ceiling made of hard, reflective materials? (Soft furnishings, curtains, and acoustic panels help reduce reverberation.)
Can you move the microphone closer to the mouth? A headset microphone is the most practical solution for noisy environments.
Cantaris includes a Calibration Wizard that maps the raw decibel values from your recording setup to real-world SPL (Sound Pressure Level) values. Calibration is specific to each equipment profile, so if you use multiple setups, each one needs its own calibration.
What You Need
A sound level meter (SPL meter) — a handheld device or smartphone app that measures decibel levels in real time. This provides the "true" reference values.
Your recording equipment (microphone + audio interface) already configured in an equipment profile.
A relatively quiet room.
Opening the Calibration Wizard
Go to Settings.
Make sure the correct equipment profile is selected in the dropdown.
Click Calibrate….
Step 1/3 — Input gain / clipping
The first screen of the wizard (titled Step 1/3 — Input gain / clipping) lets you verify that your microphone is picking up audio correctly and that the signal is not clipping.
Select the correct audio input device from the dropdown (it defaults to the one saved in your equipment profile).
Click Start Meter.
Speak or produce a sustained "aaa" sound at various volumes.
Watch the level meter:
The progress bar shows the current input level.
Below it you see three indicators: RMS (the overall level), Peak (the loudest moment), and CLIP (whether the signal was too loud and got cut off).
If "CLIP: yes" appears, your input level is too high. Reduce the microphone gain on your audio interface, move further from the microphone, or reduce your volume.
The goal is to be able to produce your loudest intended sound without clipping. Normal speaking and singing should show comfortable levels with headroom to spare.
When you are satisfied with the levels, click Stop Meter and then Next.
Step 2/3 — 5× AAAA calibration
This step (titled Step 2/3 — 5× AAAA calibration) collects paired measurements: Cantaris records a short audio sample and computes a raw dB value, and you enter the "true" dB reading from your sound level meter.
Position your sound level meter at the same distance from your mouth as it would be during normal use (or at a known reference distance).
For each trial:
Hold a steady sustained vowel sound ("aaaa") at a consistent volume.
Note the reading on your sound level meter.
Click Record 2.5s in Cantaris. The software records 2.5 seconds of audio and computes a raw dB value. This value appears on screen (e.g., "Raw dB: 62.34").
Type the value from your sound level meter into the Measured dB (external meter) field.
Click Save Trial.
Repeat this process at 5 different volume levels — for example, very soft, soft, moderate, loud, and very loud. The wizard requires at least 5 trials before you can proceed.
As you save trials, they appear in a list showing each raw/true pair.
After collecting at least 5 trials, click Next.
Step 3/3 — Result
The wizard (titled Step 3/3 — Result) computes a linear regression (y = slope × x + intercept) mapping raw values to true dB values and shows:
Status:APPROVED ✅ or REJECTED ❌.
slope and intercept of the calibration line.
mean abs error (the mean absolute error, in dB): the average error across your trials. A value of 1.0 dB or less is required for the calibration to be approved.
If the calibration is rejected (MAE > 1.0 dB), the wizard shows tips:
Reduce background noise.
Keep a consistent distance to the microphone.
Sustain a steady vowel for the full 2.5 seconds.
Make sure the signal is not clipping.
You can go back to Step 2 and re-do the trials.
Optional: Distance Recalculation
At the bottom of the results screen, you can configure distance recalculation:
Enable the checkbox if you want Cantaris to mathematically adjust dB readings to a standard reference distance.
Mic distance (cm): Enter the actual distance from the microphone to the mouth during recordings (in centimeters).
Recalc to (cm): Enter the target reference distance you want values normalized to (e.g., 30 cm is a common reference standard).
This is useful when your microphone is at a fixed short distance (e.g., 8 cm on a headset) but you want to report values as if measured at 30 cm.
Saving the Calibration
If the calibration is approved (or you want to save it regardless), click Save to Settings.
The calibration parameters (slope, intercept, error, date) are saved to the active equipment profile.
The button text changes to "Saved ✅" and the wizard can be closed.
Calibration Validity
Calibration results are shown in Settings with a green checkmark if valid, or a warning icon if expired or high error.
A calibration is considered valid when:
The MAE is 1.0 dB or less.
The calibration is less than 12 months old.
When calibration is approaching 12 months old, a warning appears. Recalibrate periodically to ensure accuracy.
Each equipment profile maintains its own calibration independently. Switching profiles loads the corresponding calibration automatically.
The microphone is half of the story — the room is the other half. A noisy room can hide the softest parts of the voice and pull intensity measurements off. Cantaris can record a short sample and report two numbers that tell you how quiet your recording environment actually is.
Room profiles
Like equipment, rooms are stored as profiles you can switch between. A Room acoustics panel sits in Settings, just below Equipment & Calibration. The dropdown lists every room you've saved. Use Add, Rename, and Delete the same way you would for equipment.
Measuring a room
Pick the room profile you want to measure from the dropdown (or click Add to create a new one).
Click Measure…. A small dialog appears with instructions.
Speak normally for about ten seconds — in short phrases, with brief pauses between them. Cantaris needs both speech and silence in the recording so it can tell them apart.
Click Start. A progress bar shows the remaining time. Stay quiet for the rest of the ten seconds after you stop speaking.
When the recording finishes, Cantaris analyses it and shows the result next to the room name.
What you get back
A short summary appears below the room panel, each line in decibels (for example "- SNR: 38.2 dB ✅"):
SNR (Signal-to-Noise Ratio) — how much louder your speech was than the silent moments.
VNR (Voice-to-Noise Ratio) — the same idea, but only the voiced parts of the speech are compared against the silence. This is the more telling number for voice work, because it ignores consonants and breath.
Reverberation (T60) — an estimate of how long sound lingers in the room.
Try a quieter room or move the microphone closer. See Section 11 for tips.
The badge is the checkmark itself (✅✅ or ✅); the words "outstanding" and "good" above are descriptive only and do not appear in the app.
Important Notes
The measurement is saved with the room profile, so the numbers stick around between sessions.
Re-measure whenever the environment changes — a different room, new fan noise, time of day with more traffic outside, and so on.
If Cantaris can't separate speech from silence in the recording (usually because you spoke without pauses or the room is too noisy), it asks you to try again. Speak in short phrases with clear breaks.
SNR and VNR are about the room, not the voice. They don't affect calibration or measurement values — they just tell you whether the environment is good enough to trust the soft and quiet parts of the voice.
Cantaris has two analysis modes that share the same person list and measurements but each show a different set of recording buttons, analyses, and visualizations on the dashboard:
Voice Range Profiling — the standard mode for VRP, SRP, and Song recordings. Covered in Parts V and VI.
Phonanium Voice Lab — aerodynamic registrations, DSI, AVQI, and perceptual ratings (GRBAS, CAPE-V). Covered in Part VII.
Both modes work on the same measurement for the selected person. You can move between them at any time without losing data — each mode simply shows the buttons and results that belong to it.
Free-form recording, audio editing, and the pre-packaged Praat analysis scripts are available in both modes — through the Recordings sub-tab in the left column and the audio editor (see Part VIII).
14. Switching Between Modes
Using the mode button
On the dashboard, look at the top of the screen. The current mode name is shown on a button (either "Voice Range Profiling" or "Phonanium Voice Lab").
Click the mode button. A small dialog opens showing one card per mode (two cards), each with its logo and a short description.
Click the card of the mode you want to switch to. The dialog closes and the dashboard rebuilds itself for that mode.
Using the keyboard
Press the Left arrow or Right arrow key (when no text field has focus) to cycle to the previous or next mode.
Press Up or Down to step through the measurement list of the current person.
What happens during a switch
If a recording or live biofeedback is running, Cantaris stops it automatically before changing mode.
The selected person and measurement stay the same.
The left, center, and right columns are rebuilt to match the new mode — for example, the right column with the profile toggles is hidden in Phonanium Voice Lab, so the center area can grow.
Setting the default mode
On the login screen, click Options, then Functions in the sidebar.
Choose your preferred default mode. Cantaris will start in that mode the next time you log in.
Hiding parts of a mode you don't use
If you only use a few of the available buttons or visualizations, you can hide the rest:
On the login screen, click Options, then Functions in the sidebar.
The dialog has a section per mode (Voice Range Profiling and Phonanium Voice Lab). Uncheck anything you want to hide.
Click Save. Hidden items disappear from the dashboard the next time it opens.
This part covers the Voice Range Profiling mode (the default Cantaris mode). All recording functions are accessed from the dashboard after logging in. Make sure you have selected a person and measurement before recording, and that the dashboard shows "Voice Range Profiling" on the mode button at the top — if it doesn't, click the button and pick it (see Section 14).
15. Adding a Person
Before you can make recordings, you need to add at least one person to your account.
Steps
On the dashboard, click the add-person button (the + near the person list). A dialog titled New person appears.
The dialog has the following fields:
Full name — the person's name (required).
Date of birth — entered as separate year / month / day fields and used to calculate the person's age. The form marks it as required; if you enter it, it must be a valid date.
Gender — optional field.
Voice classification — optional, for example "soprano", "tenor", etc.
Repertoire (comma-separated) — optional, to note what the person sings or performs.
Coaching goals (comma-separated) — optional free-text notes about what you are working on with this person.
A full name is required; enter the date of birth as well. Click Save to confirm. (An Import case button is also available, to create a person from an exported Cantaris case file.)
Cantaris creates the person record and a first empty measurement. The person is selected automatically and you can begin recording immediately.
Important Notes
Each person's data is stored in an encrypted folder under your user account. Only you (with your password) can access it.
You can edit a person's details later by opening the contextual info dialog from the dashboard.
The equipment profile that is active when you create a person is stored as the default for that person's measurements.
The Live Biofeedback tool gives real-time visual feedback of pitch and loudness while a person is producing voice. It is useful for vocal exercises, pitch training, and monitoring voice production during a session.
Starting Live Biofeedback
On the dashboard, make sure a person and measurement are selected.
Click the Live Biofeedback button on the left sidebar (below the recording buttons).
The button turns red to indicate that live feedback is active and the microphone is listening.
What You See
While live feedback is active, the piano diagram shows a real-time display:
A crosshair marker appears on the piano diagram at the position corresponding to the current pitch (horizontal axis) and loudness in dB (vertical axis). As you speak or sing, the marker moves in real time.
The piano key that matches the current pitch highlights in yellow, giving an immediate visual reference of which note is being produced.
A thin horizontal dB bar is drawn across the diagram at the current loudness level, making it easy to see the intensity even from a distance.
In the top-right corner of the diagram, the current note name (e.g., "C4", "A5") is displayed in large text, with the frequency in Hz and intensity in dB shown below it (e.g., "440.1 Hz • 85.3 dB").
When the voice stops or drops below the noise floor, the marker disappears after a brief moment.
The display updates approximately 50 times per second and uses smoothing to reduce wobble, so the marker movement appears fluid and stable.
Live CPPS Quality Coloring
If you have enabled Live CPPS biofeedback in Settings, the crosshair marker changes color in real time based on voice quality:
Red — poor voice quality (low CPPS).
Yellow — moderate voice quality.
Green — excellent voice quality (high CPPS).
The crosshair colour gives the person instant feedback on whether their voice production is clear and periodic, or breathy and irregular, while they are actively producing voice. Note that the CPPS value itself is not printed in the overlay — only the crosshair colour reflects it; the overlay text shows just the note name, the frequency in Hz, and the intensity in dB.
Area Annotations
If you have drawn area annotations on the piano diagram (rectangular zones marking regions of interest), the live biofeedback interacts with them:
When the crosshair enters an annotated area, the area rectangle highlights in green and the crosshair changes color to match. This gives instant feedback that the voice is within a target zone.
Clicking Piano Keys to Hear a Tone
You can click on any key in the piano keyboard at the bottom of the diagram. When you do:
Cantaris synthesizes a short "ah" vowel tone at the pitch of the key you clicked and plays it through your speakers or headphones.
The key briefly flashes blue to confirm your click.
This is useful for giving a person a reference pitch to match, or for checking what a specific note on the diagram sounds like.
CPPS Explained
CPPS (Cepstral Peak Prominence Smoothed) is a measure of voice quality. A higher CPPS value indicates a clearer, more periodic voice signal, while a lower value may indicate a breathier or less clear voice. It is one of the most reliable acoustic indicators of overall voice quality. For a deeper explanation of how CPPS works, see Section 32.3.
You can enable or disable CPPS analysis in Settings using the Measure CPPS toggle (in the Audio & Analysis panel on the right).
When CPPS is enabled, it is computed for each recorded profile (VRP, SRP, SONG) and can be visualized as a color-coded overlay on the piano diagram after recording.
To view the CPPS overlay, make sure exactly one profile is active (using the profile toggle buttons on the right), then click the Quality (CPPS) metric button in the toolbar. The diagram will show a heatmap where warmer colors indicate higher voice quality.
Stopping Live Biofeedback
Click the Live Biofeedback button again. It returns to its normal white appearance and the microphone stream stops. You can then proceed with recordings or other actions.
The Voice Range Profile (also called the Vocal Performance Range or phonetogram) captures the full range of pitches and loudness levels a person can produce. It maps out what the voice can do from the softest to the loudest and from the lowest to the highest pitch.
Before You Begin
Select the correct person from the person list on the left side of the dashboard.
Select the measurement you want to record into (or create a new measurement using the "Add measurement" option).
Make sure your equipment profile is set up and calibrated in Settings.
Recording Protocol
The following protocol is recommended. Each step is recorded as a separate VRP recording, which makes it easier to review individual attempts and isolate specific regions of the voice range.
Sustained, steady vowel. Click the Record VPR button. Instruction for the person: "Sustain the vowel /a/ on a comfortable pitch and loudness." Click the button again to stop recording.
Low intensity, steady pitch. Click Record VPR again. Instruction: "Try to make the softest /a/ possible, on a comfortable pitch. As soft as you can go without whispering." Stop recording.
Soft glide up. Click Record VPR. Instruction: "Now make the voice glide up, try to find the softest voice possible while gliding to the highest tone." Stop recording.
Soft glide down. Click Record VPR. Instruction: "Now glide downward, try to reach the softest and lowest tone you can make." Stop recording.
High intensity, steady pitch. Click Record VPR. Instruction: "Make the loudest /a/ that is comfortably possible. Try to sustain it for a few seconds." Stop recording.
High intensity, glide to lowest pitch. Click Record VPR. Instruction: "Start out at that loud /a/, but slowly bring it down to the lowest tone. Try to keep it as loud as comfortably possible. Your voice will become softer — that's normal." Stop recording.
High intensity, glide to highest pitch. Click Record VPR. Instruction: "Now that the voice is warmed up, try to start out loud and glide to the highest tone. It may help to do it slowly. Remember not to strain the voice." Stop recording.
Finalize softest voice on high pitch. Click Record VPR. Instruction: "Let's start very high and very soft, see if you can slowly glide the voice down." Stop recording.
Finalize softest voice on low pitch. Click Record VPR. Instruction: "Finally, make a glide to the lowest tone possible, with the softest voice you can make." Stop recording.
This protocol is a recommendation. The best protocol is one adapted to the specific person's needs and abilities. Making separate recordings for each step gives you the flexibility to play back, inspect, and compare individual attempts.
Tips for accurate measurement
Use gliding tones (glissando) rather than isolated sustained notes at each pitch. Research consistently shows that glissando produces more complete and reliable voice range profiles (Lamarche, 2009; Titze, 1994).
Soft phonation often takes 3–5 attempts to reach the true minimum. The softest sound levels can vary by 3–4 dB between attempts. Give the person time and multiple tries — this is why the protocol includes several soft steps.
Loud phonation is more stable, typically varying only about 1 dB between attempts.
Time of day does not significantly affect the Voice Range Profile in healthy voices (Van Mersbergen et al., 1999).
Over a one-year period, pitch range can naturally vary by 1–2 semitones. Keep this in mind when comparing profiles over long intervals.
Validity: The VRP provides a valid representation of vocal performance capabilities in pitch and intensity, provided that the signal-to-noise ratio is adequate and the recording protocol is applied consistently.
How Recording Works
Click the Record VPR button. It turns red with a stop icon to indicate recording is active.
The person produces voice as instructed.
Click the button again to stop. Cantaris saves the audio and immediately runs the analysis.
The results appear as data points on the piano diagram: each analyzed time frame becomes a point at the corresponding pitch (x-axis) and loudness (y-axis).
Repeat for each step of the protocol. All VRP recordings within the same measurement accumulate on the same diagram, building up the full voice range profile.
After Recording
The VRP profile toggle button on the right side of the dashboard activates automatically after a VRP recording. When active (highlighted in blue), the VRP data is visible on the piano diagram.
The statistics panel on the left shows summary values:
F0 range: minimum and maximum pitch in Hz and musical note names, and the total range in semitones.
IL range: minimum and maximum intensity in dB, and the total dynamic range.
Phonographic surface: the area of the voice range polygon, representing the overall vocal capacity.
You can view voice quality metrics by clicking the metric buttons when only VRP is active:
HNR (Harmonicity) — how clear and harmonic the voice sounds. Higher values indicate a clearer voice.
JIT (Irregularity) — how steady the vocal fold vibration is from cycle to cycle. Lower values mean a steadier voice.
Vocal ring (SPR) — the singing power ratio, indicating how well the voice projects and carries.
Quality (CPPS) — an overall voice-quality measure. Higher values indicate a clearer, more focused voice.
Each metric is shown as a color-coded heatmap overlaid on the piano diagram.
Right-clicking on the key measurement values opens a context menu with a Copy option. This copies all displayed statistics as formatted text to the clipboard, making it easy to paste the values into a person's record, an email, or a report.
Clicking an individual recording in the VRP recordings list on the dashboard opens it in the audio editor, where you can play it back, look at the waveform with pitch and intensity contours, generate a per-recording Voice Report, export the audio, or send it to Praat. See Section 30.
The Speech Range Profile captures the pitch and loudness characteristics of connected speech at different effort levels. Unlike the VRP (which uses sustained vowels and glides), the SRP measures how the voice behaves during actual speaking.
Before You Begin
Same as for VRP: select the correct person and measurement, and ensure your equipment is set up and calibrated.
Recording Protocol
The SRP consists of four recording conditions, each with its own button:
Normal speaking voice. Click the SRP Normal button. Instruction for the person: "Please count from twenty to thirty in a normal speaking voice, just as you would use your voice when having a conversation." Click the button again to stop recording.
Soft speaking voice. Click the SRP Soft button. Instruction: "Please count from twenty to thirty as softly as you can, without whispering." Stop recording.
Loud speaking voice. Click the SRP Loud button. Instruction: "Please count from twenty to thirty with a loud voice, just as you would speak if there were a lot of noise in the room." Stop recording.
Shouting voice (optional). Click the SRP Shout button. This recording is optional — always protect the person's voice from straining. Instruction: "Now, shout the sentence: 'Hey, get off there, it's dangerous!'" Stop recording.
Multiple Recordings Per Condition
You can record each condition more than once. For example, you might record Normal twice or Soft three times to capture the voice on different attempts or across a session.
Each new recording is numbered automatically (Normal 1, Normal 2, etc.).
The SRP buttons on the dashboard show a checkmark and count for each condition that has been recorded (e.g., "SRP Normal ✓2"). When only one recording exists for a condition, just a checkmark is shown without a number.
All recordings of the same condition are included in the analysis and shown on the diagram.
How Recording Works
Click the appropriate SRP button (SRP Normal, SRP Soft, SRP Loud, or SRP Shout). The button turns red with a stop icon.
The person speaks as instructed.
Click the button again to stop. Cantaris saves the audio, runs the speech analysis, and displays the results.
Each SRP condition is analyzed separately and shown with its own data on the diagram.
After Recording
The SRP profile toggle button on the right activates automatically (highlighted in green when active).
The SRP buttons on the left update to show which conditions have been recorded and how many times.
The statistics panel shows the SRP summary with separate values for each recorded condition:
Normal: average F0 (Hz), musical note, average IL (dB).
Soft: average F0, note, IL.
Loud: average F0, note, IL.
Shout: average F0, note, IL (if recorded).
You can view the Quality (CPPS) metric for the SRP by clicking the CPPS button when only SRP is active.
Right-clicking on the median speaking voice values opens a context menu with a Copy option. This copies the SRP statistics as formatted text to the clipboard, for easy pasting into external documents.
Clicking an individual recording in the SRP recordings list on the dashboard opens it in the audio editor, where you can play it back, look at the waveform with pitch and intensity contours, generate a per-recording Voice Report, export the audio, or send it to Praat. See Section 30.
Tips for reliable speech profiles
Test-retest reliability is high for most SRP variables (Sanchez et al., 2014; Ma et al., 2007; Hallin et al., 2012). However, the melodic speaking range (pitch range in semitones) and total area can vary between sessions.
Standardise speech material and duration when comparing measurements over time. What you ask the person to say, and for how long, significantly affects the size and shape of the speech profile.
In less flexible or less developed voices, the speech profile often overlaps with or approaches the edges of the voice range profile. In healthy voices, the speech profile sits well inside the voice range with room to spare.
The Song profile records and analyzes a singing performance. This can be used for assessing the singing voice separately from speaking, and is especially useful for voice professionals, singers, and vocal coaching.
Recording
Select the correct person and measurement.
Click the Song button on the left sidebar. It turns red with a stop icon.
The person sings the desired passage.
Click the button again to stop recording. Cantaris saves and analyzes the audio.
After Recording
The SONG profile toggle button on the right activates automatically (highlighted in yellow/gold when active).
The song data appears on the piano diagram, showing each analyzed frame at its pitch and loudness position.
Statistics (F0 range, IL range, phonographic surface) are displayed, similar to the VRP.
When only SONG is active, the following metric buttons are available:
Vocal ring (SPR) — the singing power ratio, showing how well the singing voice projects.
Quality (CPPS) — voice quality during the song.
Clicking an individual recording in the Song recordings list on the dashboard opens it in the audio editor, where you can play it back, look at the waveform with pitch and intensity contours, generate a per-recording Voice Report, export the audio, or send it to Praat. See Section 30.
You can view multiple profiles simultaneously by toggling them on or off:
Vocal Performance Range (VRP) — blue toggle button.
Speech Range Profile (SRP) — green toggle button.
Song Range Profile (SONG) — yellow/gold toggle button.
When multiple profiles are active, their data overlays on the same piano diagram, allowing you to visually compare the speaking voice, singing voice, and vocal range. Deactivate profiles you don't need to focus on specific data.
When exactly one profile is active, the metric buttons become available for that profile:
VRP: HNR (Harmonicity), JIT (Irregularity), Vocal ring (SPR), Quality (CPPS).
SRP: Quality (CPPS).
SONG: Vocal ring (SPR), Quality (CPPS).
Inspecting an individual recording
Each of the recordings that make up a VRP, SRP, or Song profile is listed on the dashboard. Click a recording to open it in the audio editor, where you can listen to it, look at the waveform with pitch and intensity contours, run a per-recording Voice Report, export the audio, or send the take straight to Praat for further acoustic analysis. The full set of editor controls is described in Section 30.
Comparison and Normative Data
Cantaris can compare the current measurement with another measurement or with published normative reference data. Both options are accessed from the options menu on the dashboard.
Opening the Comparison Dialog
On the dashboard, open the options menu.
Click Compare. A checkmark appears next to the menu item while comparison mode is active.
The comparison dialog opens, showing the heading "Compare with current [date]" along with the current person's name.
The dialog offers three mutually exclusive comparison modes:
Mode 1: Compare with Another Measurement (Same Person)
This is the default. The dialog shows other measurements from the current person, listed by date (newest first). Select a measurement and click Compare (or double-click it).
If the current person has only one measurement, the dialog automatically switches to cross-person mode.
Mode 2: Compare with Another Person
Enable the checkbox "Compare with measurements from another person".
A person list appears above the measurement list. Select the person you want to compare with.
The measurement list updates to show that person's measurements. Select one and click Compare.
Mode 3: Compare with Normative Data
Enable the checkbox "Compare to normative data".
The measurement list is replaced by a list of published reference datasets. The following normative VRP contours are available:
Heylen/Wuyts – Female (nonprof.) — Heylen/Wuyts (2002); includes mean, minimum, and maximum contours with a 95% prediction band.
Heylen/Wuyts – Male (nonprof.) — same format as the female dataset.
Healthy Untrained Female (synth.) — a synthesized mean contour based on Hollien, Coleman, and Titze data.
Healthy Untrained Male (synth.) — same approach for male voices.
Andersen 2021 – Young Female — young females aged 18–28, N=39.
Hacki 1999 – Male Choir — nonprofessional male choir singers, N=50.
(The year, N, and age details above come from each dataset's description; the picker shows the short label only, e.g. "Andersen 2021 – Young Female".)
Select a dataset and click Compare. A grey shaded band (and contour lines) is drawn on the piano diagram showing the normative range. The overlay persists across profile toggles and metric views.
What You See in Measurement Comparison Mode
When comparing two measurements, the piano diagram and statistics panel update to show both side by side:
When one profile is active (VRP, SRP, or SONG):
Two overlapping heatmaps are drawn with different color schemes so you can distinguish them:
The current measurement (Date 1) uses blues (VRP), greens (SRP), or warm browns (SONG).
The comparison measurement (Date 2) uses oranges (VRP), purples (SRP), or blue-purples (SONG).
For the SRP, markers for each condition (N, S, L, X) are shown with different shapes: circles for Date 1 and squares for Date 2. Individual SRP recordings are enumerated separately.
When multiple profiles are active:
Instead of heatmaps, Cantaris draws polygon outlines of each profile for each date:
Date 1: solid outlines.
Date 2: dashed outlines.
Each profile uses a distinct color pair (e.g., blue vs. gold for VRP, green vs. magenta for SRP).
This makes it easy to see at a glance how the vocal range has expanded or contracted between the two measurements.
A legend at the bottom of the diagram shows the date labels with color swatches matching the visualization.
Statistics in Comparison Mode
The statistics panel on the left switches to difference display:
F0 range: shows the difference in minimum and maximum pitch (Hz) and range (semitones) between the two measurements, with colored arrows:
Green arrow (▲) = improvement (range widened, max increased, or min decreased).
Red arrow (▼) = regression (range narrowed).
Gray (=) = no significant change.
IL range: same difference display for intensity (dB).
Phonographic surface: difference in the overall vocal area (dB·ST).
SRP medians: for each recorded condition (Normal, Soft, Loud, Shout), the panel shows the median F0 and IL for both dates side by side.
Exporting a Comparison
While comparison mode is active, you can save the comparison as its own PDF. Open the options menu and click Export comparison (pdf). An Export Comparison dialog opens where you choose:
Visualization style — Polygon outlines or Heatmap overlay.
Pages to include — Multi-profile overview and/or Individual profile pages.
Include statistics — add the difference table to the report.
Export anonymously (no name/birthdate) — leave out identifying details.
Exiting Comparison Mode
To leave comparison mode, open the options menu and click Compare again (the checkmark disappears). The diagram returns to the normal single-measurement view. Comparison mode also exits automatically when you switch to a different person or measurement.
Cantaris can generate a comprehensive PDF report of a person's measurement data, suitable for printing or sharing with colleagues.
Exporting a Single Measurement
On the dashboard, select the person and measurement you want to export.
Open the options menu and click Export report (pdf).
An Export settings dialog opens, where you can:
Set a Title and Subtitle for the report.
Choose to Include personal information (Name, Date of birth, Gender, Voice classification, Coaching goals, Repertoire) or to Export anonymously (no name/date of birth).
Pick the Profiles to export — All, or any combination of VRP, SRP, and SONG.
Toggle Show details of recordings to include or omit the per-recording detail pages.
Choose the Page format / layout: A4 (portrait), US Letter (portrait), or Diagram only (no header/summary).
Click OK, then choose a location and filename. Cantaris generates the report and saves it.
The report includes:
Cover page — with the Cantaris logo, person name, measurement date, and basic information.
Vocal Performance Range (VRP) — an overview occurrence heatmap of the full voice range, followed by a 2×2 Voice Quality Maps page (Jitter, HNR, CPPS, and SPR). Individual VRP recordings are shown on subsequent pages with spectrograms and pitch overlays (up to 8 per page).
Speech Range Profile (SRP) — a single overview page built on the CPPS quality map, with the Normal/Soft/Loud/Shout condition markers and a regression line. (Per-recording SRP detail pages are not included.)
Song Profile — an overview heatmap (with a CPPS quality inset) plus individual song recordings with pitch, loudness, and spectrogram details (2 per page).
Profile overlap page — a page that overlays the profiles in pairs (VRP vs SRP, and VRP vs Song) on the piano diagram, color-coded (blue for VRP, green for SRP, gold for Song) so you can directly compare the extent of each profile.
Statistics — a clean summary of all numeric values (F0 range, IL range, phonographic surface, per-condition SRP averages).
Exporting a Voice Lab Report
The Phonanium Voice Lab has its own PDF report. With a measurement that holds Voice Lab data selected, open the options menu and click Export Voice Lab report (pdf). The report carries the Phonanium logo and includes only the sections that actually contain data:
DSI — the four components (MPT, softest intensity, highest F0, jitter) and the composite score.
AVQI — the six components and the composite score.
Aerodynamic — MPT, S/Z Ratio, Phonation Quotient, and any pressure or airflow values you entered.
Auditory-Perceptual — the GRBAS and/or CAPE-V ratings, plus any notes.
Per-recording Voice Report
A detailed acoustic analysis is also available for any single VRP, SRP, or Song recording. Open the recording in its editor (from the recording list on the dashboard), then click the 🧾 Voice report button in the editor's action bar. Cantaris opens an on-screen, scrollable report for that recording showing extended metrics (HNR, jitter, CPPS statistics) alongside published reference values, so the numbers are easier to put in context. This report is generated on demand and is separate from the full measurement export described above.
Cantaris measures scientifically used voice parameters — pitch, intensity, voice quality, aerodynamic measures, and so on — and shows the values alongside published reference points. It is a measurement tool, not a diagnostic tool. Where a score lands relative to a reference line is information for the person reviewing the results; Cantaris doesn't draw conclusions about the voice or label it.
The Phonanium Voice Lab mode collects a set of voice measures that sit alongside the Voice Range Profile: maximum phonation time, the S/Z ratio, the Dysphonia Severity Index (DSI), the Acoustic Voice Quality Index (AVQI), and a perceptual rating panel for GRBAS and CAPE-V. Everything is recorded and saved into the same measurement as the VRP, so you can build a complete picture of the voice from one session.
To enter this mode, click the mode button at the top of the dashboard and choose Phonanium Voice Lab (see Section 14).
22. Phonanium Voice Lab Overview
Layout
When you switch to Phonanium, the dashboard rearranges itself:
The left column leads with a [ Measure | Recordings ] switch. On the Measure tab it shows recording buttons grouped under four headers: Record MPT, Record S/Z Ratio, Record AVQI, and Record DSI, with a Results panel below that fills in as you record. The Recordings tab lists this measurement's recordings (AVQI clips and any free recordings) — see Part VIII.
The center column shows the visualization for the active section: Aerodynamic, AVQI, DSI, or Auditory-Perceptual Evaluation. You pick which one is shown using the four large icons at the top of the center column.
The right column (with the profile toggles in standard mode) is hidden, so the visualization can use the full width.
Choosing a visualization
At the top of the center area, four icons let you switch between Aerodynamic, Dysphonia Severity Index, Acoustic Voice Quality Index, and Auditory-Perceptual Evaluation.
Click an icon to bring its visualization to the front. The recording buttons in the left column stay visible at all times, so you can record while looking at any visualization.
The results panel
The Results panel on the left fills in automatically as soon as recordings are made:
DSI — the composite score (the four components are shown inside the DSI visualization).
AVQI — the composite score (components are shown inside the AVQI visualization).
MPT, S/Z Ratio, Phonation Quotient — the aerodynamic numbers.
Scale and Summary — the active perceptual scale (GRBAS or CAPE-V) with a short summary of the ratings.
Copying results
Right-click anywhere in the Results panel and choose Copy results. All the Voice Lab numbers (DSI, AVQI, aerodynamic, perceptual ratings, and any notes) are copied to the clipboard as plain text, ready to paste into another document.
The Aerodynamic visualization brings together the timing-based measures — Maximum Phonation Time (MPT), the S/Z Ratio, and the Phonation Quotient — alongside optional, manually-entered pressure and airflow measures. The MPT and S/Z recordings are simple time-only registrations: Cantaris measures how long the person can sustain a sound and uses those durations to compute the values.
Selecting the visualization
Click the Aerodynamic icon at the top of the center column. The header reads MPT · S/Z Ratio · Phonation Quotient · Pressure · Airflow. Below it are cards for Maximum Phonation Time (MPT), S/Z Ratio, and Phonation Quotient (VC / MPT), followed by the manual Pressure, Airflow, and Combination cards described below.
Recording MPT
In the left column, find the Record MPT group. There is a single button labelled MPT.
Click MPT. The button turns red and recording starts.
Instruction for the person: "Take a deep breath in, then sustain the vowel /a/ for as long as you can on a comfortable pitch and loudness."
Click the button again as soon as the voice runs out. The duration is added to the MPT card as a new try.
Repeat several times. The MPT card shows each try as a horizontal bar, with the longest try highlighted in green — that is the value used as the MPT.
Recording the S/Z Ratio
In the left column, find the Record S/Z Ratio group. There are two buttons: /s/ and /z/.
Click /s/, ask the person to sustain the unvoiced "ssss" sound as long as possible, then click again to stop.
Click /z/, ask for a sustained voiced "zzzz" sound, and stop in the same way.
Repeat each as often as you like. The S/Z card shows every try; the longest /s/ and longest /z/ are picked automatically and the ratio is computed from them.
The S/Z card shows the computed ratio in green once both /s/ and /z/ have been recorded, and stays muted gray while either side is still missing. Cantaris doesn't label the ratio one way or the other — it just shows what was measured. As a published reference value, an S/Z ratio at or below 1.4 is commonly seen in healthy voices.
Phonation Quotient
The Phonation Quotient card sits at the bottom. It needs the person's Vital Capacity in millilitres.
Click into the Vital Capacity (mL) field, type the value (for example 3500), and press Enter or click outside the field.
Cantaris divides Vital Capacity by the best MPT and shows the result in mL/s. If no MPT has been recorded yet, the card just shows "no tries" until one is.
Pressure & Airflow (manual entry)
Below the timing cards, three more cards let you record aerodynamic values you measured with separate instruments. There are no recordings here — you type the numbers in directly, and they are saved with the measurement and included in Copy results and the Voice Lab report.
Pressure — Subglottal pressure, Pharyngeal/oral pressure, Translaryngeal pressure, and Phonation threshold pressure (all in cm H2O).
Airflow — Peak expiratory airflow and Mean expiratory airflow (L/s).
Type a value into any field and press Enter (or click outside the field) to store it. Leave fields you are not using blank — empty measures are simply omitted from the results and the report.
Tips
You can record multiple tries for each — only the longest is used. Extra tries are kept and shown as smaller grey bars.
If you misclick or get a bad take, switch to a different visualization and back — or simply record again. The score updates automatically each time.
The vital capacity value is saved with the measurement, so you don't need to re-enter it on a return visit.
The DSI is a single number computed from four components: highest F0, softest intensity level, maximum phonation time, and jitter. The DSI visualization shows each component, the resulting score, and a reference line so you can see at a glance where the voice sits relative to it.
Selecting the visualization
Click the Dysphonia Severity Index icon at the top of the center column.
Recording the DSI components
The left column has four buttons under the Record DSI header. Each one records the audio Cantaris needs for one component:
DSI soft — Click the button. Ask the person to make the softest sustained /a/ they can without whispering. Click again to stop. Cantaris analyzes the recording and fills in the soft intensity component.
DSI high — Click the button. Ask for the highest pitch they can comfortably produce on /a/. Click again to stop. The highest F0 value is stored.
DSI vowel — Click the button. Ask for a comfortable, steady /a/ for a few seconds. Stop the recording. Cantaris extracts the jitter value.
DSI MPT — Click the button. Ask for the longest possible sustained /a/. Click again the moment the voice runs out. The duration is stored as the MPT component.
Copying the MPT from the aerodynamic recordings
If you have already recorded MPT in the Aerodynamic section, you don't need to record it a second time:
Right-click the DSI MPT button.
From the menu, click Use Cantaris value.
The longest MPT try from the aerodynamic registrations is copied straight into the DSI MPT slot.
The same right-click menu also offers Copy from Cantaris, Import WAV file…, and Clear value. If no MPT recording exists yet, Cantaris shows a brief warning and nothing is copied.
Reading the visualization
The four component cards (Maximum Phonation Time, Softest Intensity of Voiced Speech, Highest Fundamental Frequency, Jitter PPQ5) each show the recorded value, the unit, and a small bar comparing it to the reference range.
The big number at the top of the screen is the DSI score. A horizontal reference line is drawn at 1.6 — scores at or above that line are commonly seen in healthy voices.
The score is also shown in the left-column Results panel for quick reference.
The AVQI is a multi-parameter voice-quality score combining six acoustic measures (CPPS, HNR, shimmer, shimmer in dB, LTAS slope, and LTAS tilt). It needs two recordings: a sustained vowel /a/ and a short stretch of connected speech.
Selecting the visualization
Click the Acoustic Voice Quality Index icon at the top of the center column.
Recording the two clips
Two buttons live under the Record AVQI header:
Vowel — Click the button. Ask for a sustained, comfortable /a/ for a few seconds. Click again to stop.
Speech — Click the button. Ask the person to read a short standard passage or count from one to ten in a normal speaking voice. Click again to stop.
You can re-record either clip as often as you like; the latest take always replaces the previous one. As soon as both clips exist, Cantaris runs the analysis and the AVQI score appears.
Copying audio from existing Cantaris recordings
If you have already recorded a VRP or SRP for this measurement, you can reuse the audio rather than recording new clips:
Right-click the Vowel or Speech button.
A menu appears with Copy from Cantaris. Click it.
For Vowel, Cantaris uses a vowel segment taken from a VRP recording in the same measurement.
For Speech, Cantaris uses an SRP recording from the same measurement.
If no suitable Cantaris recording exists yet, a short warning is shown and nothing is copied.
Choosing a normative language
The AVQI reference value depends on the language being spoken, and Cantaris draws it as a horizontal line so you can see where the score sits relative to it. At the top of the AVQI visualization there is a Normative data dropdown:
Click the dropdown. The list shows the available languages, each with its reference value next to it (for example "English (3.31)").
Select the appropriate language. The reference line updates immediately on the visualization, and the score's position relative to that line is shown.
The language selection is saved per person, so it stays the same for all of that person's measurements.
Reading the visualization
Each of the six components is shown as a small card with its value and unit.
The large number is the AVQI score, plotted against the reference line for the chosen language. Where the score lands relative to the line is one piece of information among several — Cantaris shows it, it doesn't interpret it.
The score also appears in the left-column Results panel.
Inspecting a Vowel or Speech recording
Click the Vowel or Speech recording in the Recordings list (Voice Lab) to open it in the audio editor. From there you can play the clip back, look at the waveform with pitch and intensity contours, export the audio, or send it to Praat. The vowel clip also offers a per-recording Voice Report. See Section 30 for the full editor description.
The Auditory-Perceptual Evaluation panel lets you record a listener's rating of the voice on either of two well-known scales: GRBAS (Hirano, 1981) or CAPE-V (Consensus Auditory-Perceptual Evaluation of Voice). No recording is involved — this is where you write down what you hear.
Selecting the visualization
Click the Auditory-Perceptual Evaluation icon at the top of the center column.
Switching between GRBAS and CAPE-V
At the top of the perceptual visualization there are two small buttons, GRBAS and CAPE-V. Click the one you want to use. The active scale is highlighted in blue. Switching scales does not delete ratings on the other scale — both are stored independently per measurement.
Rating with GRBAS
Each of the five dimensions (G — Grade, R — Roughness, B — Breathiness, A — Asthenia, S — Strain) has its own card with four numbered buttons: 0, 1, 2, 3.
For each dimension, click the button that matches your impression: 0 Normal, 1 Slight, 2 Moderate, 3 Severe.
The button fills in with a colour (green → yellow → orange → red) and the chosen value is shown in the card header.
Clicking a different number changes the rating. There is no "clear" button — clicking 0 records "Normal" rather than removing the rating.
Rating with CAPE-V
Each of the six dimensions (Overall Severity, Roughness, Breathiness, Strain, Pitch, Loudness) is shown as a single row with a slider that runs from 0 to 100, a numeric value at the right, and a small C/I button next to it.
Drag the slider to the position that matches your impression. The number at the right updates as you drag and changes colour as the rating crosses the mild / moderate / severe boundaries.
Click the C/I button to toggle between C (Consistent) and I (Intermittent). The button fills in blue when set to "I".
A small legend at the bottom of the panel reminds you what the abbreviations mean (MI = Mild, MO = Moderate, SE = Severe, C = Consistent, I = Intermittent).
Adding notes
At the bottom of the visualization there is a Notes text box. Click into it and type any free-text observations. The notes are saved automatically when you click outside the box, and they are included in the Copy results output.
Where the result appears
The Results panel on the left shows the active scale (GRBAS or CAPE-V) and a one-line summary of the ratings — for example G2 R1 B2 A0 S1 for GRBAS, or Ove45 Rou30 Bre50 ... for CAPE-V.
VIII — Recordings, Free Recording & the Audio Editor
The Praat scripts are the same scripts used in voice research. Use these tools when you want to take a short clip of a sound and inspect it more closely, run a Praat script on a recording or a selection, or capture audio that isn't part of a standard VRP/SRP/Song or Voice Lab protocol.
27. The Recordings List
Opening the list
The left column of the dashboard starts with a [ Measure | Recordings ] switch. Click Recordings to see every recording in the active measurement. (You can also reach it by clicking a profile's "Recordings" button.) The center and right columns stay as they are — in Voice Range Profiling the piano diagram remains in the center; in Phonanium Voice Lab the active visualization stays.
What the list contains
The list groups the measurement's recordings under per-profile headers — Voice Range (VRP), Speech Range (SRP), Song, and (in Voice Lab) AVQI — plus a My recordings group:
Voice Range / Speech Range / Song / AVQI — recordings made through the standard protocols appear under these headers automatically; you don't import them by hand.
My recordings — free recordings you captured with the Record button (see Section 28).
Selecting a recording
Click anywhere on a recording's row to open it in the audio editor (see Section 30). A single click opens it directly.
Per-recording actions
Each row has a ⋯ menu (click the ⋯ button on the row). The available actions depend on the kind of recording:
Free recordings (My recordings): Rename, Duplicate, Export (save as a plain WAV), Notes (a free-text notes window saved with the recording), and Delete (with confirmation).
Protocol recordings (VRP/SRP/Song/AVQI): a reduced set — Duplicate, Notes, and Delete only (no Rename or Export from the row menu).
Make sure the right person and measurement are selected, and open the Recordings sub-tab.
Click Record at the bottom of the left column. The button turns red while recording.
The person produces voice. There is no protocol — you record whatever you want to analyze.
Click the button again to stop. The new recording appears under My recordings with a default label like "Recording 1", "Recording 2", etc.
Use the row's ⋯ menu → Rename to give it a more descriptive name if you like.
No manual import needed
Protocol recordings (VRP, SRP, Song, and AVQI) from the active measurement are mirrored into the list automatically — they show up alongside your free recordings without any import step. Opening one in the editor works on a copy, so editing or analyzing it doesn't disturb the original profile data.
The pre-packaged Praat scripts are run from inside the audio editor, via its Analyze ▾ menu.
Running an analysis
Open a recording in the editor by clicking its row in the Recordings list (see Section 30).
Click Analyze ▾. A menu opens with: a scope line at the top (Whole recording, or Selection {start}s – {end}s if you have a region selected), Open in Praat, and then one item per pre-packaged Praat script.
Choose a script. Cantaris hands the audio (or just your selection, if one is active) to Praat in the background.
When the script finishes, a result viewer pops up automatically showing the generated plot.
The pre-packaged scripts are:
Fundamental Frequency (F0) — pitch over time (you can set the F0 range).
Intensity Level — calibrated sound level over time.
Cepstrogram & CPPS — a cepstral view with a CPPS summary.
Spectrogram & Voice Markers — a time–frequency view with voiced/unvoiced markers (F0 range adjustable).
Acoustic Speech Tremor Analysis — tremor across several acoustic dimensions.
Analyses are transient
Run → view → save as PDF
Analysis results are not stored in Cantaris. The result viewer's only action is Export PDF (save the plot to a folder of your choice). Once you close the viewer, the result is gone and nothing is saved on the recording. To keep a result, export it as a PDF; to see it again, just run it again.
Person and calibration data
The Praat scripts embed the person's name and date of birth into the result image, and the intensity script uses your active equipment's calibration so the dB axis matches the rest of Cantaris. None of this needs to be configured — it picks up the values automatically from the selected person and the active equipment profile.
The audio editor is the same window throughout Cantaris. It opens whenever you want to listen to, inspect, or trim a single recording, and it works the same way no matter where you opened it from. The only thing that changes is the row of action buttons at the top, which depends on the type of recording (a VRP/SRP/Song/AVQI protocol take versus a free recording) — the waveform, transport, zoom, and shortcuts are identical everywhere.
Where the editor opens from
In every case, the editor opens with a single click on a recording's row.
From a VRP, SRP, or Song recording — click the recording in the Recordings list on the dashboard.
From an AVQI recording (Vowel or Speech) — click it in the Recordings list (Voice Lab).
From a free recording — click it under "My recordings" in the Recordings list.
What you see
The waveform of the recording fills the middle of the window. Above it, you can see how long the recording is, its sample rate, and whether it is mono or stereo.
A Transport group on the left has Play and Stop buttons.
A Zoom group lets you zoom in (+), out (−), click All to fit the whole recording back into the window, or click Selection to zoom to the highlighted region (enabled once you have a selection).
A Show contours / Hide contours button overlays the pitch and intensity tracks on the waveform so you can see the exact boundaries of a vowel or word. (The contours are also loaded automatically when the editor opens; use this button to hide or re-show them.)
An Analyze ▾ menu hands the audio (or the current selection, if you have one) to Praat — it contains Open in Praat plus one item per pre-packaged Praat script (see Section 29).
A status bar at the bottom shows the cursor position, the current view range, and the active selection.
Selecting a region
Click and drag on the waveform to highlight a region. The start, end, and length of the selection appear in the status bar.
Right-click anywhere on the waveform to bring up a small menu with Select 3 s around cursor (a quick three-second window without dragging) and Clear selection.
Press Esc, or click the ⨯ button next to the selection actions, to clear the selection.
Playing
Press Space or click Play to play from the cursor.
For a free recording, when a selection is active a Play selection button appears in the action bar and plays only the highlighted region.
Click Stop at any time to halt playback.
Action buttons by mode
The buttons just above the waveform change depending on where you opened the editor from.
VRP, SRP, Song, and AVQI recordings
🧾 Voice report — opens an on-screen, scrollable report with extended acoustic measurements and normative reference values for this single recording. (For AVQI, this button only appears for vowel takes — a speech clip has no matching voice report.)
⬇ Export — save the audio as a plain WAV file to a folder of your choice.
🗑 Delete — remove the recording (with confirmation). For VRP/SRP/Song this removes the take from the profile; for AVQI it clears the Vowel or Speech slot.
Analyze ▾ — run a Praat script, or Open in Praat, on the recording or your selection (see Section 29).
Cut selection — appears as soon as you drag a selection. It removes the highlighted region from the recording (handy for trimming away a cough, a throat clear, or a long silence) and joins the audio before and after. After a cut, an ↻ Update profile button appears (for AVQI, ↻ Update AVQI) so you can refresh the analysis on the edited audio.
Free recordings
▶ Play selection — plays only the highlighted region (shown when a selection is active).
Trim to selection — keeps only the highlighted region and discards everything outside it.
Cut selection — removes the highlighted region and joins the parts before and after.
⨯ — clears the current selection.
Analyze ▾ — runs a Praat script on the whole recording, or on just the highlighted region if you have one. The current scope (whole recording vs selection) is shown at the top of the menu; the button label stays "Analyze ▾" either way.
Trims and cuts write a new working copy of the audio — the original on disk is left untouched until Cantaris finishes successfully. If it can't save the new file (for example, the disk is full), it warns you and nothing changes.
Keyboard shortcuts
The same keyboard shortcuts work in every editor: Space play, Esc clear selection, + / − zoom, F fit to window, C toggle contours. For free recordings these are also listed in a hint line in the editor's header; for VRP/SRP/Song/AVQI recordings the hint line isn't shown, but the keys work just the same.
Closing
Close the editor window when you are done — any cuts and trims you made are already saved automatically. Analysis results, however, are not stored: export any you want to keep as a PDF before closing the result viewer (see Section 29).
This chapter explains the science behind voice measurement in plain language. It covers the key terms you will encounter, how Cantaris turns a voice recording into a visual profile, how to interpret the results, and what reference values to use as benchmarks.
31.1 Key Definitions
Voice
The sound produced when the vocal folds (located in the larynx) vibrate rapidly as air from the lungs passes through them. This raw sound is then shaped by the throat, mouth, and nasal cavities into the voice we hear.
Fundamental Frequency (f0)
The base rate at which the vocal folds open and close, measured in Hertz (Hz) — cycles per second. A higher number means a higher-pitched voice. For example, 120 Hz means the vocal folds are opening and closing 120 times per second.
Pitch
How high or low a voice sounds to the listener. Pitch is the perceptual counterpart of fundamental frequency: frequency is what we measure with instruments; pitch is what we hear. They are closely related but not identical.
Sound Intensity (IL)
The physical strength of a sound, measured in decibels (dB). Decibels use a logarithmic scale — every increase of about 10 dB sounds roughly twice as loud to the human ear.
Loudness
How loud or soft a voice sounds to the listener. Loudness is the perceptual counterpart of sound intensity: intensity is what we measure; loudness is what we hear.
Semitones (ST)
A musical unit of pitch. One semitone equals one half-step on the piano (e.g. from C to C#). Voice range is expressed in semitones because it reflects how we perceive pitch changes more accurately than Hertz does.
Phonetogram
A diagram with pitch on the horizontal axis (x-axis, in semitones) and loudness on the vertical axis (y-axis, in decibels). In Cantaris, this is the piano diagram where all voice measurements appear.
Signal-to-Noise Ratio (SNR)
The difference between the volume of the voice signal and the volume of background noise, in decibels. A higher SNR means a cleaner recording. See Section 11 for SNR guidelines.
31.2 How Voice Measurement Works
When you record a voice in Cantaris, two basic acoustic properties are measured in real time:
Vocal fundamental frequency — the rate at which the vocal folds vibrate. This determines pitch. It is affected by the length, tension, and stiffness of the vocal folds.
Vocal intensity level — the strength of the sound wave. This determines loudness. It is affected by subglottal pressure (air pressure beneath the folds), how fully the folds close, and the shape of the vocal tract.
Here is what happens during a recording:
The microphone captures the sound wave and sends it to the computer as a digital signal.
Cantaris analyses small time windows of this signal to determine both the fundamental frequency and the intensity level.
A pitch detection algorithm determines how fast the vocal folds are vibrating (which note is being sung or spoken).
An intensity analysis determines how strong the sound is at that moment (in decibels).
These two values together form a coordinate — a dot on the piano diagram.
Thousands of coordinates are plotted in real time to build the full voice profile.
Why Cantaris uses unweighted (flat) measurement
Many older tools apply an “A-weighting” filter that reduces low-frequency sounds, mimicking how the ear perceives loudness. Research has shown that A-weighting significantly distorts voice measurements, especially for lower (male) voices (Gramming & Sundberg, 1987). It does not give an accurate picture of the actual sound level the voice produces. Cantaris uses unweighted (flat, dB SPL) measurement throughout, which is the scientifically recommended approach.
31.3 Interpreting the Voice Range Profile (VRP)
The VRP maps the absolute limits of the voice: the softest and loudest sounds across every pitch. Research has identified three measures that are the most robust and sensitive for tracking vocal progress:
Softest possible intensity
The quietest sound the voice can produce. This is one of the most sensitive indicators of vocal control. Difficulty producing very soft sounds often reflects tension or less complete vocal fold closure. Changes in soft phonation are often among the first measurable results of voice training.
Melodic range (total pitch range in semitones)
The total number of semitones from the very lowest to the very highest note. A wider range indicates greater vocal fold flexibility and neuromuscular control.
Total phonetographic area (surface)
The total area enclosed in the VRP, measured in semitones × dB. A larger area means more vocal possibilities. This is the most comprehensive single measure of vocal capacity, capturing both pitch range and dynamic range simultaneously.
Tracking progress
These three measures — softest intensity, melodic range, and total area — are the most reliable and sensitive indicators for tracking changes over time. They are robust (not easily affected by small measurement variations) and sensitive (they change meaningfully when the voice changes). Use them as your primary benchmarks when evaluating the effect of training programmes.
31.4 Interpreting the Speech Range Profile (SRP)
Habitual speaking pitch
The median fundamental frequency during normal speech should fall within the expected range for the person’s gender, age, and voice type. As a rule of thumb, a comfortable speaking pitch is typically about 5 semitones above the lowest note in the Voice Range Profile.
Habitual speaking loudness
The median intensity during normal speech should match the expected range for the context and gender. See the reference values below.
Interdependence / overall slope (dB/Oct)
This describes how efficiently the voice converts effort into projection — the relationship between pitch and loudness. In speech, healthy voices show about 9 dB per octave interdependence (Gramming et al., 1988). Across the full voice range, the overall slope should be at minimum 24 dB/Oct, with a target of 37.5 dB/Oct for voice professionals.
Speech area vs. voice range
Compare the speech profile area to the total voice range: a person using only a small fraction of their total range may be underusing their voice; a person whose speech area fills most of the voice range may have limited vocal reserve.
31.5 Normative & Reference Data
The following reference values are drawn from scientific research on healthy, untrained adult voices. Use them as guidelines, not absolute rules — individual variation is normal.
Speaking pitch by speech mode
Speech Mode
Men (avg. pitch)
Women (avg. pitch)
Soft speech
111.8 Hz (A2)
161.3 Hz (E3)
Normal speech
111.9 Hz (A2)
168.5 Hz (E3)
Loud speech
130.2 Hz (C3)
198 Hz (G3)
Calling/shouting
175.5 Hz (F3)
246.2 Hz (B3)
Typical pitch increases: soft to normal is within the same semitone; normal to loud is approximately 3–4 semitones; loud to calling is approximately 5 ST (men) or 4 ST (women).
Speaking loudness (at 30 cm microphone distance)
Measure
Men
Women
Normal speech
65.7–74.4 dB
62–71.2 dB
Older adults
60.2–72.7 dB
60.8–71.7 dB
If you record at a different distance, Cantaris can recalculate to the 30 cm standard. See the distance recalculation option in Section 12.
Vocal efficiency
Measure
Reference Value
Overall slope across the voice range
Minimum 24 dB/Oct
Target for voice professionals
37.5 dB/Oct
Speech interdependence (Gramming et al.)
~9 dB/Oct (range: 3.5–15 dB/Oct)
Voice Range Profile reference values
Measure
Men
Women
Softest possible phonation
≤64 dB at 30 cm
≤61 dB at 30 cm
Total melodic range
≥21 semitones
≥21 semitones
Total phonetographic area
≥1040 ST×dB
≥950 ST×dB
Quick-reference summary
What to Measure
What It Tells You
Healthy Benchmark
Habitual speaking pitch
Is the natural speaking voice at a comfortable pitch?
~5 ST above lowest possible note
Habitual speaking loudness
Is the natural speaking volume appropriate?
60–75 dB at 30 cm
Overall slope
How efficiently does the voice project?
Min. 24 dB/Oct; target 37.5 dB/Oct
Softest phonation
Vocal control and fold closure quality
≤64 dB (M) / ≤61 dB (F)
Melodic range
Vocal fold flexibility
≥21 semitones
Total phonetographic area
Overall vocal capacity
≥1040 ST×dB (M) / ≥950 ST×dB (F)
A word about norms
These values represent the lower boundary for healthy untrained voices. Trained singers and voice professionals often exceed them significantly. Conversely, someone just starting voice training may fall below them. The most valuable comparison is always a person’s own measurements over time.
31.6 References
Gramming, P. & Sundberg, J. (1987). Spectrum factors relevant to phonetogram measurement. STL-QPSR, 28(2-3), 39–61.
Gramming, P., Sundberg, J., Ternström, S., Leanderson, R., & Perkins, W. (1988). Relationship between changes in voice pitch and loudness. Journal of Voice, 2(2), 118–126.
Hallin, A. E., Fröst, K., Holmberg, E. B., & Södersten, M. (2012). Voice and speech range profiles and Voice Handicap Index for males — methodological issues and data. Logopedics Phoniatrics Vocology, 37(2), 47–61.
Lamarche, A. (2009). Putting the Singing Voice on the Map. Doctoral dissertation.
Ma, E., Robertson, J., Radford, C., Vagne, S., El-Halabi, R., & Yiu, E. (2007). Reliability of Speaking and Maximum Voice Range Measures in Screening for Dysphonia. Journal of Voice, 21(4), 397–406.
Nijkamp, M. & Maryn, Y. (2018). Comparing the accuracy of phonetograph systems.
Sanchez, K., Oates, J., Dacakis, G., & Holmberg, E. B. (2014). Speech and voice range profiles of adults with untrained normal voices. Logopedics Phoniatrics Vocology, 39(2), 62–71.
Spe, R., Wieneke, G. H., Wijck-Warnaar, I., & Dejonckere, P. H. (2004). Effects of Voice Therapy as Objectively Evaluated by the Voice Range Profiles of Dysphonic Patients. Journal of Voice, 17(4), 544–556.
Švec, J. G., Granqvist, S., Herbst, C. T., & Šrámková, H. (2015). The softest sound levels of the human voice in normal subjects. JASA, 137(1), 407–418.
Titze, I. R. (1992). Acoustic Interpretation of the Voice Range Profile (Phonetogram). JSLHR, 35(1), 21.
Titze, I. R. (1994). Workshop on Acoustic Voice Analysis. National Center for Voice and Speech.
Van Mersbergen, M. R., Verdolini, K., & Titze, I. R. (1999). Time-of-day effects on voice range profile performance. Journal of Voice, 13(4), 518–528.
This chapter describes the four acoustic voice-quality measures in Cantaris: HNR, Jitter, CPPS, and SPR. Each section explains the underlying acoustic concept, the scientific rationale for its measurement, and how Cantaris derives the value from recorded audio.
All measurements are computed using the Praat acoustic analysis engine (Boersma & Weenink, 2024) and are reported on a per-frame basis (50 ms time windows), allowing spatial mapping across pitch and intensity in the Voice Range Profile (VRP) and Song Profile (SNG).
32.1 Harmonics-to-Noise Ratio (HNR)
What it measures
The Harmonics-to-Noise Ratio quantifies how much of the acoustic energy in a voice signal is carried by periodic (harmonic) vibration versus aperiodic (noise) energy. It is expressed in decibels (dB).
A voiced sound produced by the vocal folds is never perfectly periodic. Every voice cycle contains small irregularities — air turbulence, mucus on the vocal folds, slight asymmetries in how they vibrate — that show up as noise in the signal. HNR captures the ratio between the structured, repeating part of the waveform (the harmonics) and these irregular components (the noise).
Acoustic interpretation
HNR value
Interpretation
> 20 dB
Clear, well-supported voice with strong harmonic structure
10–20 dB
Moderate breathiness or mild irregularity
< 10 dB
Substantial noise component — often perceived as breathiness or reduced clarity
A higher HNR indicates a “cleaner” voice signal. In voice coaching and training, low HNR values draw attention to a breathier, less efficient tone — for example, incomplete vocal fold closure or excess air escape — and can help guide exercises aimed at achieving a more focused, well-supported tone.
How Cantaris computes it
The software creates a Harmonicity object from the audio signal using the auto-correlation (ac) method. This algorithm works by correlating a section of the waveform with a time-shifted copy of itself. The height of the autocorrelation peak (relative to the energy at zero lag) directly reflects how periodic the signal is.
Parameters used:
Time step: 0.01 s — the interval at which new Harmonicity values are calculated.
Pitch floor: 65 Hz — the lowest fundamental frequency expected, which determines the length of the analysis window (the algorithm needs at least a few complete pitch periods).
Silence threshold: 0.1 — frames with amplitude below this fraction of the global maximum are treated as silent and excluded.
Periods per window: 4.5 — the number of pitch periods included in each analysis window, controlling the trade-off between temporal resolution and estimation reliability.
For each 50 ms analysis frame, Cantaris queries the mean Harmonicity value over that time interval, yielding an HNR in dB. In Voice Report mode, Cantaris additionally reports per-frame statistics (mean, median, minimum, maximum) using the cross-correlation (cc) variant of the algorithm.
Visualization
HNR is displayed on a scale of 0–30 dB using a red-yellow-green colour gradient, where green represents higher (better) HNR values.
Availability by profile
VRP (Voice Range Profile): Computed and displayed.
SNG (Song Profile): Not shown as a heatmap metric in the Song profile (only CPPS and SPR are). It is still computed in the per-recording Song Voice Report.
SRP (Speaking Range Profile): Not computed (speaking tasks use CPPS instead).
32.2 Jitter (Period Perturbation)
What it measures
Jitter quantifies the cycle-to-cycle variability in the duration of consecutive vocal fold vibrations. Where HNR captures the overall ratio of periodic to aperiodic energy, jitter zooms in on the microstructure of the pitch periods themselves: are successive vibration cycles nearly identical in length, or do they fluctuate?
It is expressed as a ratio (displayed as a percentage in Cantaris, typically between 0% and 2%).
Acoustic interpretation
Every vocal fold vibration cycle is slightly different from the last. In a well-controlled voice, these differences are very small — typically below 1%. As irregularity increases, the voice begins to sound rough or unstable. Monitoring jitter over time allows voice coaches and singing teachers to track whether vocal control and consistency are improving with training.
Jitter value
Interpretation
< 0.5%
Very stable phonation
0.5–1.0%
Normal range for most speakers
> 1.0%
Elevated irregularity — typically perceived as roughness or instability
How Cantaris computes it
The measurement proceeds in two steps:
Step 1 — Identify glottal pulses. The software creates a PointProcess object using the periodic cross-correlation (cc) method. This detects the moments at which each individual glottal cycle begins, based on the periodicity of the signal. Detection is bounded by a pitch floor (65 Hz) and a pitch ceiling (which varies by profile, e.g. 1600 Hz for singing voice).
Step 2 — Calculate period irregularity. From the detected pulse times, Cantaris derives the duration of each vocal fold cycle (the period). It then computes the local jitter for each pair of consecutive periods:
Periods shorter than 0.1 ms or longer than 20 ms are excluded as physiologically implausible. The final jitter value for a frame is the mean of these per-cycle irregularity ratios. In Voice Report mode, Cantaris additionally provides the median, minimum, and maximum across all cycles in the selected time range.
Visualization
Jitter is displayed on a scale of 0–2% using a reversed red-yellow-green colour gradient (lower values are green, higher values are red), reflecting that lower jitter indicates more stable, well-controlled phonation.
Availability by profile
VRP: Computed and displayed.
SNG: Not shown as a heatmap metric in the Song profile (only CPPS and SPR are). It is still computed in the per-recording Song Voice Report.
SRP: Not computed.
32.3 Cepstral Peak Prominence Smoothed (CPPS)
What it measures
CPPS quantifies the strength of the harmonic structure of a voice signal by analysing it in the cepstral domain. It is expressed in decibels (dB) and is widely regarded as one of the most robust acoustic correlates of overall voice quality (Maryn et al., 2009; Heman-Ackah et al., 2003).
The cepstrum: a brief explanation
The cepstrum is the result of taking the spectrum of a spectrum. When a voiced signal is transformed into the frequency domain (via FFT), its spectrum shows a series of harmonic peaks at integer multiples of the fundamental frequency. Taking the inverse Fourier transform of this log-spectrum (or equivalently, computing the “spectrum of the spectrum”) yields the cepstrum — a function of quefrency (a time-like variable). In the cepstrum, the regularly spaced harmonic peaks in the spectrum collapse into a single peak at a quefrency corresponding to the fundamental period (1/F0). This peak is called the cepstral peak.
Cepstral Peak Prominence (CPP) is the height of this cepstral peak above a regression line fitted through the cepstrum. The smoothed variant (CPPS) applies time and quefrency averaging to reduce frame-to-frame variability, yielding a more stable and reliable estimate.
Acoustic interpretation
A strong cepstral peak indicates that the voice has a clear harmonic structure — the vocal folds are vibrating in a regular, periodic fashion, and the harmonics are well-defined. A weak or absent cepstral peak suggests that the harmonic energy is diminished relative to noise, as occurs in breathy or unfocused voice production.
CPPS value
Interpretation
> 10 dB
Strong harmonic organisation; clear, efficient voice production
5–10 dB
Moderate harmonic structure
< 5 dB
Weak harmonic organisation; voice likely perceived as breathy or unfocused
CPPS has been shown to correlate well with perceptual ratings of voice quality and is considered a reliable, objective measure. For voice coaches and singing teachers, it provides a valuable way to track improvements in vocal efficiency and clarity over the course of a training programme.
How Cantaris computes it
The software creates a PowerCepstrogram from the audio signal — a time-series of power cepstra computed at regular intervals. The parameters are:
Pitch floor: 65 Hz (determines the maximum quefrency searched).
Time step: 0.002 s (the interval between successive cepstral frames in the cepstrogram).
Maximum frequency: 5000 Hz (the upper frequency boundary of the analysis).
Per-frame CPPS is then extracted by:
Slicing the PowerCepstrogram at the centre of each 50 ms analysis frame.
Smoothing the resulting PowerCepstrum (quefrency smoothing window: 0.001 s).
Computing the peak prominence — the height of the cepstral peak above a regression line fitted through the smoothed cepstrum (parabolic interpolation, robust fit method).
Overall CPPS (used in Voice Reports for the Speaking Range Profile) is computed using Praat’s built-in “Get CPPS” command, which applies both time averaging (0.01 s window) and quefrency averaging (0.001 s window) across the entire signal and returns a single global value. For maximum accuracy, this overall value is computed on voiced-only portions of the signal (unvoiced segments and pauses are excluded).
Configuration
CPPS is an optional measurement in the VRP and SNG profiles, controlled by a configuration toggle (“Measure CPPS”). When enabled, CPPS values are computed for every analysis frame and stored in the measurement matrix alongside the other metrics. In the SRP profile, CPPS is always computed, as it serves as the primary voice-quality indicator for speaking tasks.
The software also supports live CPPS feedback during recording, providing real-time visual feedback of cepstral peak prominence to the voice coach or singing teacher.
Visualization
CPPS is displayed on a scale of 0–20 dB using a red-yellow-green colour gradient (higher values are green).
Availability by profile
VRP: Optional (configurable).
SRP: Always computed (primary quality measure for speaking).
SNG: Optional (configurable).
32.4 Singing Power Ratio (SPR)
What it measures
The Singing Power Ratio quantifies the relative energy balance between the lower and upper parts of the voice spectrum. Specifically, it compares the spectral peak in the 2000–4000 Hz region to the spectral peak in the 0–2000 Hz region. It is expressed in decibels (dB) and is always negative (or, in exceptional cases, zero).
The concept is rooted in the phenomenon of the singer’s formant — a clustering of the third, fourth, and fifth formants around 2500–3500 Hz that gives the trained singing voice its characteristic ring, carrying power, and ability to project over an orchestra (Sundberg, 1974, 2001).
Acoustic interpretation
When a singer engages the vocal tract configuration that produces a strong singer’s formant, more acoustic energy is concentrated in the 2–4 kHz band. This causes the SPR to become less negative (closer to zero). An untrained or breathy voice typically has very little energy in this region relative to the lower harmonics, resulting in a more negative SPR.
SPR value
Interpretation
−5 to −15 dB
Strong spectral energy in the singer’s formant region; characteristic of trained, projected voice
−15 to −25 dB
Moderate upper-spectrum energy
−25 to −40 dB
Weak singer’s formant; typical of untrained or breathy voice
In Cantaris, SPR is labelled “Vocal Ring” to make the concept more intuitive for practitioners and voice teachers.
Scientific reference
The SPR measurement follows the methodology described by Omori et al. (1996) in the Journal of Voice, 10(3). Their convention defines SPR as:
Because the fundamental frequency and its low harmonics typically dominate the spectrum, the high-band peak is almost always lower than the low-band peak, yielding a negative value.
How Cantaris computes it
For each 50 ms analysis frame:
The audio segment is extracted from the signal using a rectangular window.
An FFT spectrum is computed from the segment.
The spectrum is converted to a Long-Term Average Spectrum (LTAS) at full harmonic resolution (“1-to-1” conversion, preserving all spectral detail).
The spectral peak in the low band (50–2000 Hz) is identified.
The spectral peak in the high band (2000–4000 Hz) is identified.
The SPR is calculated as the difference: high-band peak minus low-band peak.
The low band starts at 50 Hz (rather than 0 Hz) to avoid DC offset artefacts. The use of LTAS at 1-to-1 resolution (rather than a smoothed or averaged spectrum) ensures that individual harmonics are preserved, giving an accurate representation of the peak energy in each band.
Visualization
SPR is displayed on a scale of −40 to −5 dB using a red-yellow-green colour gradient. Less negative values (closer to −5 dB) are shown in green, indicating stronger vocal ring.
Availability by profile
VRP: Computed and displayed.
SNG: Computed and displayed.
SRP: Not computed (not relevant for non-projected speaking tasks).
32.5 Summary of Measurements Across Profiles
Measure
Description
Unit
VRP
SRP
SNG
HNR
Harmonics-to-Noise Ratio
dB
Yes
—
Report only*
Jitter
Cycle-to-cycle period irregularity
%
Yes
—
Report only*
CPPS
Cepstral Peak Prominence Smoothed
dB
Optional
Always
Optional
SPR
Singing Power Ratio (“Vocal Ring”)
dB
Yes
—
Yes
*In the Song profile, only CPPS and SPR are mapped onto the heatmap. HNR and jitter are not part of the Song profile matrix, but they are still computed for the per-recording Song Voice Report.
Why these four measures?
These four metrics together provide a comprehensive picture of voice quality from complementary acoustic perspectives:
HNR and CPPS both assess the overall periodicity and harmonic organisation of the voice, but through different mathematical lenses (time-domain autocorrelation vs. cepstral analysis). CPPS is generally considered more robust to variation in fundamental frequency and intensity, which is why it is the measure of choice for speaking tasks.
Jitter captures the microstructure of vocal fold vibration at the individual cycle level, detecting instability that may not be apparent in the broader HNR or CPPS measures.
SPR addresses the spectral distribution of energy — a dimension of voice quality that relates to resonance, vocal tract tuning, and projection rather than to vocal fold regularity per se.
Together, they enable voice coaches, singing teachers, and vocal trainers to evaluate whether a voice is periodic (HNR, CPPS), stable (Jitter), and resonant (SPR) — the three pillars of efficient, well-produced voice. By tracking these measures over time, practitioners can objectively document a student’s or performer’s vocal development and tailor training accordingly.
32.6 References
Boersma, P., & Weenink, D. (2024). Praat: Doing phonetics by computer.https://www.praat.org
Heman-Ackah, Y. D., Michael, D. D., & Goding, G. S. (2003). The relationship between cepstral peak prominence and selected parameters of dysphonia. Journal of Voice, 16(1), 20–27.
Maryn, Y., Roy, N., De Bodt, M., Van Cauwenberge, P., & Corthals, P. (2009). Acoustic measurement of overall voice quality: A meta-analysis. Journal of the Acoustical Society of America, 126(5), 2619–2634.
Omori, K., Kacker, A., Carroll, L. M., Riley, W. D., & Blaugrund, S. M. (1996). Singing power ratio: Quantitative evaluation of singing voice quality. Journal of Voice, 10(3), 228–235.
Sundberg, J. (1974). Articulatory interpretation of the “singing formant.” Journal of the Acoustical Society of America, 55(4), 838–844.
Sundberg, J. (2001). The Science of the Singing Voice. Northern Illinois University Press.