Why sound is more than a single frequency
A pure sine wave can be described largely by frequency, amplitude and phase. A human voice is not a pure sine wave. Vocal-fold vibration produces a fundamental frequency and a family of harmonics, while the vocal tract reshapes that spectrum through resonances. Breathing, articulation, microphone position, room acoustics and recording electronics all become part of the captured signal.
That is why a useful voice-analysis tool should not reduce a recording to one mystical number. The more defensible approach is to preserve several separate descriptors and ask what each one actually represents.
What a microphone measures
A smartphone microphone converts changing air pressure into a digital waveform. The resulting samples are measurements of the recording chain, not direct measurements of the vocal folds, brain, autonomic nervous system or emotional state. Even when an acoustic feature is associated with a clinical or physiological phenomenon in research, that association does not transform the microphone into a diagnostic sensor for that phenomenon.
Recording conditions matter. Distance from the microphone, room reflections, automatic gain control, codec processing and background noise can all change calculated values. A 2025 systematic review and meta-analysis of smartphone recordings found that smartphones can be useful for acoustic voice-quality assessment, but the literature also highlights the importance of acquisition conditions and device differences.
Fundamental frequency: the repeating part of voiced sound
The fundamental frequency, usually written as F0, is an estimate of the dominant repetition rate of voiced vibration. It is closely related to perceived pitch, but the two are not identical. A recording can report an F0 median, mean, range and variability.
For longitudinal work, stability may be more useful than a single absolute value. If the same person records the same sustained vowel or phrase under similar conditions over weeks, changes in F0 distribution can become a reproducible research variable. The comparison should remain within the same protocol whenever possible.
Jitter and shimmer: small cycle-to-cycle variations
Jitter describes short-term variation in the period of vocal cycles; shimmer describes short-term variation in amplitude. Both have a long history in acoustic voice analysis. They are sensitive to analysis method, segmentation, signal quality and phonation task.
That sensitivity is important. A number should not be treated as a universal health score merely because it has a clinical-sounding name. A broad review of healthy-adult voice parameters found substantial dependence on demographic and methodological factors. Modern reviews also emphasise that perturbation measures are only one part of a wider acoustic assessment.
HNR: harmonic structure versus noise
The harmonics-to-noise ratio (HNR) estimates how strongly periodic harmonic structure stands above non-periodic components. The implementation matters: different algorithms, windows and frequency ranges can produce different values.
In a research app it is therefore better to expose HNR as a labelled estimate than to silently convert it into “voice health”. Longitudinal consistency under one protocol is often more informative than comparing a single measurement with a threshold taken from another laboratory.
CPP: a cepstral way to describe periodic organisation
Cepstral Peak Prominence (CPP) measures the prominence of periodic organisation in the cepstral domain. It has become an important acoustic voice-quality measure and is generally considered more robust than some traditional perturbation measures, especially for connected speech. Recent systematic reviews nevertheless point to methodological heterogeneity and the need for standardised recording and analysis protocols.
PICALLW Voice Studio therefore labels CPP as a research estimate. The value can support within-protocol comparison, but the app does not claim a clinical diagnosis from it.
Spectrum: where acoustic energy is distributed
The Fourier transform decomposes a time-domain signal into frequency components. From that spectrum we can calculate descriptors such as spectral centroid, rolloff, flatness, entropy, flux and slope. These features say something about how energy is distributed, how tonal or noise-like a signal is, and how the spectrum changes over time.
They are widely useful in audio analysis because they do not require us to know why a spectrum looks a certain way. They describe the signal first; interpretation comes later.
Formants: resonances of the vocal tract
Formants are spectral resonances strongly influenced by the geometry of the vocal tract. F1 and F2 are particularly important in vowel acoustics. Linear predictive coding (LPC) is one common method for estimating them.
Formant estimation can fail when the signal is noisy, when the sound is not vowel-like, or when analysis settings do not match the speaker. This is why Voice Studio marks its LPC formants as approximate rather than presenting them as guaranteed anatomical measurements.
MFCC: compact descriptors of spectral shape
Mel-frequency cepstral coefficients (MFCCs) compress spectral shape into a small set of coefficients. They are widely used in speech and audio classification because they provide a compact numerical representation that can be compared statistically or supplied to machine-learning models.
An MFCC coefficient has no simple standalone meaning such as “stress” or “emotion”. Its value is in patterns across coefficients, recordings and labelled datasets. Without a validated model and an appropriate dataset, MFCCs should remain descriptive features rather than psychological conclusions.
From resonance to Chladni patterns
Cymatics is a broad popular term for making vibration patterns visible. The classical physical example is the Chladni plate: a thin plate is driven near one of its natural modes, and grains move across the surface until characteristic patterns emerge around nodal regions. Different resonant modes produce different nodal geometries.
The basic phenomenon is real classical mechanics. Modern research continues to study the detailed motion of grains on vibrating plates and membranes. Importantly, particle behaviour can depend on acceleration, particle size, surrounding fluid and other conditions; the familiar “sand always goes to the nodes” picture is useful but not universal in every physical regime.
A simulation is not the experiment
A phone app cannot infer the exact motion of a real plate merely from a microphone frequency. A real Chladni pattern depends on plate geometry, boundary conditions, elastic properties, thickness, drive position, drive amplitude, damping and the granular medium.
A useful simulation must therefore expose at least some of those assumptions. PICALLW Voice Studio lets the user select or adjust the plate model, physical size, thickness, damping, excitation location and particle behaviour. The resulting image is a computational visualisation of that model — not a photograph of a physical plate.

Why microphone-driven cymatics is still interesting
Even with that limitation, microphone-driven visualisation can be useful. A live voice contains several harmonics and a changing spectrum. Mapping parts of that spectrum into a vibrating-plate model gives the user a visual way to explore resonance, mode structure and spectral change.
The scientific value is not that the pattern reveals a hidden human “energy”. The value is that one can define a signal-processing rule, keep the plate model fixed, repeat the same vocal task and then compare how the model responds to measurable changes in the sound.
Measured, calculated, modelled and interpreted
| Layer | Example in Voice Studio | What it means |
|---|---|---|
| Measured | PCM/WAV microphone signal | The actual sampled audio waveform from the recording chain. |
| Calculated | F0, spectrum, MFCC, HNR, jitter, shimmer, CPP | Mathematical descriptors derived from the waveform. |
| Modelled | Resonant plate and particle pattern | A simulation using selected physical assumptions. |
| Interpreted | “This pattern means stress” | A higher-level conclusion that requires independent evidence. |
Why repeated measurements are more interesting than one score
The strongest research direction for Voice Studio is longitudinal rather than diagnostic. A standardised 20–40 second voice protocol can be repeated alongside other measurements. Instead of asking whether one recording reveals a person's condition, we can ask whether the person's own acoustic features change reproducibly over time.
This creates testable questions: Does F0 variability change after a controlled breathing task? Does CPP remain stable across days? Are changes in acoustic stability associated with HRV changes when recording conditions are held constant? If Voice Studio is later paired with GDV Studio or Polar H10 data, those relationships should be treated as hypotheses until they survive repeated and independent validation.
What a repeatable protocol should control
- same phone and microphone when possible;
- same microphone distance and orientation;
- same room or comparable acoustic environment;
- same vocal task and approximate loudness;
- same analysis version and settings;
- recording quality checks before interpretation;
- several repeated sessions rather than one unusual sample.
This may sound less spectacular than claiming that a voice pattern reveals the whole person, but it produces much better data.
Where PICALLW Voice Studio fits
PICALLW Voice Studio is an Android research and experimentation app built around four connected workflows: microphone-driven visualisation, a frequency generator, recordings, and Voice Lab analysis. It keeps the cymatics model and acoustic analysis in one environment so the same signal can be heard, visualised, saved and measured.
Public Android builds will be distributed through the official PICALLW download page.
Explore the application and public releases.
Sources and further reading
- The Physics Classroom — standing waves and Chladni plate patterns
- Abramian et al. — Chladni patterns and diffusion of bouncing grains (Physical Review Research, 2025)
- Teixeira et al. — healthy adult voice baseline parameters
- Systematic review and meta-analysis of smartphone recordings for acoustic voice assessment
- Heman-Ackah et al. — cepstral peak prominence and voice quality
- Systematic review of CPP/CPPS normative values and methodological standardisation