Metaphors We Listen With (timbre semantics)
2021–present
Timbre is notoriously difficult to define, and even harder to talk about. Yet people do talk about it, especially musicians, composers, producers, instrument makers, and with surprising consistency. Previous studies have converged on a small number of recurring semantic dimensions, which can be interpreted broadly in terms of brightness/sharpness (or luminance), roughness/harshness (or texture), and fullness/richness (or mass). Saitis and Weinzierl (2019) provide a comprehensive review of the field.
These semantic descriptions of timbre embody conceptual representations, allowing listeners to talk about subtle acoustic variations through other, more commonly shared corporeal experiences—metaphors we listen with. The luminance-texture-mass (LTM) model describes how listeners across different languages tend to reach for the same kinds of metaphor.

Disembodied Timbres: a study on semantically prompted FM synthesis
Most of what we know about what is called “timbre semantics” comes from studies using acoustic orchestral instruments. Hayes, Saitis, and Fazekas (2022a) asked whether the same conceptual vocabulary applies to sounds with no recognisable physical source: the “disembodied” timbres of digital synthesis.
In a novel experimental paradigm, experienced sound designers programmed an FM synthesiser in response to semantic prompts, and provided semantic ratings on the sounds they created. We collected 1,407,604 publicly available posts from a popular synth forum, and looked for adjectives co-occuring with the terms sound, sounding, tone, and timbre. An initial list of 96,277 adjectives were independently pruned by two raters down to a list of 27 unipolar semantic scales, including “bright,” “thick” and “rough” selected as synthesis prompts.
Exploratory factor analysis of the semantic ratings recovered five dimensions. The first two broadly echoed the LTM model: luminance and texture merged into a single “sharpness” factor, while mass appeared as a second independent factor. Three additional dimensions emerged, namely clarity, percussiveness, and rawness, which appear to reflect specific qualities of FM timbres that listeners discriminated.
Semantic prompts left measurable imprints on synthesiser controls. Participants tuning for “brightness” and “roughness” made similar adjustments to the modulator parameters, the FM settings that govern the distribution of spectral energy, consistent with the acoustic correlates observed for these descriptors. “Thickness”, by contrast, was pursued through the amplitude envelopes, shaping the sustain and temporal decay of the sound. Word affect also left its mark: adjectives with stronger emotional valence (positive or negative) tended to produce sounds with more energy in the higher frequencies and greater inharmonicity, echoing earlier findings.
Perceptual and semantic scaling. Using FM sounds created in the prompted synthesis study, and the same 27 unipolar semantic scales, Hayes, Saitis, and Fazekas (2021) further collected standard pairwise dissimilarity and semantic ratings. Multidimensional scaling of the dissimilarity data revealed that musicians and synthesiser-experienced listeners perceptually organised the sounds differently from non-experts, suggesting that prior experience shapes the perceptual structure of electronic timbres in ways not typically seen with acoustic instruments. Semantic ratings, by contrast, were remarkably consistent across expertise groups.
timbre.fun: A gamified interactive system for crowdsourcing a timbre semantic vocabulary
Based on the prompted synthesis task, Hayes, Saitis, and Fazekas (2022b) developed the timbre.fun game. Debuted at the 2021 Edinburgh Science Festival, it attracted nearly 800 users from 35 countries, yielding hundreds of tagged sounds. Even with this more casual, diverse sample, the emergent structure of the data aligned meaningfully with our prior controlled findings.
Very interestingly, and somewhat unexpectedly, the emotional arousal connotation of timbral metaphors (using validated word affect norms) proved to be a reliable predictor of the acoustic character of the sounds that people created (binomial prediction accuracy: synth parameters 73.1%, acoustic PCs 71%; p < 0.001).
Timbre semantic associations vary both between and within instruments
Reymore, Noble, Saitis, Traube, and Walmark (2023) looked at the variations in timbre within an instrument and between different instruments, and the corresponding semantic associations. To address these variations, we designed an experiment to examine the effect of register on instrumental timbre semantics, with additional analysis relating specifically to pitch height. Four of the instruments (violin, bass clarinet, trombone, and vibraphone) were selected based on the ACTOR CORE (Composer-Performer Orchestration Research Ensemble). Vibraphone sounds were bowed, rather than struck, to maintain consistency of excitation type. We then added flute, oboe, trumpet, and cello in order to balance the range of the stimuli and to maximize the variability of orchestral timbres tested. We note that the vibraphone was included precisely because it is an outlier—the only percussion instrument, idiophone, non-default technique, and non-standard orchestral member, offering useful variability for studying how instrument type interacts with register-semantic relationships. We used 20 semantic scales (sets of descriptive adjectives) derived from Reymore & Huron (2020):
| deep, thick, heavy | brassy, metallic | woody | pure, clear, clean | hollow |
| smooth, singing, sweet | raspy, grainy, gravelly | muted, veiled | resonant, vibrant | watery, fluid |
| project, commanding, powerful | ringing, long decay | sustained, even | percussive (sharp beginning) | focused, compact |
| nasal, buzzy, pinched | sparkling, brilliant, bright | open | shrill, harsh, noisy | airy, breathy |
Register and instrument influence most scales. For certain sets of terms, the results showed similar patterns across registers for all the instruments, for instance, deep/thick/heavy was consistently rated highest in the low register, whereas sparkling/brilliant/bright received the highest ratings in the high register. For other terms, relationships between register and semantic associations depended on the instrument, for example, the trombone was rated most smooth/singing/sweet in its higher register, whereas the trumpet received increased ratings for smooth/singing/sweet in its middle register. There was little variance among registers for the descriptors brassy/metallic and sustained/even across all the instruments. Unlike the other instruments, the vibraphone displayed very little semantic variation across registers.
Pitch height explains more variance than register. Data were analysed in several ways, including exploratory hierarchical clustering. These clusters were not easily explained by instrument, instrument family, or Hornbostel-Sachs categories for organology (apart from the vibraphone, which stands alone in that its low, medium, and high notes are clustered together). Rather, they are better interpreted using pitch height, with a high cluster (G5–C7), a medium-high cluster (C5–G5), a medium-low cluster (G3–G4), and a low cluster (C2). Two post hoc models were considered for each semantic scale, one using only pitch height as a fixed effect and one using only register as a fixed effect. We found that the former produced significantly better goodness-of-fit than the latter, i.e., explaining more variance in semantic ratings.