SYSTEM NOTICE

Auto translation by AI. Be sure, accuracy, nuances and authorial intent may not be fully reflected.
見出し画像

[Sound Science #60] Why do we feel that live performances and recorded music sound "different"? The secret of "sound onset" that the brain distinguishes

Hello, this is NUARL.

Have you ever had the experience of hearing a live performance at a cafe on a holiday afternoon, or being stopped in your tracks by the tone of a street musician's acoustic guitar? If you love music, haven't you been captivated by such a moment at least once?

The sound of an instrument playing right in front of you has an energy that vibrates the air and seems to envelop your body directly. However, music heard live versus music heard through earphones—even though it is the same song, haven't you ever felt that the "vividness" or "sense of air" is somehow different?

No matter how much recording technology or playback environments evolve, we unconsciously realize that it is"sound coming from audio equipment (such as earphones or speakers)". Actually, an element called "sound onset (transient)" is deeply involved in this.


The elements that humans perceive in sound are extremely complex.

It is said that sound has three elements: "pitch," "loudness," and "timbre," but our brains are even more sensitive to"the change at the moment the sound begins".

In the world of psychoacoustics, it is known that if you isolate and play only the part where the instrument's sound is sustained (the steady-state portion), even professional musicians find it difficult to distinguish between a piano and other instruments.

The reason we can instantly judge "this is a piano" or "this is a guitar" is thanks to the complex information packed into the first few tens to hundreds of milliseconds when the sound starts.

Research on timbre perception also scientifically shows thatthe sound onset time (attack time) is an extremely important dimension for humans to identify timbre.

The moment a piano key is struck, the moment a guitar string is plucked. By capturing the explosive change in energy at the moment the sound is born, the brain unconsciously imagines the material and even the resonance of the performance space. This initial sound onset of just a few milliseconds is a crucial clue that makes the brain judge whether the sound is"the real thing".


The human ear perceives even the slightest differences in sound.

Furthermore, the human ear is surprisingly sensitive to time delays. Research on the temporal resolution of human hearing has been published for a long time, including in academic journals such as those of The Acoustical Society of America (ASA). As Viemeister's (1979) research shows,the human auditory system has the ability to detect minute fluctuations in volume on the order of milliseconds.

In addition to this, it is believed that our brains detect the extremely small difference of a few ten-thousandths of a second in the sound reaching both ears (interaural time difference) to identify where the sound is coming from and the spaciousness of that environment.

A classic study by Klumpp & Eady (1956) supporting this showed that the limit of the interaural time difference (ITD) that humans can perceive is about 10 microseconds (one hundred-thousandth of a second). The human brain calculates the slight time difference in sound reaching the left and right ears and recognizes it as a three-dimensional space, such as "a guitar played diagonally to the right".


Things to be solved on the playback device side

Now, let's apply this to the perspective of audio equipment.

When trying to reproduce"vivid sound"with earphones, the key ishow fast and accurately the driver (the part that converts electrical signals into sound) can follow the input signal.

Music signals are sent to the earphones as electrical waves, but the driver must convert them into physical "vibrations" to push the air.

The barrier here is the weight (mass) of the driver itself and the law of inertia. If the physical movement is slow and the onset is dull, the brain unconsciously sees through it as "this is a recorded sound," and it sounds as if a veil has been placed over it.

No matter how high-resolution the sound source is, if the movement of the driver that ultimately produces the sound is sluggish, the original sharp onset (transient) is rounded off, and the reality is lost.

Pursuing this"speed of onset"to the limit within wireless earphones, where the output of the built-in amplifier is small, is a very difficult and simultaneously challenging theme for audio manufacturers.


Aiming to reproduce live performance

The "MEMS speaker" featured in the NUARL νClip wireless earphones has an extremely high response to this very "sound onset."

MEMS speakers are a completely new type of transducer (a device that converts electrical signals into sound) manufactured from silicon wafers. Compared to conventional speakers, they have no coils or magnets, making them extremely light in physical weight, and they start up with overwhelming speed in response to input electrical signals.

Their potential lies in their ability to vividly depict the "energy of the moment sound is born," which has been difficult to reproduce until now, and to create a realistic sense of air that makes the brain mistake it for "live sound."

Especially in open-type earphones, where the distance from the driver producing the sound to the eardrum is long, the response speed of this driver has a significant impact on the reproduction of sound realism.


☕️ Summary

The reason we feel that live performances and recorded sound are "different" is not just a matter of sound quality or resolution.

· The key to timbre perception is the "onset": The brain judges timbre and "authenticity" based on the complex information packed into the first few milliseconds (attack) when a sound begins.

· Hearing that captures 1/100,000th of a second: Our ears calculate the extremely minute time difference between the sounds reaching both ears (interaural time difference) to perceive three-dimensional spatial breadth.

· Physical responsiveness creates realism: If the onset of the playback driver is sluggish, the original sharp transient is rounded off, and the brain unconsciously identifies it as "recorded sound."

When we are captivated by the atmosphere of a live performance, our brain is making full use of this extremely short "time axis" information, not just simple sound volume or pitch, to perceive the reality of the music in front of us.

Even for songs you listen to casually, just by paying attention to the sound onset, you might see a completely different landscape.

Next time you listen to music, please try listening to the "moment a guitar string is plucked" or the "moment a drum is hit with a stick." How does it sound to your ears?


Thank you very much for reading this far.

The official NUARL LINE account delivers quizzes about sound, information about products, and more. If you are interested, please take a look.

👉 Click here for the official LINE

If you'd like, please give us a like and follow. Your likes are the driving force behind writing the next article. ✨


📚 [References]

Viemeister, N. F. (1979). Temporal modulation transfer functions based upon modulation thresholds. The Journal of the Acoustical Society of America, 66(5), pp. 1364-1380. https://doi.org/10.1121/1.383531 McAdams, S., Winsberg, S., Donnadieu, S., De Soete, G., & Krimphoff, J. (1995). Perceptual scaling of synthesized musical timbres: Common dimensions, specificities, and latent subject classes. Psychological Research, 58(3), pp. 177-192. https://doi.org/10.1007/BF00419633 Klumpp, R. G., & Eady, H. R. (1956). Some measurements of interaural time difference thresholds. The Journal of the Acoustical Society of America, 28(5), pp. 859-860. https://doi.org/10.1121/1.1908493