A Talk on Difference Tones That Musicians Don't Know
Article from May 25, 2026 (Updated 2026.5.28)
Even in the modern era, musicians are about the only ones who cannot distinguish between a "difference tone (a type of combination tone)" and the missing fundamental phenomenon.
The most a musician barely knows about "difference tones" is that they were "discovered" in the 18th century by Tartini and are used for tuning. If they are a bit more knowledgeable, they might know the term "Tartini pitch/note" (Note 1).
First, it is important to know that the discoverers (Sorge and Tartini) were mistaken in thinking that "difference tones" were "physical interference caused by two sounds in the air." You can consider that the true source of "difference tones" is almost non-existent in the air (though it does exist to some extent).
I apologize for saying something that immediately contradicts intuition, but the true nature of the phenomenon called a "difference tone" is that we are listening to sounds produced by our own "ears" (*).
* Since Helmholtz's experiments using a harmonium, the existence of "instrument-derived objective difference tones" has been claimed, but at present, there is no evidence to sufficiently support it.
1. The prophecy of a young physics enthusiast buried for 30 years
In 1948, a young man published a paper. The author was Thomas Gold, a young man still in his 20s. He was an engineer who majored in mechanical sciences and was affiliated with the Department of Zoology at the University of Cambridge. He was completely outside the field of physiology.
The title of the paper was "Hearing. II. The physical basis of the action of the cochlea." Its claims were unacceptable to the physiological community at the time.
Let me summarize. Hearing was already known to have high frequency resolution, but Gold noticed that the energy balance did not add up when trying to achieve that resolution with the "passive" cochlea model (which is only shaken by input sound), which was the most influential theory in the physiological community at the time (*). The viscous damping that occurs in the fluid-filled cochlea was too great. It was too large to realize the observed quality factor Q (Q value).
* Actually, Gold was a radar and radio communications expert who had served as the design and manufacturing manager for British Navy radar development during World War II, and it seems he felt something was off based on that experience and knowledge. (Regardless, it's amazing)

Because of this viscous damping, the Q value of a passive cochlea can only reach about 20 at most. Yet, when actually measuring the human ear, the Q value is as high as 60 to 250. Gold reached the conclusion that something must be supplying energy to resist this damping!
So, Gold proposed the "regeneration hypothesis." He argued that there must be an electromechanical action working inside the cochlea that supplies electrical energy from the outside to cancel out the damping. To explain this so that even modern musicians can understand it, the cochlea is not a "passive" microphone, but rather like a vacuum tube radio (regenerative) that was popular before the war.

Gold himself cites the regenerative circuit of a radio as an example. A regenerative circuit is a detection circuit that increases sensitivity and selectivity by adding positive feedback. He argued that similar feedback must be working in the cochlea as well.
Gold went even further. He stated that when this feedback exceeds the loss, specific elements (*) in the cochlea enter self-oscillation. At least some "tinnitus" is likely this self-oscillation. And he added, "If tinnitus is due to actual mechanical vibration, a portion of the acoustic energy should be radiated to the outside. A high-sensitivity measuring instrument would be able to detect this and prove its mechanical origin. This would be almost definitive proof of this theory."
Actually, Gold tried to verify this himself at the time but failed. The microphone's sensitivity was insufficient.
* At the time, Gold thought that each part of the basilar membrane vibrated, and that its self-oscillation caused ringing. Ringing is a phenomenon where, for example, if you trace a wine glass with a wet finger, the glass continues to vibrate in a self-sustaining manner. The sustain of a guitar or piano is based on the same principle. Note that this prediction of Gold's was unfortunately incorrect; it was later discovered that the outer hair cells in the cochlea were actually acting as active amplifiers. The identification of this fact did not happen until the 1980s or later.
As was the case in the 1940s when Gold's paper was written, the physiology of hearing in the mid-20th century was dominated by Georg von Békésy's "passive" cochlea model. Békésy was an authority at a level that even won him a Nobel Prize in 1961. What he observed was the cochlea of cadavers of mammals, including humans. However, because the tissue was dead, the outer hair cells did not vibrate spontaneously. It was natural for Gold's theory to be overlooked by the standard research methods of physiology at the time.
──30 years later.
In 1978, British physicist David Kemp, who worked at a laboratory attached to an otolaryngology clinic in London, inserted a high-sensitivity microphone into the ear canal (!) and succeeded in measuring the phenomenon of "otoacoustic emissions," where the cochlea reacts to stimulation and radiates sound back outward. Kemp accomplished in 30 years what Gold had designed in 1948 but failed at due to a lack of sensitivity.
Gold's "cochlear amplifier hypothesis" was proven to be completely correct. The ear was not a "passive" microphone, but an organ that "actively" produces sound. The "nonsense" of a "young physics enthusiast who came out of nowhere from a different field" had overturned the authority of a Nobel laureate. Isn't this giant-killing dramatic enough to be a novel or a movie?
This "drama" surrounding "difference tones" may seem to be overshadowed by Gold's later achievements, which were far too immense, but when I think about how the lack of response from the physiology community at the time disappointed Gold and became the catalyst for him to pursue a path as an astrophysicist, I find it deeply moving.
2. What are Otoacoustic Emissions (OAE)?
The phenomenon that Kemp measured is now called "otoacoustic emissions" (OAE). It is a phenomenon where the outer hair cells (OHC) in the inner ear's cochlea vibrate "actively" and "radiate back" sound waves toward the ear canal.

There are two main types of OAE.
■ Spontaneous Otoacoustic Emissions (SOAE)
A phenomenon where outer hair cells radiate sound without any external sound stimulation. It is detected in approximately 50-70% of people with normal hearing (with significant variation depending on various conditions), is detected over a wide range of the audible spectrum, and in adults, it tends to be concentrated around 1-2 kHz.
The reason many people do not usually perceive this SOAE is thought to be because the signal is at a borderline level of whether it can be perceived, it is masked (drowned out) by external sounds (such as noise), and it is "processed" as a steady signal (which can be explained through neurology, perceptual psychology, and cognitive science). It also seems that people whose signal levels fluctuate easily tend to perceive SOAE.
■ Distortion Product Otoacoustic Emissions (DPOAE)
A phenomenon where, when two sounds are played simultaneously, countless distortion components are generated within the cochlea and radiated. Among these, the quadratic difference tone (QDT) and the cubic difference tone (CDT) have received particular attention due to the ease of measurement.
・Quadratic difference tone (QDT)

QDT is often explained as the true identity of the "third sound" that Tartini heard. It is explained simply as the "frequency of the difference between two sounds" ( f₂ - f₁ ). When two sounds vibrate the basilar membrane simultaneously, the OHC responds non-linearly, creating a motion component equivalent to the "product of the two waves" (quadratic non-linearity). Incidentally, QDT is known for being difficult to perceive, and the prevailing theory today is that what Tartini was hearing was the CDT, which will be explained next.
・Cubic difference tone (CDT)

CDT is relatively easier to perceive than QDT. Outer hair cells actively expand and contract in response to sound, amplifying the vibration of the basilar membrane, but this expansion/contraction characteristic itself is non-linear and saturates with strong input. Therefore, a component different from QDT is born from the interaction of the two sounds. In short, CDT is born from "stacked multiplication buffs." When calculated, it sounds exactly at the frequency of the "difference between the octave above the lower sound and the other sound" ( 2f₁ - f₂ ). (The octave above is twice the frequency. One octave below is 0.5 times.)
Both are phenomena that originate from the nonlinearity of the OHC.

To explain this again for musicians: when a just major third (E4 and G4) is played, the "ear" produces C2 and C4 (C2 is a QDT, so it is difficult to hear).

When a perfect fifth (C4 and G4) is played, the "ear" produces C3 (one octave below C4) as a "third tone."

It is due to this property that Tartini attempted to derive the basis of harmony from "combination tones" (QDT), and that "combination tones" (QDT/CDT) came to be applied to tuning in later generations.
3. The problem of not being able to distinguish between combination tones and MF
Combination tones and the missing fundamental (MF) are often confused historically, and the background to this is the fact that combination tones generated by the nonlinear process of the cochlea can become tones of the same frequency as the MF.
This was shown numerically in 1967 from observations of neural responses by the American auditory researcher and engineer Julius L. Goldstein. He independently reached the same conclusion as Gold and realized that the cochlea is an organ that essentially responds nonlinearly (Note 2).
"Combination tones" are not just wave interference, but arise from the active and nonlinear process of the cochlea itself. The "third tone" that Tartini heard was not created in the air, but inside his own ear.
The reason why musicians in particular tend to confuse combination tones and MF is likely this.
■ Cases where the numerical values match
For example, when two tones of the harmonic series (300Hz and 500Hz) of a fundamental (100Hz) are played, a CDT of 100Hz (equivalent to the fundamental) is produced.

QDT 500Hz - 300Hz = 200Hz
In this case, confusingly, the "tone perceived as MF" and the "tone produced as a combination tone (CDT)" happen to have the same frequency. And annoyingly, because the pitch heard is the same even if the explanation of the phenomenon is different, they are indistinguishable at first glance.

A perfect fifth (just, f₁:f₂ = 2:3) is the only interval where the frequencies of QDT and CDT structurally coincide, and the MF also approximates this. It is the point of intersection for QDT, CDT, and MF, which have independent generation mechanisms.

A specific example of an incorrect explanation is shown in the following video. Although it was posted by someone with the title of visiting professor at a music university, they likely do not know about combination tones (CDT) and thus commit a typical error (creating a situation where CDT and MF can occur simultaneously, and attributing it solely to MF).
【The 'Missing Fundamental' where you hear a C even though a C is not being played】 Posted on 2017/11/13
[Hiroshi Fujimaki] This is the channel of Hiroshi Fujimaki, who composes and arranges music for anime, movies, and commercials. Tokyo College of Music, Film and Broadcasting Course (2nd graduating class); Visiting Professor, Senzoku Gakuen College of Music, Music and Acoustic Design Course; Author of Yamaha composition books and the Fujimaki Method.
■ Similarity of the phenomena themselves
Both combination tones and MF appear as an experience where 'you hear a low sound that was not input.' It is natural that musicians cannot distinguish between them at a practical level.
Although it is often explained that one is the physics of air and the other is central perceptual processing, they are different phenomena that look similar but are experimentally distinguished. Musicians confuse them because the values often match, and when playing a harmonic series, it is possible for both combination tones and MF to occur simultaneously.
To avoid these confusions, checking via one of the following methods is effective.
Present the two tones separately and completely divided into the left and right ears using headphones
Apply a masker (noise) to the same frequency band as the combination tone
What disappears with the above method is a 'combination tone'; if it is MF, it will still be heard. If you want to experience real combination tones or MF, there are site articles with good demos.
NTT Communications Science Lab 'Illusion Forum'
Also, for MF, the demo in this note article is excellent.
4. 'Inner sound' called tinnitus
In the context of otoacoustic emissions (OAE), 'tinnitus' cannot be left out. And here, Gold's 1948 paper shines brilliantly once again.
As mentioned above, while discussing the 'regeneration hypothesis,' Gold pointed out that when feedback exceeds loss, the element enters a self-oscillation state, and wrote that 'at least some tinnitus is this self-oscillation, and is not always due to central nervous system damage.' Gold had already positioned tinnitus as a pathological form of OAE in 1948. Modern research supports this, and in fact, spontaneous otoacoustic emissions (SOAE) are variously called 'objective tinnitus' or 'physiological tinnitus,' and have become established as a form of tinnitus.
If SOAEs arise from the active vibration of a normal cochlea, the boundary between tinnitus and SOAEs is a gradient, and it could be said that an "audible SOAE" is tinnitus.
The ear does not just receive external sounds; it is constantly emitting sounds from within.
5. Cage's Experience in the Anechoic Chamber
Here, I would like to look back on a famous episode in music history.
Around 1951 (the exact year is unknown), composer John Cage entered an anechoic chamber to experience absolute silence. Cage himself repeatedly spoke about that experience.
It was after I got to Boston that I went into the anechoic chamber at Harvard University. Anybody who knows me knows this story. I am constantly telling it. Anyway, in that silent room, I heard two sounds, one high and one low. Afterward I asked the engineer in charge why, if the room was so silent, I had heard two sounds. He said, 'Describe them.' I did. He said, 'The high one was your nervous system in operation. The low one was your blood in circulation.'
It was after I got to Boston that I went into the anechoic chamber at Harvard University. Anybody who knows me knows this story. I am constantly telling it. Anyway, in that silent room, I heard two sounds, one high and one low. Afterward I asked the engineer in charge why, if the room was so silent, I had heard two sounds. He said, 'Describe them.' I did. He said, 'The high one was your nervous system in operation. The low one was your blood in circulation.'
Biographer David Revill also records this event in his book, 'The Roaring Silence'.
What is important is that when Cage sat down in the stillness of the chamber, he was surprised. Far from its being devoid of sound, he could hear two sounds, persistent and quite loud - a constant singing high tone and a throbbing low pulse. Cage supposed there was something wrong with the room. Puzzled, he quit the chamber and asked the engineer in charge why, if the room was soundproof, sounds were creeping in. "Describe them," the engineer demanded. Cage complied. The engineer told him they were not any fault of the chamber; they were the sounds made constantly by his own body - the high sound the ringing of his nervous system, the low noise his blood in circulation.
What is important is that when Cage sat down in the stillness of the chamber, he was surprised. Far from its being devoid of sound, he could hear two sounds, persistent and quite loud - a constant singing high tone and a throbbing low pulse. Cage supposed there was something wrong with the room. Puzzled, he quit the chamber and asked the engineer in charge why, if the room was soundproof, sounds were creeping in. "Describe them," the engineer demanded. Cage complied. The engineer told him they were not any fault of the chamber; they were the sounds made constantly by his own body - the high sound the ringing of his nervous system, the low noise his blood in circulation.
This experience is famous for convincing Cage that "silence does not exist" and for serving as the catalyst for his 1952 work, '4'33"'.
Regarding this, Revill records in the same book that several doctors refuted the engineer's explanation (the sound of the nervous system operating and the sound of blood circulation), suggesting that what Cage heard may have been "tinnitus".
Currently, one of the leading candidates for the identity of the higher sound Cage heard is SOAE. It is thought that approximately 50-70% of adults with normal hearing have SOAEs, and in an environment like an anechoic chamber where external sounds are blocked, these may become easier to bring into consciousness.
Here, I would like to focus on the chronology. Cage visited the anechoic chamber in 1951. Gold predicted otoacoustic emissions in 1948. It is possible that Cage subjectively "heard" in the silence of the anechoic chamber the sounds that Gold could not pick up due to a lack of sensitivity.
6. Post-hoc Overinterpretation of '4'33"'
The explanation that it is "a work that encourages noticing environmental sounds within silence" is the most superficial reception of '4'33"'. Cage himself spoke of it in that way, but that seems like an incomplete explanation.
The most crucial aspect of '4'33"' lies in the structure where the performer continues to choose to "not produce sound" throughout the three movements. There is no coincidence there, but rather an intentional repetition. In the repeated refusal, what is asked of the audience is not "is environmental sound music?" but first, "what is musical sound?" and "what is music?" (I think Cage fails to explicitly ask this).
The audience at the premiere of '4'33"' was "braced" for and continued to expect to "listen to music" in the institutional setting of a concert hall. The work is designed so that this "bracing" and expectation itself become the object of inquiry. As a result, for the audience to hear external, accidental sounds as "music" is a bad bet, and it does not seem that he was aiming only for that. Cage must have also been betting on the audience hearing something else, based on his experience in the anechoic chamber. In any case, if it was enough for the social construct of "musical sound" to be visualized, it was a winning bet for Cage.
By the way, Cage did not anticipate otoacoustic emissions (OAE). However, if that perspective were introduced, the reading of "4'33"" mentioned above could have been reinforced by physiology. In the silence where socially constructed "musical sounds" are "withheld," the audience (about 50% to 70% of them) encounters sounds that exist before culture and before institutions (SOAE). The cochlea, which continues to oscillate even without input, fills the ear canal with subtle sounds it produces itself. The deeper the external silence, the closer the audience gets to those sounds. This provides a possibility to approach the question "What is a musical sound?" from both social and physiological perspectives. The fact that "the auditory organ itself emits sound" was quietly present at the bottom of the hole that Cage's question left dug, as an answer that no one at the time knew.
Actually, there was a successor to Cage who utilized OAE in music!Maryanne Amacher (1938–2009). To confirm the timeline:
1948, when Gold wrote in a paper that "the ear must emit sound externally"
1951, when Cage heard sounds resembling SOAE in an anechoic chamber
Amacher called OAE "ear tones" and had already begun attempts in the 1970s to treat the cochlea itself as a musical instrument
1978, when Kemp demonstrated OAE
Amacher learned of Kemp and Gold's research in 1992 and discovered that the phenomenon she had been treating intuitively for nearly 20 years had a scientific name. Cage sublimated his accidental experience in an anechoic chamber into a philosophical question. Amacher intentionally incorporated the same phenomenon into her compositions. It is interesting that the trajectories of Cage and Amacher, two people, intersect around the same physiological fact, involving chance and intention, intuition and verbalization.
Of course, this is a retrospective interpretation, and I do not mean to claim that "Cage knew about OAE." However, that is the important point; Cage's philosophy that "there is no such thing as silence" later acquired a physiological basis beyond what he had intended. I think this is a good example of being able to positively grasp the idea that "a work reaches places the author did not intend."
7. Deduction, Deduction, Deduction
In Japan, the debate over whether "music theory" is necessary or unnecessary often reignites, and every time it does, I recall the words of an acquaintance: "People who say theory is unnecessary are people who do not perform deduction." Yes, that is indeed an attitude that closes off the path of deduction.
Tartini, Gold, and Cage—these three individuals differ in era and field, but they share the fact that they began their deductions from the conviction that "there must be something there." Tartini tried to deduce harmony theory from combination tones, Gold deduced the activity of the cochlea from energy balance, and Cage deduced the impossibility of silence from the sounds he accidentally experienced in an anechoic chamber.
As a result, Tartini failed in his demonstration and arrived at philosophy, while the seeds sown by Gold sprouted 30 years later thanks to Kemp, though Gold himself went on to pursue astrophysics. Cage never knew the true nature of his anechoic chamber experience during his lifetime, but his works certainly expanded the scope of "music."
The true nature of the "combination tone" we have looked at in this article—the fact that its source was inside the ear—is actually something that cannot be ignored when considering a major theme in the history of Western music theory: the perception of consonance and dissonance. This is because even if two sounds observed in the air are not pure intervals (for example, in equal temperament), the possibility arises that our ears are creating pure sounds and "hearing" them. This leads directly to the question of what consonance is and what the feeling of consonance is, and it may also become a seed for new deductions.
Seeds sprout in places unknown to those who sowed them
I suppose that is about it.
Footnotes
Note 1:Regarding the discovery of "combination tones," the German organist Sorge's 1745 work was nine years earlier than Tartini's 1754 publication. However, Tartini claimed to have heard combination tones as early as 1714. Note that Helmholtz (1863) pointed out that Tartini consistently estimated "combination tones" to be one octave too high.
Tartini tried to construct a harmony theory based on these combination tones, but it stalled due to many problems, such as the inaccuracy of his experiments, a fundamental failure in minor triads, and arguments that did not reach the scientific standards of the time. In his later years, he turned to criticizing the idea that music could not be explained by Enlightenment science, and one could say he was a person who made a more decent "retreat" than Rameau, who remained a naturalist and failed.

"If you sustain two accurately harmonized sounds, a third sound (a combination tone) will be clearly audible. It will be at the pitch indicated by the filled-in note below."
In the introduction (Figure 1.4) of his book 'Human and Machine Hearing' (2017), Richard F. Lyon (Dick Lyon) cites Tartini's 1754 plate as 'one of the earliest instances of recognizing nonlinear effects in hearing.'
Note 2:Richard F. Lyon's cochlear model, 'CARFAC: Cascade of Asymmetric Resonators with Fast-Acting Compression,' numerically demonstrates that slight asymmetric distortion in the transfer characteristics of outer hair cells can generate QDT, and that these components can match the frequency of the missing fundamental (MF). Difference tones and MF may be physically related phenomena.
Key References
Richard F. Lyon (Dick Lyon), 'Human and Machine Hearing: Extracting Meaning from Sound' (2017)
D. T. Kemp, 'Stimulated Acoustic Emissions from Within the Human Auditory System' (1978)
David Revill, 'The Roaring Silence: John Cage: A Life' (Arcade Publishing, 1992). Available for purchase on Amazon. E-book available.
The following note article is quite detailed regarding Maryanne Amacher. Even so, there is significant variation in how her name is written in the Japanese-speaking world.
