SYSTEM NOTICE

Auto translation by AI. Be sure, accuracy, nuances and authorial intent may not be fully reflected.
見出し画像

Measuring the Same Song Twice, Only Six Songs Changed Their Answer — How to Find Code That Fails Silently


Last time, I measured the 22 songs on "Nebula Romance" and wrote about where the seams were cut (the previous article is here).

Following that, I was planning to measure the 52 songs on the best-of album "Perfume The Best 'P Cubed'" (hereafter referred to as P3). It contains 15 years of music, from 2005 to 2019, all in one place. I thought that if I arranged them by year, I would be able to see how Perfume's sound has changed.

Then I remembered. This best-of album was entirely remastered by Yasutaka Nakata.

"Remastering the audio source" was in the news at the time, and I was aware of it. However, I hadn't considered what that meant for my analysis. What I was measuring wasn't the sound from 2005, but the sound after Nakata had worked on it in 2019.

So, I thought I would compare it to the original sound. I have all the CDs. That's a given for a fan.

Measuring the same song twice

I imported 50 songs from the original discs and compared them one by one with the remastered versions.

Since they are the same songs, the difference should theoretically only be due to the mastering. The performance and vocals are the same. The sequence of sounds is the same. The only difference is the finishing process.

The results were clean.

The only thing that changed was the volume of the sound, which increased by an average of 0.52 decibels. The timbre, the dynamics within the songs, and the harmony did not move within the measurable range. The only statistically significant difference was in volume.

What was interesting was that the way it was increased was not uniform.

Songs from 2005 to 2009 were increased by an average of about 1 decibel, while songs from 2015 to 2019 were only increased by 0.24 decibels. The older the song, the more it was boosted. Conversely, there are 12 songs where the volume actually decreased.

A best-of album is something you listen to through all 52 songs. It's hard to listen to if the volume varies from song to song. I think that's why the older songs were boosted to match. However, they weren't perfectly leveled, and the differences between eras remain. I can only speculate on the reasoning behind that balance.

There was one other thing that went well. "Nananananairo" is a new song from 2019, so it was not subject to the P3 remastering. The official announcement also stated, "All songs except for the new songs."

For this song alone, the difference was exactly zero. The volume, brightness, and dynamics all produced the exact same numbers. I was able to confirm that if the audio source is the same, the difference will be zero. This served as proof that the differences found in the other 50 songs were real.

Up to this point, everything had been going smoothly.

Only six songs had a different key

In my analysis, I estimate the key for each song. Things like C major or A minor. This is the foundation of the analysis, where I measure the brightness of a song based on whether it is major or minor.

Since they are the same songs, naturally, the same key should have appeared before and after remastering.

However, for six songs, I got different answers.

For example, "Polyrhythm" was identified as C minor on the original disc and F minor on the remastered version. "One Room Disco" was D major and B minor. As for "Computer City," it was G# minor and E major—completely different keys.

A song's key does not change during mastering. If that were to happen, it would be a major incident.

So, did I compare different audio sources? I checked that too. When I measured the length of the songs, the difference was one to three seconds. This is due to the difference in how the silence at the beginning and end was trimmed; if they were different versions, they would differ by dozens of seconds. It was the same audio source.

I also compared the content of the sound itself. When I lined up the numbers representing how much each of the twelve notes was sounding and calculated the correlation, it was 0.99. They were almost identical.

The same audio source, and the content is almost the same. Yet, only the judgment differs.

In other words, it was my tool that was wrong.

Decided by a 0.2% difference

To find the cause, I measured how close the judgments were.

I looked at the difference between the strongest and second-strongest of the twelve notes. If this difference is large, you can say, "This is clearly this key," but if the difference is small, a slight difference in conditions can cause the first and second places to flip.

Looking at the six songs that disagreed, the closest one, "Computer City," had a difference of 0.05% . "Perfect Star Perfect Style" was 0.2%, and "Polyrhythm" was also 0.2%.

The key of the song was being decided by a 0.2% difference. It was only natural that the ranking would swap just by turning up the volume a little bit.

Three of the six songs were in a relative key relationship. Like C major and A minor, they use all the same notes, and only which one is viewed as the center differs. These are the two keys that can be played using only the white keys of a piano. This means the algorithm was struggling exactly where it is most prone to confusion.

Changing to a smarter algorithm didn't fix it

The first thing I thought of was changing the estimation method.

What I was using was a standard method created from psychological experiments. Since there are newer, improved versions, I replaced it with those and measured again.

  • Original method: 44 out of 50 songs matched (88%)

  • Another standard method: 44 out of 50 songs matched (88%)

  • Yet another improved version: 43 out of 50 songs matched (86%)

It didn't change. In the case of the improved version, it actually got worse by one song. And the six songs that disagreed were common to all methods.

It was here that it finally made sense. This is not a question of which algorithm is superior. Songs where the first and second place are close will remain close no matter what method you use to measure them. The difficulty in judgment was not the fault of the tool, but the nature of the songs themselves.

Admitting that you don't know what you don't know

So, I changed my approach. I was making mistakes because I was trying to judge every single song, so I should only judge the songs that can be judged.

I narrowed the scope to songs where the difference between the first and second place was above a certain threshold, and re-measured the consistency rate.


When I set a cutoff at two percent, they match perfectly.

The target count drops from fifty songs to thirty-five. That is seventy percent. But for the remaining thirty-five songs, I am guaranteed that even if I re-measure them with different mastering, I will get the same answer.

Instead of trying to make a clever judgment, I let go of what could not be judged. That was the answer.

Note that this cutoff only applies to key determination. Since key estimation is not involved in measuring volume or timbre, the result from the earlier remastering that "only the volume changed" was based on all fifty songs.

What happened to the past articles?

I was quite panicked at this point. Every article I have written for this analysis is based on this key estimation.

I re-examined the twenty songs of "Nebula Romance" using the new criteria.

The number of songs for which judgment should be withheld was two. Both are from the first part.

The impact on the conclusion is as follows.

First Part ・Counting all 10 songs: Major 7 : Minor 3 ・Counting only the 8 confirmed songs: Major 5 : Minor 3

Second Part ・Counting all 10 songs: Major 2 : Minor 8 ・Counting only the 10 confirmed songs: Major 2 : Minor 8 (unchanged)

The second part was unscathed. All ten songs meet the criteria. In the first part, two songs are withheld, but the trend of there being more major keys remains.

The claim I had been making that "the first part and the second part are mirror images" has weakened. But it has not collapsed. Honestly, I felt relieved.

More than the fact that the conclusion did not change after recounting, I think the fact that I was able to recount is more significant. What I could previously only say was "7 to 3," I can now say is "7 to 3, although 2 of those are close calls, so if we withhold them, it is 5 to 3."

My ears had noticed first

There is actually a prologue to this series of stories.

I was the one who first remembered the remastering issue, and this verification began when I asked the generative AI, 'Best-of albums have all tracks remastered, right?' The AI, which had been looking at the data, did not know that.

The same thing has happened many times. It was my memory as a fan, not the numbers, that realized 'Electro World is a re-recorded version included on the best-of album' and pointed out that 'this song is different between the album version and the single version.'

Numbers do not tell you what those numbers represent. The ones who know that are the ones who have been listening.


From here on, I will write about 'how to doubt code written by AI,' which I saw through this verification. In fact, in this analysis, there were parts that returned incorrect answers without issuing any errors or warnings. And it wasn't just once.

As a story about music, it concludes here. Please continue only if you are interested in the technical aspects.

ここから先は

3,416字

¥ 280

この記事が参加している募集

この記事が気に入ったらチップで応援してみませんか?