SYSTEM NOTICE

Auto translation by AI. Be sure, accuracy, nuances and authorial intent may not be fully reflected.
見出し画像

[Suno Mixing: Doubling] How to Create Stereo Vocals with Depth

Tokii here.

Suno vocals are
"perfect in terms of singing ability and expression, but compared to professional tracks..."

・Lacking impact
・Wanting thicker harmonies
・Wanting stereo vocals with depth

Do you ever feel that way?

This is often not due to prompts or a lack of sense, but rather
a current characteristic of Suno.

The "number of layers" in generated vocals is low (often just 2 layers: main + backing), and the "localization" (position of the sound) is too uniform (vocals tend to be centered). This is often the cause.

In this article, I will explain how to use the classic music production technique "doubling" to make Suno vocals dramatically thicker, punchier, and more stereophonically deep.

It is super easy and highly effective, so please give it a try ♫

📝Author
Tokii (aka Tokima Tokio) | Graduate of Osaka University of the Arts, Department of Fine Arts
・Musician (2009–2021)
・Video Creator/Director (2022–)
Notable work: Avicii - Without You (Tokima Tokio Remix) (2017)


First, here are the results ♫

The original vocal (Track 1) is heard from the center, but
from 8 seconds onwards, the additional vocals (Tracks 2 and 3) are heard from both sides.

Earphones or headphones recommended 👂️

What do you think?

The method is very simple; just by placing additional vocals on the left and right, you can achieve this stereo effect.

[Track 1] Main vocal localization is center (heard from the middle)
[Track 2] Pan the additional vocal to (L)
[Track 3] Pan the additional vocal to (R)

*If the additional vocals are the same material, shift the timing of Tracks 2 and 3 slightly (about 10–30ms) forward or backward.


Practice: 2 ways to make Suno vocals stereo

Eh, Midjourney

First, "doubling" is a technique to add thickness to the sound by layering another take of the same melody (=making it double). is.

And, once you have "doubled" it, placing those sounds left and right (=panning) creates a stereo feel.

Doubling: Layering another take of the same melody *Irrelevant to mono/stereo
Panning: Deciding where to place the sound left or right

Below, I will introduce 2 ways to combine these two to make Suno vocals stereo.

Method 1: Pan separate takes left and right (spreads naturally)

Recommended case: When you want to thicken the lead melody

Vo1 (pink) -> PAN: C (center), Vo2 (green) -> PAN: 100L, Vo3 (blue) -> PAN: 100R

[Steps for Method 1]
1. Prepare two more of the same vocals using Suno's "Cover" function, etc.
-> If they are different takes of the same lyrics and melody, extracting them from one song is also fine 👍
2. Line up the different takes in Suno Studio (or DAW)
3. Pan one side to L (approx. 80-100) and the other to R (approx. 80-100)

Benefits】
・Produces a natural thickness as if actually singing
・Sound is less likely to disappear during monaural playback *Will explain later

[Demerits]
・Need to prepare multiple source materials
・If the texture of the different takes varies, it can easily become cluttered

By the way, the opening video uses this Method 1.

Method 2: Duplicate and shift "slightly" (time-saving and easy to adjust)

Recommended case: Center lead + want to add air to the outside

Duplicate the same material, pan left and right, and delay one side slightly

[Steps for Method 2]
1. Duplicate the additional vocal to make two tracks
2. Line up the duplicated vocals in Suno Studio (or DAW)
3. Pan one side to L (approx. 80-100) and the other to R (approx. 80-100)
4. Delay one side slightly (approx. 10-30ms)
5. Lower the volume of the faster one slightly (approx. -1 to -3dB) (countermeasure for shifting)*Will explain later

[Benefits
]
・Only one source material needed
・Can control localization accurately

Demerits]
・"Phase cancellation" due to phase shift can occur(phase cancellation)
*Will explain later

♫: The result of Method 2 is like this ↓

Method 2 is easier to control because you can adjust the delay time, and since the same sound is playing on the left and right, it is less likely to get cluttered and easier to keep clean.
However, as a point of caution, "phase cancellation" due to phase shift can be mentioned.

However, as a point of caution, "phase cancellation" due to phase shift can be mentioned. (phase cancellation) is mentioned.

Important: Checkpoints to avoid failure

The following two points are points you definitely want to check.

✅️1: Prevent "shifting" due to the Haas effect

Since the ears have a property where the "sound heard first" is heard more strongly,
you can achieve balance by making the volume of the faster sound slightly smaller (-1 to -3dB).

The Haas effect (precedence effect) is an acoustic psychological phenomenon where, when listening to almost the same sound from multiple directions, the sound is heard from the direction that arrived earlier in time. Normally, if the delay is within 30-40 milliseconds (ms), the brain combines the sounds and recognizes them as one sound, and the sound image moves/localizes to the side that arrived first (the faster one).

✅️2: Beware of "phase cancellation" due to phase shift *Especially for Method 2

When the left and right sounds overlap, the sound in a specific frequency band may disappear and become "suddenly thin".

Phase cancellation is a phenomenon where the 'peaks' and 'valleys' of sound waves overlap, causing them to cancel each other out, which can make specific frequencies disappear or make the overall sound thin.

When sounds in phase overlap, the volume increases; when sounds out of phase overlap, they cancel each other out.

The 'cancellation' caused by phase misalignment becomes noticeable during monaural playback, but you might not notice it in a stereo environment while producing.

Therefore, always listen in a monaural environment to check if there are any parts where the sound becomes extremely thin.

☝️ How to check in mono (any one of these is fine)
・Temporarily set your device to mono to check
・Check by setting it to mono in your DAW, etc.
・Listen through smartphone speakers (not strictly mono, but it makes it easier to notice thinness)

If the sound becomes thin, shortening the time shift usually makes it more stable.

If the low end sounds muddy, or if you want to lightly layer it as a backing vocal, you can apply a high-pass filter to the layered track and cut the low end (aim for 150–300Hz) to achieve a cleaner finish.

Cut the low end with a high-pass filter (HPF)
I cut it "a lot" this time because I wanted a cleaner sound


Application example: Layering ad-lib vocals to create a big chorus

You can apply this technique to place main vocals on the left and right, and add ad-libs singing in the center. (Currently, this is difficult to do in a single generation...)


Conclusion:

Mixing is a deep subject, and there is no end to it once you start getting particular.

However, knowing these basics will change how you "hear" things.
As a result, I think the range of your expression and the range of your instructions to the AI will also expand.

With AI becoming capable of so much, the things humans actually need to do might decrease. Even so, I believe that "adding a touch" with intention becomes the "difference" in a work and the "meaning" behind creating it.

Please try this technique in your own songs.
I hope this helps!

Tokii

いいなと思ったら応援しよう!