Japanese Lyric-Melody Fitting and Vocal Production
What listeners hear in a song is the multiplicative effect of lyrics and melody: a melody is not built from the organization of notes alone — it comes into being inseparably fused with the lyrics. This skill covers the full chain from "interlocking the two temporal structures of words and music" to "delivering a finished product from the recording studio."
1. Fitting the Temporal Structures of Words and Music (Pitch Accent / Segmentation / Syllable-Count Correction)
1. Core principle: interlocking two kinds of time
- Reading lyrics aloud naturally produces rhythm and inflection — lyrics possess their own "lyric time"; a melody likewise has its own mechanics and inherent rhythm.
- If you stuff words into a melody by character count alone, the two rhythms will inevitably clash: words get cut off in the wrong places and the meaning becomes hard to convey.
- Merely matching the rhythm is not enough — the melody will lack appeal. A good melody = one that respects the rhythm and inflection of the lyrics while making the two temporal structures reinforce each other, complement each other, and sometimes deliberately oppose each other.
- Key points from a comprehensive correction demonstration:
- Deliver words clearly with short notes so they are easy to understand;
- At moments of intense emotion, use rhythms with a hooked, syncopated feel and leaping intervals;
- Let the closing phrase start on a strong beat and end with force.
2. Intonation (イントネーション, pitch accent)
- Every word has an inherent contour of pitch and stress. Principle: the melody's rises, falls, and dynamics should match the pitch accent of the lyrics.
- When they disagree, listeners struggle to grasp the meaning by ear alone — the song becomes one whose "lyrics don't come through," and in extreme cases is heard as something else entirely:
- Example: 「雨が晴れて虹の橋が架かり」 gets heard as 「飴が腫れて2時の端が係」 (the meaning is completely scrambled).
- Pitch accent should follow hyōjungo (標準語, Standard Japanese); note that Standard Japanese ≠ Tokyo dialect — when unsure, consult an akusento jiten (アクセント辞典, pitch-accent dictionary).
- Deliberately departing from the pitch accent is an expressive device: it creates a "verbal hook" that grabs attention and expresses big emotions; taking the particle up to a high note can build emotion in surging steps.
- Precondition: the adjacent words carry correct pitch accent and the overall meaning still comes through; syllables prone to mishearing should yield priority to the melodic climax.
- Phrase-ending treatment gives the sentence its meaning:
- Gobi no jōkō (語尾の上行, rising phrase ending) = question, appeal, irresolution, hope;
- Gobi no kakō (語尾の下行, falling phrase ending) = answer, conclusion, assertion, resignation.
3. Rules for handling special syllables (tango no rizumu 単語のリズム, the inherent rhythm of words)
Ignoring a word's inherent rhythm makes it hard to sing and robotic-sounding, and in extreme cases unintelligible (「飛行機に乗って」 → heard as 「ひこきに、のて」).
- Chōon (長音, long vowels) (情熱→ジョーネツ, 希望→キボー):
- Standard treatment: give the long vowel a long note value; alternatively, give it its own note and re-articulate (iinaosu 言い直す) for emphasis.
- Warning: arbitrarily stretching a non-long syllable changes the meaning — 「彼が好き、好きだから」 stretched in the wrong place is heard as 「カレーが好き、スキーだから」; make the bunsetsu (文節, phrase unit) align with the stretched note.
- Sokuon (促音, the small っ glottal stop) (走って→ハシッテ):
- Standard treatment: represent the sokuon with a rest;
- Variant: first stretch the vowel before the sokuon, then cut it off abruptly with an accent (sung in practice as 「はしーぃって」「わらーぁって」).
- Hatsuon (撥音, the moraic ン) (弾んだ, 気分):
- It can merge with the preceding sound into one syllable, or be given its own note; but ン is hard to sing strongly and long — when giving it its own note, make it short and weak for a natural result.
- A long ン on a strong beat is easily buried by the backing track; the trick is to take it off the chord tones (コードトーン) so it floats above the accompaniment.
- Museion (無声音, devoiced vowels) (the フ in 深い, the ク in 高く — in Japanese, voiceless consonants such as t, k, f, s, p combined with the vowels u/i are often devoiced):
- Devoiced sounds carry no perceptible pitch — if a melodic note lands on a devoiced syllable, its pitch is effectively lost, and the effect of a repeated melodic figure disappears.
- Assign short notes to devoiced syllables, or the musical meaning of the phrase will change.
- Consecutive identical vowels (the と + お in 「君のこと+思ってる」):
- Left untreated, the two notes fuse into one long vowel and the second vowel gets swallowed (heard as 「君のこと持ってる」).
- Two remedies: insert a rest between the two notes to prevent fusion, or change the pitch between the two notes to prevent fusion.
4. Bunsetsu (分節, phrase segmentation) and melody
- Japanese lyrics form a bunsetsu structure of words and particles. When the segmentation of the lyrics matches the segmentation of the melody, the lyrics are easy to understand.
- Misaligned segmentation makes meaning hard to convey or even changes it: 「君が来るまで待ってる」 with wrong timing = indistinguishable from 「君が車で待ってる」.
- Deliberately mismatched segmentation is a common device (preconditions: the meaning is preserved, and pitch accent and stress position are respected):
- Breaking before a particle keeps the listener in suspense and builds momentum step by step;
- Cutting in the middle of a word creates tension and draws attention.
- Three emphatic treatments (example: 「私はここで」):
- Stretch 「わたしー」 to create tame (タメ, held-back tension) that induces anticipation for the next word;
- Segment 「わ・た・し」 syllable by syllable, like spelling it out for emphasis, pressing each character home;
- Unnaturally isolate only 「わ」, jolting the listener to attention.
- Principle: strongly segment the word you want emphasized, and it will rise to the surface.
- Wholesale mismatch of segmentation: singing the entire song in even eighth notes, letting musical phrasing take priority over linguistic segmentation, produces a sense of single-minded devotion / earnestness — "even so, I must sing these words";
- In this case, secure word boundaries through pitch movement (kashi no bunsetsu o onkō no undō de kakuho 歌詞の分節を音高の運動で確保) to keep the meaning from becoming too obscure;
- Overused, it pushes the lyric meaning too far into the background and can even turn comical.
5. Fuwari (符割り, syllable-to-note mapping)
- Principle: "one character per note" (一音一字) is the ideal.
- The recent trend of "cramming in more characters to create a sense of speed" is popular — but if a lyricist writes this way unconsciously, the recording session descends into chaos, and the one who gets hurt afterward is the lyricist.
- The same melody automatically changes its fuwari depending on the words (a real example: the same melody 「ソラファファ」 set to six different four-character words):
- 「あきかぜ」 = standard one character per note;
- 「そよかぜ」 = そ→よ approaches diphthongization; よ gets absorbed;
- 「きたかぜ」 = き is pronounced short and the stress shifts to た;
- 「ありがとう」 = the two kana とう merge to occupy a single note;
- 「さようなら」 = the う is omitted (sung as さよなら);
- 「こんにちは」 = こん merges into a single note.
- When writing or setting lyrics you must anticipate this elasticity of Japanese phonology.
- The best insurance: ask the lyricist to record themselves singing it in their own voice and send it over — a poor recording environment or a quiet voice does not matter; otherwise the lyrics will circulate in the world with the wrong fuwari.
6. Four methods of syllable-count correction (adjusting the melody to the words in kyoku-sen workflow)
- Melody note count < lyric syllable count (too many words): subdivide note values in the melody, or append notes before/after.
- Melody note count > lyric syllable count (too few words): merge note values into long notes.
- Counts match but the rhythm doesn't fit: change note values/rhythm, regroup.
- Melodic contour doesn't match the lyric's pitch accent: adjust pitch and rhythm while respecting segmentation.
- The "leave-room" correction method: don't finish the melody in fine detail up front — keep room for correction and refine it in dialogue with the lyrics:
- First draft: fix only the positions of key words; place the rest roughly;
- Second draft: split/merge notes according to the words;
- Third draft: adjust pitch accent, put the key words on the melodic apex, and align beats with words to close.
- Caution: these operations can destroy structural effects such as repetition and contrast — if a repeated structure gets wrecked by accommodating the lyrics, go back and revise the lyrics instead.
7. Ji-amari (字余り) / Ji-tarazu (字足らず) — deliberate mismatch
- Ji-amari (じあまり, too many syllables): more words than the melody can hold → specific words get delivered in rapid-fire early articulation, producing a breathless desperation and urgency (earnestness / setsujitsu 切実).
- Compressing the notes at the head of the phrase creates a distinctive "pitching forward" momentum and a forward-driving intensity;
- Comparison by compression position: head compression = strongest momentum; mid-phrase compression = momentum plus emphasis on the following words (but segmentation is easily misheard); compression at a segmentation boundary = the most well-behaved and smooth but lacking momentum; end-of-phrase compression = sounds like a hasty wrap-up, the words "spill out";
- Repeating the ji-amari rhythmic figure through the phrase reinforces the feeling of desperate forward drive;
- Warning: the impression is extremely strong and easily contrived; if the compressed, emphasized word fails to earn sympathy it slides into comedy; overdone, consecutive equal note values push melodicism into the background and the melody turns "speech-like" (which can be exploited in reverse for narrative passages).
- Ji-tarazu (じたらず, too few syllables): too few words to fill the melody → specific words get stretched and articulated across multiple pitches, drawing special attention and allowing deep feeling to be sung into them (pouring emotion into the long notes).
- Side effect: overstretching pushes the meaning into the background and the passage takes on an instrumental quality like scat (スキャット, wordless syllabic singing).
2. The Musical Elements of Lyrics (Rhyme / Rhythm / Rhetoric)
In (韻, rhyme)
- In (韻) = repeating words containing specific syllables, giving the lyrics a sense of rhythm; it includes both exact-sound repetition (umi 海 / umi 産み) and repetition of matching vowel combinations (umi 海 / kuni 国 = the u-i rhyme).
- Classification: tōin (頭韻, head rhyme / alliteration) placed at the start of a bunsetsu, and kyakuin (脚韻, end rhyme) placed at its end; employing rhyme is called in o fumu (韻を踏む).
- Setting rhyme to melody: at rhyming positions keep rhythm and beat placement exactly identical while the pitch movement differs each time — rhyme naturally makes the melodic rhythm cohere.
- Warning: stretching a rhyme too long produces a stiff, foreordained impression — let elements other than rhythm (pitch, etc.) flow and vary to preserve a sense of freedom.
- To emphasize the rhyme: give the rhyming parts an identical melodic figure for strong formal clarity, while words and music move freely elsewhere.
- Rap (ラップ) usage: strip pitch content from the non-rhyming parts (approaching spoken delivery), and make the rhyming parts leap out with jumping intervals.
Shichigo-chō (七五調) and the Japanese-rhythm alarm
- Shichigo-chō (alternating seven and five morae), haiku 5-7-5, waka 5-7-5-7-7, the 3-3-7 clapping rhythm (三三七拍子), matsuri-bayashi (祭囃子, festival music) rhythms, and other traditional syllable-count meters can be used deliberately to evoke a Japanese flavor.
- Reverse warning: these phrases are so familiar to Japanese ears that they slip in unnoticed, producing unintentionally Japanese-sounding melodies — be especially vigilant when writing rhythm-driven melodies like rap (shichigo-chō rhythm and the 「生麦生米生卵」 tongue-twister rhythm both appear unconsciously).
Rhetoric (レトリック) and melody
- Heichi (並置, juxtaposition/parallelism): lining up parallel phrases (「春がきて、夏がすぎ、秋になり、冬を越え」):
- Have the melody follow it by repeating the same rhythm or figure — the effect is doubled;
- Advanced: don't copy exactly — repeat while varying pitch and beat placement, with the pitch climbing stepwise to drive the climax.
- Tōchi (倒置, inversion): word-order inversion for emphasis, as in 「出会ったのは冬」:
- You must be conscious of which word the inversion is meant to emphasize — use a rest as tame (タメ, held-back tension) to set up the key word, or delay the key word's entrance by a full measure.
- Juxtaposition + inversion combined (「好きだ、あなたが、好きだ、いつまでも」):
- Turn the repeated phrase into a catchphrase (キャッチフレーズ, hook line) stamped in with the same long-note figure; or hammer it with an evenly divided triple hit to emphasize the instrumental quality;
- Here, having the backing play in unison (ユニゾン) is also highly effective.
- Connective logic determines the tone of the next section: look at which setsuzokushi (接続詞, conjunction) links the sections, and decide from it how the music should change:
- Junsetsu (順接, sequential connection — そして/だから): it is most natural for the melody to continue and develop as is, repeating similar melodies in an understated narration; alternatively, receive it on a grander scale, with the second phrase raised in register for a "cause → effect" climax;
- Gyakusetsu (逆接, adversative connection — だけど/しかし): change the melody's character and make a large sectional break — the first phrase lingers in a narrow range, the second moves by leaps across a wide range, presenting opposition; or connect them with same-character melodies in an understated way, expressing an objective, bird's-eye view of the whole;
- Heichi (併置, parallel placement — ところで): placing unrelated things side by side, while usually still keeping some relationship between the two phrases — same-character melodies imply the bird's-eye view; a rising first phrase answered by a falling second phrase implies causality.
- Summary: don't merely "trace the lyrics matching syllable counts" or "fill in a backdrop with general mood" — create a state in which words and music genuinely interact.
3. Shisen / Kyokusen Workflows (Lyrics-First / Music-First)
Shisen (詞先, lyrics-first)
- Write the lyrics first, then conceive the melody while singing the words — character counts, note counts, pitch accent, and pitch rhythm align naturally.
- The whole song need not be lyrics-first: writing only the catchphrase parts (such as the opening of the chorus) lyrics-first is also very effective;
- Even pre-deciding only the sokuon positions (embedding 「ずっと」「もっと」「きっと」 at phrase heads) already pays off — phrase design that exploits sokuon and chōon yields strong expression with words and music tightly fused.
Kyokusen (曲先, music-first)
- The melody comes first and the words are fitted afterward; at the stage the words emerge, manipulate and correct the melody according to the words' pitch accent (using the "four methods of syllable-count correction" above).
Fitting words by melody type
- Repeated-note melodies (rap-like, often appearing in hirauta / A-melo verse sections):
- They have a definite pitch — usually the 1st or 5th degree relative to the song's key; phrase ends often leap down a perfect 4th;
- Articulation at the downward leap is inevitably weak → place particles or unaccented characters there, and it sits cleanly on the ear;
- In a swung rhythm 「タタタタ→タッカタッカ」, the 「カ」 positions are likewise good spots for unaccented characters.
- Melodies with English scratch lyrics (localizing songs bought from overseas, often delivered with an English karauta 仮歌 guide vocal):
- Jamming one Japanese character onto every note destroys the atmosphere entirely. Countermeasures:
- On notes where the English naturally produces a sokuon-like clipped feel, set Japanese words that exploit that punchiness;
- At strong-consonant positions, choose Japanese words that likewise carry strong consonants;
- Trump card: split eighth notes into sixteenths and the like, increasing the note count to adjust;
- Where Japanese simply cannot sit on a high sustained (shirotama 白玉, whole/half-note) passage, considering English lyrics outright is often the shortcut.
- Jamming one Japanese character onto every note destroys the atmosphere entirely. Countermeasures:
- Setting words in the chorus (サビ, sabi): the formula is to choose words of strong universality, words expressing grand things so the chorus "sounds like a chorus" — but nothing interesting is born without breaking the formula.
The iron rule of English syllables
- Biggest pitfall: syllable count does not match note count. "free" is 1 syllable; it cannot be split into "f-ree" and sung over 2 notes (a melodic note cannot sit where there is no vowel nucleus); the Japanese loanword 「フリー」, however, is 3 morae and can be split.
- "freedom" = free-dom, 2 syllables, can take 2 notes; 「フリーダム」 = 5 morae. Dictionaries mark syllable division points — check them when setting words.
- Translating English lyrics into Japanese ≈ writing new lyrics: English delivers a whole word in 1-2 syllables, carrying far more information for the same number of notes; a literal translation will always run short of syllables — you need habuku chikara (省く力, the power to prune), cutting down to leave space between the lines, or simply paraphrase freely (= rethink from zero).
Fake (フェイク, improvised vocal embellishment)
- Its essence is improvisation; in principle leave it to the singer's free invention.
- Having a non-native-English singer perform Fakes laced with English does not work. If you just "want that flavor":
- Design suitable English for the singer in advance, woven in from the overall meaning of the lyrics;
- Or first let the singer freely sing whatever syllables fit, then swap in English words with similar pronunciation — the result is a Fake that is easy to sing.
Karauta (仮歌, guide vocal demo) and compe (コンペ, song competition)
- Professional workflow: soliciting songs from multiple composers for a given singer and selecting>compe (コンペ); competition entries mostly come with a karauta carrying temporary lyrics (singing la-la-la is now rare).
- The selecting side intends to judge purely on the music, but unconsciously absorbs the world-view of the lyrics — the melody's appeal is constituted together with the feel of the words. The lyrics of the karauta actually matter a great deal.
- Once the song is selected, the lyricist is commissioned (this too is often a lyrics competition).
4. Vocal Recording (Mic / Positioning / Direction / Pitch-Correction Philosophy)
Mic selection
- Photography analogy: (1) mic = lens; (2) mic preamp (マイク・プリアンプ) + audio interface = camera body; (3) DAW = film. The same mic sounds different through different preamps.
- Dynamic vs condenser:
- Dainamikku (ダイナミック, dynamic): rugged, cheap, needs no power → the live-performance standard;
- Kondensā (コンデンサー, condenser): delicate and fragile, expensive, needs phantom power → the studio standard; "the sound is simply better" (it captures even the breath, with even response across highs and lows).
- How to choose: have the singer actually sing and listen through the mic. Checkpoints:
- Does the hardness/fullness of the sound suit the song; does the midrange you want get captured; is the singer comfortable singing into it.
- Pairing by arrangement: soft arrangement → soft-voiced mic; hard sound → hard-character mic; rock punch → personality in the low-mids; R&B texture → sparkling (キラキラ) highs.
- When an engineer is present, just say what you feel in plain words: "I want it more gorgeous," "I want a shinier sound," "I want it to feel more like summer" — any phrasing works.
Positioning, distance, environment
- Face the mic as much as possible (staring at the lyric sheet records a "sound with soft focus" — plan the geometry of music stand and mic in advance).
- But don't over-demand and make the singer stiff — better to let the singer move their body and capture a vocal with good groove (the "thump" of a foot tapping is absorbed by the studio carpet).
- Kinsetsu kōka (近接効果, proximity effect): the closer to the mic, the more the low end is emphasized — keep a moderate distance; when you want a rock flavor, deliberately recording up close is also an option.
- Height relationship: mic in front of the mouth is the baseline; mouth above the mic → picks up more lows; mouth below the mic (mic angled from above) → picks up more highs.
- Accommodate habits: for someone who usually sings holding a guitar, or who is used to singing seated, set up according to their habit.
- For a singer who turns neurotic and can't sing well because the monitoring is too revealing: deliberately downgrade the mic or increase the reverb in the cue mix.
- Cue-mix monitoring: how loud the singer's own voice is in the cans is the singer's call, but too loud makes pitch hard to track; strong rhythm parts give the vocal groove but obscure the pitch-reference parts.
- Heavy compression in the monitoring chain is not recommended: the singer can't hear their own dynamics, which hinders growth; "singing harder and harder without any tactile response" paradoxically produces a thin voice.
Vocal direction (ボーカル・ディレクション)
- Definition: the process of objectively judging the singing; traditionally handled by the record company director (ディレクター), nowadays often taken on by the arranger.
- Six duties:
- Choose the mic/preamp and prepare the environment;
- Advise on the singing itself;
- Verify pitch/rhythm accuracy;
- Plan the harmonies (ハーモニー);
- Select and decide the takes;
- Perform the treatment (トリートメント, vocal cleanup/editing).
- The most important is (2), coaching the singing — in the digital environment pitch and rhythm can be fixed later; what the machine cannot do is the feel of the words and the vocal expression that only a human can sing.
- There are only two musical criteria:
- Does it suit the song;
- Does it suit this singer (似合っているか).
- Verbal-direction technique:
- When a grand arrangement gets a mumbled vocal, don't say "sing louder" (the fader can fix volume) — what you really want may be "a brighter voice," so saying "please sing it brighter" works better;
- Image-based instructions like "sing it smiling," "corners of the mouth up," "sing toward the far distance" are effective.
- Think consonants and vowels separately (essential knowledge for both direction and treatment):
- Japanese has only 5 vowels — あ/い/う/え/お; the consonant precedes the vowel (「か」 = K + A; romaji makes this clearest);
- When you want impact, rather than "louder," saying "sing the consonants harder" works better;
- When editing, distinguish "is this a consonant-timing problem or a vowel-length problem"; in the mix, merely controlling consonant level changes the vocal's expression.
- The magic consonant "S":
- "S" ≈ white noise produced with the voice — it is essentially noise;
- Therefore you can splice takes in the middle of an S with no audible seam; its length can be stretched or shrunk freely without sounding wrong; grafting an S from a different take is also basically seamless;
- The sense of the beat lives in the vowel, not the consonant: when 「さ」 (= S + A) is sung on beat 2, the S occurs at the tail of beat 1 — what aligns with the beat is the vowel A;
- Practical takeaway: in direction/editing, spotting "there's an S here" means "this is a good splice point."
- Goal: a recorded vocal track that makes everyone happy — you cannot push forward ignoring the singer's feelings; you must also respect the intentions of the composer, lyricist, arranger, record label, and management.
- Production-chain analogy (film → pop music): script = lyrics/music; sets and costumes = arrangement (アレンジ); actor = vocalist (ボーカリスト); hair and makeup = on-the-spot vocal direction; cinematography = recording (録音); editing = treatment (トリートメント); MA = mixdown (ミックス・ダウン). "On-the-spot direction + recording + treatment" together constitute vocal direction (ボーカル・ディレクション) in the broad sense.
Pitch-correction philosophy (ピッチ補正)
- "The pitch is correct, therefore the singing is good" is incomplete; rather, good singing is mostly in tune — correct pitch is "kind" to the listener (no strain in the brain's pitch comparison).
- Key mixing lessons:
- A vocal that sings flat will not "come forward" no matter how far you push the fader;
- A vocal slightly sharp has better nuke (抜け, cut-through/presence) — good singers unconsciously sing slightly sharp.
- Therefore there is no need to correct everything to exact pitch:
- Where the accompaniment is sparse, flat or sharp moments can simply be the "inflection" itself;
- It is fine to be sharp where you want emphasis — not pulled up after the fact, but deliberately preserved: the places where the singer "couldn't help singing sharp in the heat of it";
- Not full alignment, but deliberately leaving parts untouched — that is natural correction.
- Tools:
- Industry standard ANTARES Auto-Tune: good sound, fine-grained editing; its weakness is formant (フォルマント) shifts — for large corrections, use SERATO Pitch'n Time and SOUND TOYS PurePitch alongside;
- Melodyne on the rise: easy editing, rhythm and volume can be edited together, visual checking of doubles/triples — especially needed on R&B multi-track sessions; but some engineers still say "a vocal that has passed through Melodyne stops coming forward."
- The three treatment (トリートメント) steps and their order:
- Clean up the splice points (つなぎ) between multiple takes;
- Align the rhythm;
- Fix the pitch.
- Why this order: if you move the rhythm (shift timing) first and splice afterward, the splices have to be redone; with Melodyne, once the splices are joined, rhythm and pitch are done in one pass.
- Treatment is time-consuming and needs good monitoring — the engineer taking the data home to work on it is also common.
- Removing plosive pops (ポップ・ノイズ, pop noise) (transient low-frequency noise, the "momentary EQ" method):
- Find the pop and loop-play just that short segment;
- Insert a steep high-pass filter and raise the cutoff gradually from low to high until the "pop" disappears (real example around 300 Hz; exact value 309.6 Hz);
- Offline-render and export only that segment, paste it back into the original track at the same position, and crossfade (クロスフェード) both ends to join it.
- Taming a "runaway" vocal (wild level swings):
- Compression example: WAVES Renaissance Compressor, Attack 24.9 ms / Release 16.6 ms / Ratio 4.72:1 + WAVES L1;
- Threshold is the key: set it "riding the limit" — too shallow and the level differences won't shrink; too deep and the singing turns unnatural;
- If compression still isn't enough: finish with fader automation lifting the small-waveform parts;
- In the DAW era, using no compression at all and leveling entirely with automation is perfectly viable — it can even draw complex curves compression cannot follow, with zero unnaturalness — recommended: level with automation first.
5. Vocal Mix Chain and Harmony Vocal Design
Lead-vocal mix chain (goal: the vocal "steps forward" out of the backing track)
Process in order:
- Compression (WAVES Renaissance Compressor): Ratio about 4:1 (4.27:1), Attack 16 ms, Release 160 ms (slightly slow attack);
- How to set it: don't clamp down — the reduction meter barely flickers in normal passages, and you set the threshold so it compresses 3-4 dB when the level rises in the chorus, the aim being "to slightly average out the vocal's swells."
- Limiter (WAVES L1): raise the overall loudness — the Maximizer can lift loudness while preserving the singing's inflection (the parts about to get buried come up in level).
- EQ (WAVES Renaissance Equalizer):
- High-frequency boost to emphasize the edge (エッジ) component of the voice;
- Also boost around 125 Hz in the lows (real example: 151 Hz peak, +6.9 dB) — this band "doesn't sound much like a voice" yet lifts the vocal's latent thickness and presence (a technique approaching the texture of Western singers); watch the level rise and beware of clipping.
- Short delay (AVID Mod Delay II): on a send — yields a chorus-like sense of width; too much and it "smears" (nijimi-kan にじみ感) — set it riding the threshold of audibility.
- Reverb (AVID Reverb One): finally, apply a thin layer.
- Final balance: after the vocal processing is done, if you want the backing more prominent, pull the whole vocal down slightly.
- The R&B "glued-to-the-front" method (multiband compression): WAVES C4 (4 bands, mainly emphasizing lows and highs) = acting as EQ and compression at once; the vocal "sticks flat to the front";
- Cost: the vocal expression in choruses and the like gets flattened — for pop outside R&B this risks destroying the vocal's expression;
- When doing this kind of processing, use automation to fine-tune volume and preserve the song's inherent inflection.
- De-Esser = "a compressor focused on the sibilance band" (real example: FREQ 7.0 kHz / RANGE -5.0 dB); overdone it produces a mumbling "thick-tongued" vocal — proof, in reverse, of how much consonants matter.
The four harmony (hamo ハモ) types
Terminology: hamo (ハモ) = short for harmony (ハーモニー); ue-hamo (上ハモ) = harmony above the lead; shita-hamo (下ハモ) = harmony below the lead.
- Utaiage ue-hamo (歌い上げ上ハモ, full-voiced upper harmony — upper third, etc.):
- Purpose: keep the main melody intact while making the whole brighter and more brilliant;
- The contradiction: being higher than the lead it naturally stands out, yet must not overpower the lead — singing it weakly loses the brilliance, so keep the full-voiced delivery and solve it with EQ;
- EQ concept = "make the harmony escape from the lead's frequency bands": cut the lows heavily, also carve out the "core" of the mids, and if necessary actually emphasize the highs.
- Upper harmony hugging the lead:
- Purpose: not brilliance but creating air, performing the chord feel;
- The key is changing the vocal delivery: not the projected voice (hatta koe 張った声) but a light, breathy-clean delivery;
- The volume drops, but do not compensate by moving closer to the mic (proximity effect boosts lows) — control it with gain (ゲイン).
- Understated parallel lower harmony (shita-hamo):
- Image = "the shadow cast by the lead vocal" — match delivery and diction to the lead as closely as possible;
- To reduce its presence, cut the lows (real EQ example: low cut around 110 Hz, 123 Hz -10 dB, 247 Hz -4.9 dB, 621 Hz -4.2 dB — an overall low-mid reduction).
- Ai-no-te (合いの手, interjected answer phrases — short 2-3 voice unison lines between vocal phrases):
- Positioned as "a different object" from the lead; reach for EQ first to change the texture — radio-voice EQ is fine, or even just cutting the lows transforms the atmosphere;
- To increase the unity of multiple voices: align levels with faders + compression → apply one shared spatial effect to the group → even a shared modulation effect across the whole group (chain: EQ → compression → chorus → delay).
- Think of reverb as part of the set: types (1) and (2) can use reverb/delay rich in high-frequency content to supplement brilliance; for (3), reducing the direct sound and increasing the reverb component lowers its presence.
Harmony panning strategy
- Single (シングル): with no special reason, place it center; create distance and depth relative to the lead with volume, reverb, and delay; for harmonies in extreme high/low registers, consider switching to a different mic from the lead's.
- Double (ダブル, the same harmony recorded twice):
- They can be split left/right; the more extreme the pan, the more it stands out and is emphasized even at low volume — throw the ones you want emphasized further out;
- When both upper and lower harmonies have doubles, throw only one of the pairs to the extremes — throwing both makes too many elements;
- If nothing else in the arrangement is wide and only the harmonies are, it sounds odd (exploiting this deliberately is allowed); consider the panning of the reverb along with it.
- Triple (トリプル): understand it as "single + double" — the single handles depth, the double handles width; pulling the single's volume down keeps it cleaner.
Doubling the lead vocal (ダブル)
- Trade-off: doubling a lead with superb diction dilutes the feel of the words; but the stereo solidity unique to a double carries a warmth and thickness that machine copying cannot produce.
- Workflow points:
- The lead take must be selected first — the double is recorded while listening to and hugging the lead;
- But slow take-picking leaves the singer waiting — recording the double immediately, while the singer still remembers the diction, works better (a time-versus-quality trade-off);
- For songs with intricate diction: record the double diligently in segments of a few lines at a time.
- Mic strategy (e.g., hard-voiced lead, full-voiced double): switch mics after the lead is fully recorded; or set up two mics and have the singer turn to face the other one — ready to record at any moment.
- Checking technique: pan the two vocals hard left/right, or mute the lead track, and strictly check whether the double hugs the lead.
- Direction principle: don't let the lead's diction get erased — having the singer record the double with a delivery that emphasizes the diction lands it just right.
Directing harmony recording (ディレクション)
- Phenomenon: lead singers who sing the main melody well often struggle unexpectedly on harmonies — the harmony's subtle similarity to the main melody makes it hard to memorize.
- Guide track (ガイド) — the most common solution:
- Create a new track and play the harmony line with a soft timbre with a weak attack (アタック);
- Tricks: deliberately play it an octave higher for easier hearing; loop a 1-bar line as a 2-bar loop; you can pan it to one side of the headphones;
- As the singer gets comfortable, gradually lower the guide track's volume before the real take;
- Track setup: lead vocal muted, guide track (soft timbre) turned up, backing turned down, looping over the target section;
- Once roughly learned, you can also drop the guide and quietly play back one basically-OK harmony take, recording while the singer listens along (hearing their own voice builds confidence);
- No need to stop after every line — keep loop-recording and stop when a good take appears.
- Diagrams also work: draw pitches as curves/arrows, mark long notes, flag emphasis points; write short lines in big letters on paper and hand them over; gesturing support from the control room is also common.
- Most effective: sing it to them — an instrument sound and a singing voice are entirely different things, and asking a non-instrumentalist lead singer to "mentally substitute" an instrument tone with a voice is extremely difficult, especially when the placement of the lyric syllables is intricate.
6. The Demo Three-Piece Set and Studio Data Standards
The advance-preparation set for the lead singer (three-piece set)
- Demo audio file (デモ音源ファイル) (xxxx_demo.mp3) — when including a karauta (仮歌, guide vocal), the pitch (ピッチ) and rhythm must be unmistakably clear! Otherwise the session bogs down in: "Is this rhythm on the beat (ジャスト, just), pushed ahead of the beat (食って入る, kutte hairu — anticipating), dotted and bouncing (付点で跳ねる), or just unsteady singing?" — the production intent fails to transmit and the composer ends up disappointed in themselves.
- Karaoke file (カラオケ・ファイル) (xxxx_oke_130bpm.mp3) — put the tempo in the file name; place a count-in (カウント)/click at the head (very convenient when the singer imports it into their own DAW to practice); for songs that open a cappella (アカペラ), include a pitch-reference sound along with the click (a chord or the opening single note, either works).
- Lyric sheet (歌詞カード).
- For a lead singer who gets pulled along by hearing too much karauta (which is actually a good sign — a singer should build their interpretation through their own filter):
- Omit the karauta and play the main melody with an instrument timbre instead — flute, organ, electric piano, analog synth, etc.; avoid timbres similar to the backing and choose one that stays clearly audible without being buried;
- Or hand over sheet music (the author considers this the proper form to begin with): the singer reads the composer's intent from the page and expands their own delivery from it.
Audio format standards
- Composition stage:
- Music production uses uncompressed audio; WAV is the most universal; mainstream specs: 24bit/48kHz or 24bit/44.1kHz;
- Format recommendation grades: WAV 24bit at 44.1/48kHz = ◎ (88.2/96 = ○); AIFF 24bit next; 16bit tends to sound "thin"; AIFF 32bit float (Logic's internal format) is not universal;
- Setting 96kHz gains nothing: most soft synths only support 44.1/48kHz, and converting bit depth or sample rate upward never raises quality above the original ("putting a small thing in a big box doesn't make the small thing bigger") while taxing the computer;
- Converting 32bit float to 24bit produces virtually no audible quality change — no need to worry.
- Studio stage:
- The professional-studio standard DAW is AVID Pro Tools; the studio may be set to 24bit/96kHz — when recording real instruments later, 96kHz captures higher quality than 48kHz;
- Such decisions are settled in consultation with the engineer and producer; some studios only support 48kHz — confirm in advance.
- Mastering awareness: converting a high-resolution format to the CD spec 16bit/44.1kHz inevitably changes the sound (a kind of degradation); the process that minimizes that change = mastering (マスタリング); build the habit of checking the sound separately at rough mix / after mixdown / after mastering — this power of judgment is precisely the creator's job.
Iron rules for data brought to the studio
- Soft synths must be bounced to audio beforehand: MIDI data is only performance information — without the same synth and patch at the studio, "nothing makes a sound."
- Two options for external hardware synths: (1) record their audio into the DAW at home (one synth per track, recorded one at a time); (2) haul the whole unit to the studio (possibly better sound, but laborious).
- Iron rule: "1 track = 1 audio file":
- Export every track from bar 1 beat 1 through the end of the song as one continuous audio file (the position information of scattered audio regions exists only inside your own project file);
- A track silent throughout except in the
…(truncated)