フォルマントシフト & 口腔容積保護機構 Formant Shift & Acoustic Protection Engine
「何故、声をピッチシフトすると口に物を含んだような声になるのか?」への回答とアプローチ Preventing Unnatural Acoustic Artifacts in Real-time Formant Shifting
概要と搭載アプリケーション Overview & Applications
本アルゴリズムは、音声やボーカル音源のキー(ピッチ)を変える際、声質(フォルマント)を維持したり、逆にキャラクター(女声↔男声)を自由に変えるために独自開発された処理モジュールです。単なる全帯域の周波数伸縮ではなく、人の声道(Vocal Tract)の物理的・音響的特性に合わせた保護措置を組み込んでいます。
This module preserves natural vocal timbres or alters vocal characteristics (e.g., female to male) during pitch shifting. Instead of uniform frequency scaling, it incorporates acoustic protection mechanisms based on human vocal tract physics.
1. 口に物を含んだような違和感 ➔ F1/F2 緩和ロジック(口腔スペース保護) 1. Muffled Mouth Artifact ➔ F1/F2 Shift Mitigation
【問題】
一般的なフォルマントシフターで女声のピッチを下げて男声に近づけようとした際、「何かを口に含んで喋っているような不自然な重み」や「モコモコした籠もり感」が発生する。
音声スペクトル包絡(フォルマント)を全体的に均一の倍率で低域シフトさせると、「声帯から口唇に至るすべての口腔内共鳴空間の体積が機械的に一括拡大」されてしまうことが原因です。
【対処(アプローチ)】
全周波数帯域を一律にシフトさせるのではなく、F1(第1フォルマント:顎の開き)およびF2(第2フォルマント:舌の位置)が集中する低中域(〜約3,000Hz以下)のシフト量を緩やかに抑制するダイナミック・ウェイト制御を導入しました。
これにより、声質のキャラクター(骨格感)を変えつつも、口腔内の空間的体積が過剰に膨張するのを防ぎ、言葉の明瞭性と自然な発音を確保します。
[Problem] Uniformly down-shifting spectral envelopes to lower pitch creates a muffled "talking with food in mouth" artifact due to mechanical expansion of the entire simulated vocal tract volume.
[Solution] Our engine introduces dynamic weight scaling that moderates shift ratios in the F1/F2 region (~3 kHz and below), preserving natural pronunciation and diction without unnatural acoustic cavity expansion.
2. 音色変化・高音潰れ・ノイズ強調 ➔ 各自動保護ロジック 2. Timbre Distortion & Noise ➔ Automated Safeguard Filters
【問題】 フォルマント移動処理を常に一律で適用すると、以下の3つの音質劣化が発生する。
1. ドラムや子音のアタック領域で周波数が平滑化され、アタックが「にじみ」パーカッションの音色が極端に変化する。
2. ピッチを極端に上げた際、ナイキスト周波数付近で包絡線が密密になり高音が潰れ、金属性ノイズが発生する。
3. 無音〜極小音量部分でフォルマント補正が過剰に働き、残留ノイズが強調される。
- アタック保護(Transient Protection): ゼロ交差率(ZCR)やエネルギー急変から立ち上がり領域を検出し、アタック瞬間のみフォルマント干渉を自動減衰。
- 高域破綻防止(High-Frequency Safeguard): 高域成分(>50% Nyquist)の帯域重みを自動減衰させ、クリアな高域を維持。
- 弱音保護(Low-Energy Safeguard): 微小信号(-80dB〜-40dB)に対してSmoothstepアルファ制御を適用し、ノイズ強調をシャットアウト。
[Problem] Continuous envelope filtering causes attack blurring in transients, high-frequency crowding near Nyquist, and noise amplification in silent parts.
- Transient Protection: Reduces envelope interference during initial attack transients.
- High-Frequency Safeguard: Attenuates shifted weights near Nyquist to prevent metallic distortion.
- Low-Energy Safeguard: Smoothstep alpha fading between -80dB and -40dB prevents noise boosting.
3. 保護が裏目に出る場面(二重声問題) ➔ スルー(バイパス)ロジック 3. Artifacts in Vocal Restoration ➔ Protection Bypass (Formant Keep)
【問題】
「ピッチシフトした音源の声を元に戻す(キー変更された曲のボーカルフォルマントだけを原音に近づける)」目的でフォルマントシフトを用いる場合、上記の自動保護ロジックが裏目に働く。
変調された声と保護されたドライ音(原音)が混ざり合ってしまい、「声が2つ重なって聞こえる二重声(ダブルボイス)現象」が発生してしまう。
【対処(アプローチ)】
用途に応じて「自動保護ロジックをあえて完全にスルー(バイパス)し、フォルマント包絡線のみをストレートに適用・保持する設定(Formant Keep /
スルー制御)」を選択できる自由度を設計に組み込みました。
これにより、二重声を完璧に回避しながら目的のボーカルキー補正を達成できます。
[Problem] When restoring pitch-shifted vocal tracks back to original formants, automatic protection logic can mix dry features with shifted envelopes, causing a bizarre double-voice artifact.
[Solution] Our engine provides an explicit protection bypass switch ("Formant Keep" mode) to apply pure envelope correction without mixing or double-voice artifacts.
4. ケプストラム分析の2つのトレードオフ ➔ 破綻防止 & 動的ピッチ保護機構 4. Dual Cepstral Trade-offs ➔ Artifact Prevention & Dynamic Pitch Protection
フォルマント包絡線の精度を決める「ケプストラム分析次数(係数)」の調整には、「① 次数の大小による静的トレードオフ」と「② 適用範囲(固定 vs 動的)による時間軸トレードオフ」という2つの構造的な課題が存在します。
- 次数が大きすぎる場合: 声道のフォルマントだけでなくピッチ成分(基音・倍音ピーク)まで包絡線に巻き込んでしまいます。その結果、ヴァイオリンやミュートトランペット等の複雑な倍音を持つ楽器で位相が不自然にずれたり、ソプラノの高音域ボーカルで声が二重に聞こえる音質破綻が発生します。
- 次数が小さすぎる場合: 包絡線が平滑化され過ぎてしまい、フォルマントシフトの効果が聴感上ほとんど現れなくなります。
- オーディオ全体で固定適用した場合: 楽曲セクションごとの音質ギャップが発生します(例: ボーカルセクションは綺麗にシフトするが、バイオリンソロに入ると倍音の位相が狂って不自然になる)。
- 分析窓(フレーム)ごとに動的変更した場合: 窓ごとに次数が急変することで、再生音全体に「チカチカ・パチパチ」とした音質のブレ(ちらつきノイズ・フリッカー)が発生します。
- ピッチ領域の動的自動保護(①の解消): フレームごとにケプストラム領域の基音ピークを検出。強烈なピッチ成分を検知した際、分析次数をピッチ周期の直前領域へ動的制限し、倍音の位相崩れや二重声を自律的に遮断します。
- Fast Attack / Smooth Release 制御(②の解消): ピッチ検出時は即座に保護を発動(Fast Attack)して不自然な破綻を防止し、検出解除時は IIR スムージングで滑らかに復元(Smooth Release)することで、フレーム切替時のちらつきノイズを完全に抑制。
- 手動ケプストラム係数調整の開放: ユーザーが音源の性格に合わせて「ピッチ保護優先」か「ケプストラム解像度(フォルマント効果)優先」かを自在に選択できる自由度を提供しています。
Setting the cepstral order requires balancing two structural trade-offs: (1) Magnitude Trade-off (Order Size) and (2) Temporal Scope Trade-off (Global vs Dynamic).
- Orders Too High: Fundamental and harmonic pitch peaks leak into the spectral envelope. Instruments with complex harmonics like violins or muted trumpets suffer unnatural phase shifts, while high soprano vocals exhibit double-voice artifacts.
- Orders Too Low: Over-smoothing occurs, reducing audible formant shift resolution.
- Global Static Setting: Creates section mismatches (e.g. clean vocal shifting, but harmonic distortion during a violin solo).
- Naive Dynamic Frame Switching: Frame-by-frame order jumps introduce jarring acoustic flickering and popping transients.
- Dynamic Pitch Domain Protection (Resolves Axis 1): Real-time quefrency inspection detects pitch peaks and dynamically clamps active coefficients just below the detected pitch peak to eliminate phase distortion and doubling.
- Fast Attack / Smooth Release (Resolves Axis 2): Fast attack prevents pitch corruption immediately, while smooth IIR release eliminates frame-to-frame switching flicker.
- Manual Cepstral Order Control: Empowers users to balance pitch protection priority against maximum envelope detail.
5. JavaからC++への移植とKissFFT換装 ➔ ネイティブ高速化とFFT移植の壁 5. Java to C++ Migration & KissFFT Adoption ➔ Native Optimization
【問題】
初期のプロトタイプはJava(JTransformライブラリ)で構築していたが、モバイル端末での処理速度向上とバッテリー消費抑止のため、ネイティブC++(JNI)への完全移植が必要に。
しかし、JavaのJTransformとC++の軽量ライブラリ「KissFFT」では、複素数の配列構造や正規化スケール係数、実数FFTのパッキング仕様が全く異なり、移植後に計算結果が一致せず音質が崩壊するトラブルに直面。
【対処(アプローチ)】
AI(LLM)と何度も対話を重ねて原因を徹底調査。FFT順変換から対数スペクトル・逆FFT(ケプストラム抽出)に至る各ステップのデータフォーマットと正規化処理を一工程ずつ検証・修正しました。
泥臭い検証の末、KissFFTに最適化した高精度なC++ネイティブエンジンへの換装を完了。AndroidのCPUパフォーマンスを限界まで引き出し、リアルタイムでの超高速フォルマント補正を実現しました。
[Problem] The prototype built in Java (using JTransform) needed full C++ native migration for mobile CPU/battery optimization. However, switching to KissFFT caused major calculation mismatches due to differences in complex array packing, normalization scaling, and real FFT layouts.
[Solution] Through extensive iteration with AI (LLM) assistant, we audited every step from forward FFT to log-spectrum cepstrum inverse FFT. We successfully re-architected the pipeline into a high-performance, lightweight KissFFT-based C++ native engine.
6. アプリの目的に応じた実装のグラデーション(UX & 処理負荷の最適化) 6. Product-by-Product Implementation Gradation
DS Softでは、すべてのアプリに一律の処理を強制するのではなく、「アプリの目的・UI操作感・リアルタイム性能(CPU負荷)」のバランスを考慮し、フォルマントシフト機能の組み込み度合いをあえてグラデーション(選択的適用)させています。
設計思想: リアルタイム処理の限界よりも「音質とカスタマイズ性」に全振り。
F1/F2緩和、アタック保護、弱音保護、ケプストラム逆算、手動係数調整まで、すべてのパラメーターを開放したフラッグシップ仕様。
設計思想: 動画再生のリアルタイムUXを損なわないCPU負荷と高音質の両立。
「ケプストラムの逆算適用」およびボーカルキー変更時の二重声を防ぐ「Formant Keep(保護スルー制御)」に絞って高機能化。
設計思想: ファイル読み込みとスライダー操作の「超高速反映」に全振り。
処理の遅延要因となる重い自動保護ループを排し、「ケプストラム係数の手動補正」など軽量かつダイレクトに効く処理のみを厳選実装。
Rather than enforcing uniform processing across all software, DS Soft tailors engine implementations to match individual app goals, user interfaces, and CPU load profiles:
Prioritizes maximum quality and full parameter customization over real-time constraints.
Balanced real-time video playback with Cepstral Inverse Scaling and Formant Keep mode.
Optimized for instantaneous file loading and slider updates, retaining manual cepstral tuning.
おわりに:実際の音でお試しください Conclusion: Experience the Sound
これらの理論や仕組みが実際に音としてどのような効果をもたらしているか、ぜひご自身の耳でお試しください。
その上で、もしご意見やご感想をいただけましたら、それが今後の開発の原動力になります。
Experience how these algorithms and design ideas actually shape the sound in our
applications.
If you have any feedback or thoughts, we would love to hear from you—it directly fuels our
ongoing development!