faster-whisper-xxl.exe --language English --model large-v2 --ff_vocal_extract mdx_kim2 --vad_method pyannote_v3 --standard <input>
Use Japanese, then translate: faster-whisper-xxl.exe --language Japanese --task translate --model large-v2 --ff_vocal_extract mdx_kim2 --vad_method pyannote_v3 --standard <input>
1. https://github.com/Purfview/whisper-standalone-win
This might seem pretty obvious, but I wasn't expecting it to matter as much since the end result was compressed again (and yet again on my car's Bluetooth) using a different codec.
This isn't on anything close to audiophile gear either. My car stereo is thoroughly mediocre and my home stereo is speakers I literally found in the trash.
Now I'm (belatedly) rebuilding my music collection in FLAC...