Local Whisper + Local Llama = <500ms end-to-end
Whisper large-v3 runs on your Mac's Metal GPU or your PC's DirectML/CUDA. Llama 3.1 8B is quantized and fine-tuned for interview vocabulary. First token in ~460ms — 10× faster than cloud copilots users complain about on Trustpilot.