The /generate endpoint created the voice prompt before loading the
user's requested model size. Since create_voice_prompt() internally
calls load_model_async(None), it fell back to the hardcoded default
of "1.7B", causing the 1.7B model to be downloaded even when the
user explicitly selected 0.6B.
This reorders the operations so the requested model is loaded first,
ensuring create_voice_prompt() and generate() use the correct model.
Co-Authored-By: Claude Opus 4.6 <[email protected]>