"stream": true to get tokens as they generate. Buffered JSON is the default.
Streaming uses the same auth, model slugs, and credit billing as a non-streamed call. The HTTP timeout on inference routes is long enough for a multi-minute completion.
OpenAI
stream=True / stream: true.