Plain Whisper accepts a finished file and returns text after a few seconds or minutes — good enough for transcribing a recorded call, but not for live captions or a voice assistant. Streaming transcription processes audio in pieces as it arrives and shows text with a fraction-of-a-second delay. Let's see how to build such a pipeline on your own VDS.
How streaming transcription differs from batch processing
The basic Whisper setup, described in the article on transcribing audio and video on a VPS, works with a whole file: the model sees the entire context at once and produces the most accurate result. A stream, on the other hand, never ends — a microphone, a call, a broadcast — and you cannot wait for a complete recording.
The solution is to cut the stream into short 2-5 second segments, recognize each one nearly independently, and stitch the result together. The price for speed is that the model sees less context, so accuracy at segment boundaries is a bit lower than with batch processing.
Chunking audio and detecting speech (VAD)
Cutting the stream strictly by a timer is a bad idea: a word can end up split in half between two chunks. Voice Activity Detection (VAD) finds pauses in speech and cuts exactly there. The silero-vad library detects phrase boundaries in milliseconds and needs no GPU.
pip install silero-vad faster-whisper sounddevice numpy
python3 -c "import torch; model, utils = torch.hub.load('snakers4/silero-vad', 'silero_vad')"
A typical scheme: a buffer accumulates audio, VAD marks the end of a phrase, the accumulated chunk goes straight to the recognition model, and the buffer is cleared.
faster-whisper in streaming mode
faster-whisper fits streaming better — a CTranslate2 implementation that recognizes short segments 2-4 times faster than the original PyTorch Whisper code. On a CPU server with 8 cores, the small model processes a 3-second chunk in about 0.5-0.8 seconds — enough for captions without a noticeable delay.
from faster_whisper import WhisperModel
model = WhisperModel("small", device="cpu", compute_type="int8")
segments, info = model.transcribe("chunk.wav", language="en", beam_size=1)
for seg in segments:
print(seg.start, seg.end, seg.text)
beam_size=1 together with the small or base model is a required trade-off for streaming: the full-size large-v3 with a wide beam search adds a delay of several seconds per chunk, which breaks the idea of real time.
A WebSocket server for receiving audio
The client — a browser or an app — sends audio in pieces over WebSocket, the server accumulates a buffer, cuts it by VAD, and sends the text back over the same connection. A minimal server with websockets and asyncio:
import asyncio, websockets
async def handle(ws):
buffer = bytearray()
async for chunk in ws:
buffer.extend(chunk)
if len(buffer) > 48000 * 2 * 3: # ~3 seconds of 16-bit audio
text = transcribe_chunk(bytes(buffer))
await ws.send(text)
buffer.clear()
asyncio.run(websockets.serve(handle, "0.0.0.0", 8765).__aenter__())
In production, use the VAD from the previous section instead of manual buffer-size accumulation — it cuts on speech pauses, not a fixed byte count.
Latency versus accuracy: what to choose
| Chunk size | Latency | Accuracy | When to use it |
|---|---|---|---|
| 1-2 sec | minimal | lower, words get cut | live captions, assistant |
| 3-5 sec | moderate | good | calls, meetings |
| 8-10 sec | noticeable | high | lectures, podcasts |
| whole file | after recording | maximum | archival transcription |
For most real-time caption tasks, a sensible trade-off is a 3-second chunk with the small model on CPU or medium on GPU.
Where to plug the finished transcript
Real-time text is useful in three scenarios: live broadcast captions, voice input for a chatbot, and searching call recordings. If the transcript is meant as input for an LLM agent — for example, so a bot can answer by voice to a question spoken aloud — it is convenient to route the result to an OpenAI-compatible endpoint through LocalAI, combining speech recognition and answer generation in one API. To control latency in production, it helps to hook up AI service monitoring and track each chunk's processing time as a separate metric.
- Live captions — a WebSocket server plus VAD plus the
small/basemodel. - Voice input for a bot — the same pipeline, but accumulating until the end of a phrase before sending it to the LLM.
- Searching call recordings — you can go back to batch processing: latency is not critical, accuracy matters.
Checklist for launching streaming transcription
Before rolling streaming out to production, check four things: what latency the scenario actually allows, whether CPU is enough or a GPU is needed, whether VAD cuts words at boundaries, and what happens when a connection drops — the buffer must not be lost silently.
- Determine the acceptable latency for the scenario: 1-2 sec for an assistant, 5-10 sec for lecture captions.
- Pick the model size for your resources:
base/smallon CPU,medium+ on GPU. - Set up VAD instead of timer-based cutting — it leaves fewer broken words at the seams.
- Add reconnection and buffer recovery for when the WebSocket connection drops.