Fastify WebRTC Listen Audio Stream to OpenAI Realtime Transcribe

This blog guides you through creating a scalable audio transcription pipeline using Fastify, WebRTC, and OpenAI Whisper.

Blog cover image
2101050's avatar
2101050
14 views

Fastify WebRTC Listen Audio Stream to OpenAI Realtime Transcribe


1. Introduction: Building a Real-Time Audio Transcription Pipeline

Today’s technology landscape is increasingly shaped by the need for real-time processing and analysis of live multimedia. Audio transcription, especially real-time speech-to-text, is revolutionizing fields from accessibility to productivity, live broadcasting, online conferencing, and more.

This blog will walk you through building a fully functional, scalable, and performant browser-to-backend-to-AI audio transcription pipeline, using:

  • WebRTC/Web Audio API (for in-browser microphone capture and streaming)
  • Fastify (Node.js server for high-performance streaming and session management)
  • OpenAI Whisper (powerful speech-to-text engine, supporting real-time-like workflows)

Typical use-cases:

  • Live captioning for conferences and classes
  • Meeting assistants for automated minutes and accessibility
  • Real-time note-taking in telemedicine, support, or interviews
  • Content creation: podcasts, live shows, and streaming subtitling.

Pipeline Overview

Text
1[Browser Mic (WebRTC)] ---> [Fastify Server (WebSocket)] ---> [OpenAI Whisper API] ---> [Realtime Transcript Output]
2

This guide will show you how to implement each component, wire them together, and follow best practices for production reliability.


2. Client Setup: Capturing and Streaming Audio via WebRTC

Capturing Microphone Audio with Browser APIs

Start by using getUserMedia to request the microphone stream. This stream is then processed using the Web Audio API so that you can capture PCM chunks for streaming:

Js
1navigator.mediaDevices.getUserMedia({ audio: true }) 2 .then(stream => { 3 // Setup for audio processing 4 processAudioStream(stream); 5 }).catch(err => { 6 alert(`Microphone error: ${err.message}`); 7 }); 8

Processing the Audio Stream:

Js
1function processAudioStream(stream) { 2 const context = new AudioContext(); 3 const source = context.createMediaStreamSource(stream); 4 const processor = context.createScriptProcessor(4096, 1, 1); 5 source.connect(processor); 6 processor.connect(context.destination); 7 8 processor.onaudioprocess = e => { 9 const input = e.inputBuffer.getChannelData(0); // Float32Array 10 const int16Data = floatTo16BitPCM(input); 11 sendToServer(int16Data); 12 }; 13} 14 15function floatTo16BitPCM(input) { 16 const buf = new Int16Array(input.length); 17 for (let i = 0; i < input.length; i++) { 18 buf[i] = Math.max(-32768, Math.min(32767, input[i]*32767)); 19 } 20 return buf.buffer; 21} 22

Streaming Audio with WebSockets:

Js
1const ws = new WebSocket('wss://your-backend/audio-stream'); 2ws.onopen = () => { /* connection is ready */ }; 3function sendToServer(buffer) { 4 if (ws.readyState === WebSocket.OPEN) { 5 ws.send(buffer); 6 } 7} 8ws.onmessage = event => { 9 const msg = JSON.parse(event.data); 10 if (msg.transcript) displayTranscript(msg.transcript); 11}; 12

Best Practices

  • Chunk audio in small packets (e.g., 50–200ms) for minimal latency.
  • Request permissions clearly and handle denials with helpful instructions.
  • Handle audio device errors gracefully.

References:


3. Building the Fastify Server: Streaming, Buffering, and Pipelines

Setting Up Fastify with WebSocket Support

First, ensure you've installed Fastify and its websocket support plugins:

Bash
1npm install fastify fastify-websocket ws axios form-data 2

Minimal WebSocket Audio Server Example:

Js
1const Fastify = require('fastify'); 2const fastifyWs = require('fastify-websocket'); 3const axios = require('axios'); 4const FormData = require('form-data'); 5 6const app = Fastify(); 7app.register(fastifyWs); 8 9app.get('/audio-stream', { websocket: true }, (conn, req) => { 10 let buffer = Buffer.alloc(0); 11 conn.socket.on('message', async data => { 12 buffer = Buffer.concat([buffer, data]); 13 if (buffer.length > 32000) { // Roughly 1s at 16kHz 16bit mono PCM 14 const transcript = await transcribe(buffer); 15 conn.socket.send(JSON.stringify({ transcript })); 16 buffer = Buffer.alloc(0); 17 } 18 }); 19}); 20

Buffering and Session Handling

Adapt buffering size depending on your target latency vs API limits. For security and scale:

  • Use JWT or API tokens for authentication.
  • Use wss:// (over HTTPS) for all audio data.
  • Implement rate limits and error handling for abusive or broken connections.

4. Integrating with OpenAI Whisper: Real-Time Transcription

Relaying Batched Audio to Whisper

OpenAI Whisper API supports standard audio file formats. Here's how you transmit to it from Node.js:

Js
1async function transcribe(audioBuffer) { 2 const form = new FormData(); 3 form.append('file', audioBuffer, { filename: 'audio.wav' }); 4 form.append('model', 'whisper-1'); 5 const response = await axios.post( 6 'https://api.openai.com/v1/audio/transcriptions', 7 form, 8 { 9 headers: { ...form.getHeaders(), 10 'Authorization': `Bearer ${process.env.OPENAI_API_KEY}` } 11 } 12 ); 13 return response.data.text; 14} 15

Chunk, POST, Await, and Respond: For every buffered audio segment, POST to Whisper and immediately relay the transcription to the client.

Whisper API Notes:

  • Supports batch (1–10s) audio for "pseudo real-time" workflows.
  • Streaming endpoints may be in beta—monitor OpenAI Whisper docs for updates.

5. Best Practices: Reliability, Security, and Monitoring

  • Latency vs throughput: Shorter chunks reduce latency, but may lower accuracy for very short phrases.
  • Backpressure handling: If API calls slow, maintain audio queue and warn users of delays.
  • Authentication: Always require user tokens before accepting audio streams.
  • TLS/HTTPS: Never send unencrypted audio data in production.
  • Resource cleanup: Detect and close idle connections, free any associated buffers.
  • Compliance: Inform users about cloud audio usage (GDPR/etc. as needed).
  • Error Reporting: Surface API or network errors cleanly on the client and in server logs.

References:


6. End-to-End Example and Testing

Launching the Full Stack

  1. Start the Fastify server with your OpenAI API key.
  2. Open your client HTML/JS in a modern browser.
  3. Grant microphone permissions. Start speaking. Inspect transcript output, which should appear within 1-3 seconds per phrase.

Debugging tips:

  • Validate your audio format using FFmpeg if conversions are failing.
  • Monitor buffer sizes and latency with console debugging.
  • Check OpenAI API quotas and handle 429 or 413 errors by backing off and chunking less frequently.

Adapting from Public Projects

Reference:


7. Conclusion & Recommended Resources

With Fastify, WebRTC, and OpenAI Whisper, you can build robust, scalable real-time speech-to-text solutions for modern web applications. Expand further by integrating diarization, translation, or speaker identification, and contribute improvements back to the open-source ecosystem.

Further Reading:

Recommended Articles

Discover more articles you might find interesting

Implementing LangGraph REST API with FastAPI
Technical Insights

Implementing LangGraph REST API with FastAPI

This guide provides a comprehensive implementation plan for building a LangGraph REST API using FastAPI, covering environment setup, agent definitions, endpoint creation, testing, and deployment.

2101050
Jun 18
153
Read More
DeepSite v2 Practical Guide
Technical Insights

DeepSite v2 Practical Guide

A comprehensive guide to DeepSite v2, covering its features, installation, and advanced workflows.

2101050
Jun 21
112
Read More
Fastify OpenTelemetry: Logging, Metrics, and Tracing in Practice
Technical Insights

Fastify OpenTelemetry: Logging, Metrics, and Tracing in Practice

Learn how to implement logging, metrics, and tracing in Fastify using OpenTelemetry.

2101050
Jul 11
106
Read More
Creating Diverse Logo Designs with Flux Model and ComfyUI
Technical Insights

Creating Diverse Logo Designs with Flux Model and ComfyUI

Learn to leverage the Flux model and ComfyUI for unique logo designs through effective prompts and examples.

2101050
Jan 10
93
Read More
Formatting Dates in TypeScript to UTC
Technical Insights

Formatting Dates in TypeScript to UTC

A guide on how to format dates in TypeScript to the specific format YYYY-MM-DDTHH:mm:ss+00:00.

2101050
Dec 19
83
Read More
Implementing a Custom Chat Model with LangChain
Technical Insights

Implementing a Custom Chat Model with LangChain

This guide provides a comprehensive blueprint for creating a custom chat model by subclassing LangChain's BaseChatModel, including configuration, method overrides, and error handling.

2101050
Jun 17
78
Read More