Building a Real-Time Speaking Language App with Expo and OpenAI Agents

Step-by-step guide to create a language app that uses speech recognition and OpenAI's conversational agents.

Blog cover image
2101050's avatar
2101050
38 views

Building a Real-Time Speaking Language App with Expo and OpenAI Agents (2025)

Introduction

This guide demonstrates, step-by-step, how to build a robust, real-time speaking language app using Expo (React Native) and OpenAI’s conversational agent APIs. The final product will allow users to speak into their device, transcribe their input to text, interact conversationally with an OpenAI agent, and hear the agent’s response read aloud—enabling instant, interactive AI-powered language experiences on mobile.


Project Setup and Core Dependencies

Overview

Get started with a modern Expo project. Clone a real-world example (joblesspoet/react-native-openai-speech-totext), install the necessary dependencies, and prepare your environment for both iOS and Android using EAS Dev Builds if using custom native modules.

Key Steps

  • Create/clone project:
    Bash
    1expo init real-time-agent-app 2cd real-time-agent-app 3
  • Install dependencies:
    Bash
    1expo install expo-speech 2npm install react-native-voice axios 3
  • Obtain OpenAI API key: Store securely in a .env file and use libraries like react-native-dotenv for access.
  • Configure permissions: Android: Add <uses-permission android:name="android.permission.RECORD_AUDIO" /> iOS: Update Info.plist for NSSpeechRecognitionUsageDescription and NSMicrophoneUsageDescription.

Reference Structure

Project folders like /src/hooks, /src/components, /src/api organize speech, API, and UI logic for maintainability.


Implementing Speech Recognition and Handling Audio Input

Practical Voice Capture

Speech-to-text recognition is achieved using react-native-voice in custom Expo Dev Builds. A React hook (e.g., useSpeech) manages permission, speech event binding, and cleanup:

Typescript
1import Voice, { SpeechResultsEvent } from '@react-native-community/voice'; 2 3export const useSpeech = () => { 4 /* ... state setup ... */ 5 useEffect(() => { 6 Voice.onSpeechResults = (data: SpeechResultsEvent) => setResult(data.value?.[0] ?? ''); 7 /* ... other handlers and cleanup ... */ 8 }, []); 9 return { speaking, result, startSpeechToText, stopSpeechToText }; 10}; 11

Usage:

Jsx
1<Button onPress={speaking ? stopSpeechToText : startSpeechToText} title={speaking ? "Listening..." : "Speak"} /> 2<Text>{result}</Text> 3

Testing and Pitfalls:

  • Always test on hardware, as emulators may lack mic support.
  • Handle permission denials gracefully and prompt user recovery.
  • Manage overlapping states to avoid mic/TTS conflicts.

Integrating with OpenAI Agents: Text Streams and Conversational Context

Conversational AI Interaction

Interact with OpenAI’s chat/completions API or the Assistants API for session-aware conversation:

Typescript
1import axios from 'axios'; 2export const sendToOpenAI = async (messages) => { 3 const res = await axios.post("https://api.openai.com/v1/chat/completions", { 4 model: "gpt-3.5-turbo", 5 messages 6 }, { headers: { Authorization: `Bearer ${OPENAI_API_KEY}` } }); 7 return res.data.choices[0].message.content; 8}; 9

Workflow:

  1. Maintain message context as an array of user/assistant turns.
  2. Send the updated context (“messages”) after each user utterance.
  3. Stream replies (SSE/eventsource) for low-latency interaction, or get full replies at once.

Error handling:

  • Detect network/API failures and display clear feedback.
  • Show loading spinners on pending agent replies.

Text-to-Speech (TTS) and Real-Time Response Playback

Making AI Talk

Use Expo’s built-in expo-speech for instant voice feedback:

Javascript
1import * as Speech from 'expo-speech'; 2Speech.speak(agentReply, { 3 language: 'en-US', // or any supported locale 4 rate: 1.0, 5 pitch: 1.0, 6}); 7

Code Patterns:

  • Trigger TTS automatically after assistant reply updates.
  • Allow manual replay and stopping via a ‘Speak’ button.
  • Clean up with Speech.stop() when navigating away or interrupting playback.

Example UI integration:

Jsx
1<Button onPress={() => Speech.speak(reply)} title="Speak" /> 2

Building the User Interface and Orchestrating the Real-Time Cycle

UI Orchestration

Combine all components—speech recognition button, message chat UI, TTS playback, and agent message state—into a cohesive, responsive user experience.

ChatScreen/Container Pattern:

  • Maintain all messages and states in hooks/context.
  • Provide a clear “mic” button for input, show recognition results live.
  • Use a chat list (FlatList) for history, with user and assistant bubbles.
  • Attach “Speak” playback controls to assistant replies.
Jsx
1<FlatList data={messages} /* ...renderItem=MessageBubble... */ /> 2<SpeechButton listening={listening} onPress={toggleSpeech} /> 3

Focus on clear visual feedback—loading spinners, error messages, speech/TTS state highlights, and scroll-to-latest on new messages.


Testing, Deploying, and Extending Your Speaking Language App

Testing

  • Device testing: Validate all speech, audio, and permissions scenarios on real Android/iOS.
  • Edge cases: Offline, interrupted, permission-denied, and rapid input states.
  • Automated E2E: Tools like Detox help automate repeated UI/state flow checks.

Deployment

  • Use Expo EAS for final builds: eas build --platform android / --platform ios
  • Secure API keys; ideally, proxy agent calls through a backend in production.
  • Prepare all privacy, permission, and data usage disclosures for App Store/Play Store submission.

Extending Functionality

  • Multi-language input/output: Let users select and persist language pairs.
  • Session/threaded chat: Use OpenAI Assistants threads for persistent context.
  • Image generation: Integrate DALL·E endpoints for “draw by description” features.
  • Accessibility: Provide full screen-reader support, large voice/visual toggles, and clear error feedback.
  • History and export: Store chats locally and enable export for learning review.

References & Further Reading

Recommended Articles

Discover more articles you might find interesting

Implementing LangGraph REST API with FastAPI
Technical Insights

Implementing LangGraph REST API with FastAPI

This guide provides a comprehensive implementation plan for building a LangGraph REST API using FastAPI, covering environment setup, agent definitions, endpoint creation, testing, and deployment.

2101050
Jun 18
153
Read More
DeepSite v2 Practical Guide
Technical Insights

DeepSite v2 Practical Guide

A comprehensive guide to DeepSite v2, covering its features, installation, and advanced workflows.

2101050
Jun 21
112
Read More
Fastify OpenTelemetry: Logging, Metrics, and Tracing in Practice
Technical Insights

Fastify OpenTelemetry: Logging, Metrics, and Tracing in Practice

Learn how to implement logging, metrics, and tracing in Fastify using OpenTelemetry.

2101050
Jul 11
106
Read More
Creating Diverse Logo Designs with Flux Model and ComfyUI
Technical Insights

Creating Diverse Logo Designs with Flux Model and ComfyUI

Learn to leverage the Flux model and ComfyUI for unique logo designs through effective prompts and examples.

2101050
Jan 10
93
Read More
Formatting Dates in TypeScript to UTC
Technical Insights

Formatting Dates in TypeScript to UTC

A guide on how to format dates in TypeScript to the specific format YYYY-MM-DDTHH:mm:ss+00:00.

2101050
Dec 19
83
Read More
Implementing a Custom Chat Model with LangChain
Technical Insights

Implementing a Custom Chat Model with LangChain

This guide provides a comprehensive blueprint for creating a custom chat model by subclassing LangChain's BaseChatModel, including configuration, method overrides, and error handling.

2101050
Jun 17
78
Read More