Building a Real-Time Speaking Language App with Expo and OpenAI Agents (2025)
Introduction
This guide demonstrates, step-by-step, how to build a robust, real-time speaking language app using Expo (React Native) and OpenAI’s conversational agent APIs. The final product will allow users to speak into their device, transcribe their input to text, interact conversationally with an OpenAI agent, and hear the agent’s response read aloud—enabling instant, interactive AI-powered language experiences on mobile.
Project Setup and Core Dependencies
Overview
Get started with a modern Expo project. Clone a real-world example (joblesspoet/react-native-openai-speech-totext), install the necessary dependencies, and prepare your environment for both iOS and Android using EAS Dev Builds if using custom native modules.
Key Steps
- Create/clone project:
Bash
- Install dependencies:
Bash
- Obtain OpenAI API key:
Store securely in a
.envfile and use libraries likereact-native-dotenvfor access. - Configure permissions:
Android: Add
<uses-permission android:name="android.permission.RECORD_AUDIO" />iOS: UpdateInfo.plistforNSSpeechRecognitionUsageDescriptionandNSMicrophoneUsageDescription.
Reference Structure
Project folders like /src/hooks, /src/components, /src/api organize speech, API, and UI logic for maintainability.
Implementing Speech Recognition and Handling Audio Input
Practical Voice Capture
Speech-to-text recognition is achieved using react-native-voice in custom Expo Dev Builds. A React hook (e.g., useSpeech) manages permission, speech event binding, and cleanup:
Typescript
Usage:
Jsx
Testing and Pitfalls:
- Always test on hardware, as emulators may lack mic support.
- Handle permission denials gracefully and prompt user recovery.
- Manage overlapping states to avoid mic/TTS conflicts.
Integrating with OpenAI Agents: Text Streams and Conversational Context
Conversational AI Interaction
Interact with OpenAI’s chat/completions API or the Assistants API for session-aware conversation:
Typescript
Workflow:
- Maintain message context as an array of user/assistant turns.
- Send the updated context (“messages”) after each user utterance.
- Stream replies (SSE/eventsource) for low-latency interaction, or get full replies at once.
Error handling:
- Detect network/API failures and display clear feedback.
- Show loading spinners on pending agent replies.
Text-to-Speech (TTS) and Real-Time Response Playback
Making AI Talk
Use Expo’s built-in expo-speech for instant voice feedback:
Javascript
Code Patterns:
- Trigger TTS automatically after assistant reply updates.
- Allow manual replay and stopping via a ‘Speak’ button.
- Clean up with
Speech.stop()when navigating away or interrupting playback.
Example UI integration:
Jsx
Building the User Interface and Orchestrating the Real-Time Cycle
UI Orchestration
Combine all components—speech recognition button, message chat UI, TTS playback, and agent message state—into a cohesive, responsive user experience.
ChatScreen/Container Pattern:
- Maintain all messages and states in hooks/context.
- Provide a clear “mic” button for input, show recognition results live.
- Use a chat list (
FlatList) for history, with user and assistant bubbles. - Attach “Speak” playback controls to assistant replies.
Jsx
Focus on clear visual feedback—loading spinners, error messages, speech/TTS state highlights, and scroll-to-latest on new messages.
Testing, Deploying, and Extending Your Speaking Language App
Testing
- Device testing: Validate all speech, audio, and permissions scenarios on real Android/iOS.
- Edge cases: Offline, interrupted, permission-denied, and rapid input states.
- Automated E2E: Tools like Detox help automate repeated UI/state flow checks.
Deployment
- Use Expo EAS for final builds:
eas build --platform android/--platform ios - Secure API keys; ideally, proxy agent calls through a backend in production.
- Prepare all privacy, permission, and data usage disclosures for App Store/Play Store submission.
Extending Functionality
- Multi-language input/output: Let users select and persist language pairs.
- Session/threaded chat: Use OpenAI Assistants threads for persistent context.
- Image generation: Integrate DALL·E endpoints for “draw by description” features.
- Accessibility: Provide full screen-reader support, large voice/visual toggles, and clear error feedback.
- History and export: Store chats locally and enable export for learning review.






