Flutter voice assistant: fast speech recognition, interruption, and a large knowledge base #209883
Replies: 5 comments 1 reply
|
This sounds like a fantastic project. Building real-time voice apps in Flutter is tricky but incredibly rewarding right now. I can share some insights, especially regarding the WebRTC architecture and the specific bug you are hitting.
What is likely happening is either the emulator's OS is suppressing the virtual microphone while the speaker is active, or the microphone is picking up the narration audio (loopback), turning it into a muddy signal that the OpenAI VAD (Voice Activity Detection) rejects as background noise rather than human speech. Immediate Diagnostics: The Headphone Test: Build the app to a physical Android device and test it while wearing wired/Bluetooth headphones. If barge-in works perfectly with headphones but fails when using the phone's external speaker, your issue is AEC. The mic is hearing the app's own narration. WebRTC Constraints: Make sure your getUserMedia constraints explicitly enable hardware echo cancellation: {"audio": {"echoCancellation": true, "noiseSuppression": true}}. Audio Focus: Check your Android audio_session or just_audio configuration (if you are using those). Ensure the audio routing is set to "communications" mode (like a phone call), not "media" mode (like Spotify). Communications mode activates the phone's hardware echo cancellers.
Traditional pipelines (STT -> LLM -> TTS) are much easier to debug because you can see the text at every step, but they naturally stack latency (often 1.5 to 2 seconds). WebRTC S2S gets you down to ~300-500ms (Time to First Audio).
OpenAI will automatically fire a speech_started event over the WebSocket/WebRTC data channel. Your Flutter app must listen for this event and instantly command your local audio player to stop() and flush the remaining audio queue. OpenAI handles the context automatically—they figure out exactly which word the assistant was cut off at, truncate the transcript, and append the user's new interrupted thought.
Pre-process your respiratory care study guides into small chunks (paragraphs). Create vector embeddings for these chunks (OpenAI has a cheap embeddings API for this). Store them in a vector database (like Supabase, Pinecone, or even locally on the device using SQLite VSS). The Flow: User asks a question -> You quickly query the vector database for the top 3 most relevant paragraphs (takes ~100ms) -> You use OpenAI Realtime's "function calling/tools" feature to seamlessly inject those 3 paragraphs into the context -> The AI answers out loud.
Good luck with the app! Respiratory care education is a great use case for this tech. Let us know what happens when you test it on a physical device. good luck..... |
|
I would start with the missing speech events before changing the overall architecture. Active RTP counters/energy are useful, but do not establish that the injected words survive the capture/processing path. 1. Separate turn detection from response generation. For your application-controlled setup, check the effective For one reproduction, log incoming data-channel event types and timestamps before your UI/state-machine filters: 2. Run a small comparison, changing one thing at a time. Use the same harmless phrase (e.g. “stop reading now”) before narration, during narration, and immediately after stopping playback, in the same session. Repeat on a physical device with wired headphones, then speaker output. A headphone/speaker difference narrows routing/echo/processing suspects; it does not prove AEC is the cause. On the emulator, document how you inject speech. Host-speaker playback, host-microphone capture, and an injected audio track are different paths. Android documents that the virtual microphone uses host audio only when its host-input option is enabled. For a capture diagnosis, a short synthetic sample from the actual outgoing capture path, if your instrumentation supports it, is stronger evidence than a second recorder using a different source. Avoid adding a competing microphone recorder to the test. Record route, audio-focus changes and native audio source/mode around playback startup. 3. Identify who owns the narration. Is the chapter read through the Realtime remote audio track, or through a separate Flutter/Android TTS/player? OpenAI documents automatic unplayed-audio truncation for its WebRTC/SIP output on interruption. My implication for your app: a separate chapter player still needs its own stop/queue-clear logic, and any chapter-position context must be managed by your app. For a manual stop of Realtime output, the client event reference specifies A useful sanitized reproduction would contain the acknowledged turn-detection settings, that event timeline, playback owner, injection method, and the before/during/after results. That can distinguish client event filtering from a capture-path or service-side question without publishing your app source. |
|
Update October 10 — thanks for the diagnostic suggestions. We found and fixed a reproducible client-side narration bug, but have not yet verified the complete live-microphone interaction. Reproduced failure: a speech-start/VAD event speculatively paused the guide reader. When the final transcript matched our own recently narrated text, the echo filter rejected it; the reader nevertheless remained paused. This is distinct from a genuine missing speech-start event. Applied repair:
Results: the original regression failed before the repair. Afterward, 99 affected voice/host tests and 82 surrounding reader/native-control/progress tests passed; static analysis was clean. These fixtures exercise reader ownership but use synthetic transport, so they do not prove recognizer accuracy, audible interruption or echo cancellation. The fresh debug build is installed on the retained emulator. Live acceptance is still pending. A broader regression also had Windows packaging/inventory errors, so we are not claiming a release-ready app. The separate chat-answer path remains a concern: it can stop output on raw speech-start before the final echo decision, and it does not share the guide reader's recovery ownership. Earlier live tests were mixed: narration was audible, one Stop worked, another was misrecognized, and unsolicited self-interruptions occurred. Those are historical observations, not a successful test of this new build. What would you test next to separate residual self-echo from genuine user barge-in? Is an ownership-fenced speculative pause/recovery appropriate for a separately owned narrator, and what pattern would you use for streamed chat output that cannot simply be resumed? We'd appreciate a small licensed Flutter/WebRTC reference or a precise event-timeline checklist. We can provide a sanitized minimal reproduction; no credentials, recordings or private guide content are included here. |
|
Hi! Since you're already using OpenAI Realtime with Flutter, I'd focus on debugging the audio input before changing your architecture. A few suggestions:
For latency, measure the time between the user's last spoken word and the first audible response. Native speech-to-speech is a good starting point, while STT → LLM → TTS offers more control. Useful references: OpenAI Realtime documentation: Flutter WebRTC (MIT License): I'd start by comparing VAD event logs from the emulator and a physical device. That should help narrow down where the issue originates. Hope this helps! |
|
Thank you for the concrete diagnostic checklists, especially the distinction between playback ownership and cancellation. I am following up with a request for public reference code so this can be evaluated against a known baseline. Does anyone have a maintained, openly licensed Flutter/Android example that demonstrates all of the following?
A link to specific sample files or a small reproducible harness would be more useful than another broad architecture change. I am seeking free advice/reference examples; this is not a request to disclose private app source or a claim that a particular SDK is defective. Thank you. |
Uh oh!
There was an error while loading. Please reload this page.
🏷️ Discussion Type
Question
Body
I'm a respiratory-care professional developing a Flutter Android education app with coding assistance. I'm seeking volunteer technical guidance on making its voice assistant responsive while supporting a large domain of knowledge. I am using Grok, ChatGPT, Muse, and similar apps as examples of the conversational experience I want: quick responses, strong speech recognition, and the ability to interrupt while the assistant is talking.
I'm interested in documented architecture patterns and properly licensed code examples or reference implementations. I understand the full implementation of a commercial assistant may be private. I'm keeping my full app source private as well; this is a request for free technical advice, not paid services.
What I need help understanding:
There is also a concrete issue in my current implementation:
Current setup: Flutter/Dart Android app; flutter_webrtc pinned to 1.6.2+hotfix.3; WebRTC speech session using OpenAI Realtime; semantic VAD with application-controlled responses and speech interruption configured. Client/server configuration does not intentionally disable turn detection during narration.
What diagnostics would distinguish Android/WebRTC capture or audio-processing suppression from a provider-side VAD/event problem? Which audio source/mode, audio-focus, echo-cancellation, or emulator-input checks would you try first?
I can prepare a small sanitized reproduction and redacted diagnostics based on what would be most useful. Public replies with suggestions or relevant experience would be appreciated. Thank you.
Guidelines
All reactions