
Closed
Posted
Paid on delivery
I’m ready to move from concept to prototype on a voice-first mobile application that will run natively on both iOS and Android. The core interaction is simple to describe yet technically demanding: a user presses the mic, speaks, their speech is converted to text with minimal latency, that text is sent straight to an LLM, and the model’s response is returned immediately inside the app. The entire flow must feel seamless and natural for a general-public audience who expect consumer-grade polish. Here is the flow I need implemented: • Microphone capture with push-to-talk (or similar) that works reliably across current iOS and Android versions. • On-device or cloud speech-to-text with high accuracy; if you have experience with Whisper, Google Speech, or Apple’s Speech framework, please highlight it. • Secure, efficient hand-off of the transcribed text to an LLM endpoint (OpenAI, Claude, or another you recommend) and return of the model’s answer. • A clean, lightweight UI that shows the running transcript and the LLM’s response in real time. Quality targets • End-to-end round-trip (speak → reply on screen) under three seconds on a decent connection. • Error handling that guides the user gracefully when connectivity drops or speech isn’t recognised. • Codebase structured for future expansion (e.g., adding text-to-speech or user accounts later). Deliverables 1. Fully functioning MVP for both iOS (Swift / SwiftUI preferred) and Android (Kotlin / Jetpack Compose preferred). 2. Clear setup notes and build scripts so I can run the project locally. 3. A short video demo and testflight / APK links for initial user testing. If you have shipped voice-driven or conversational AI apps before, or have benchmark data around speech-to-text latency and LLM throughput, that’s exactly the expertise I’m after. Let’s create an experience that feels like talking to the future.
Project ID: 40602807
71 proposals
Remote project
Active 1 day ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
71 freelancers are bidding on average ₹56,118 INR for this job

Hello, I understand you're looking to build a polished, voice-first mobile prototype for iOS and Android, focusing on a seamless push-to-talk to LLM response flow. My expertise in high-accuracy speech-to-text (including Whisper and Apple's Speech framework) combined with efficient LLM integration can deliver the sub-three-second, consumer-grade experience you envision. I am prepared to build a scalable MVP using SwiftUI and Jetpack Compose, ensuring robust error handling and a future-proof architecture. I’m Waqas from Eclairios, a professional software engineer with over 7 years of experience in app and web development. I have successfully completed 128 projects, earning a 5.0 rating from satisfied clients. I specialize in mobile apps (Android, iOS, Flutter), website development, custom APIs, and backend solutions. My goal is to deliver high-quality, scalable solutions that meet your business needs. Why hire me? ★ 100+ Projects Completed with 5-star rating. ★ 3 months of free post-launch support ★ Expertise in advanced technologies and systems Let’s connect and discuss how I can help you with your project. Best regards, Waqas
₹44,310 INR in 7 days
8.4
8.4

This looks like a great fit, Hi, I will deliver a native iOS (SwiftUI) and Android (Jetpack Compose) MVP with push to talk mic capture, Whisper for speech to text, and a secure LLM relay returning responses under three seconds round trip. You will get TestFlight and APK links, build scripts, and a short video demo. On a recent voice AI build, streaming the Whisper transcript in chunks while the LLM call fired in parallel shaved over a second off perceived latency. I will wire yours the same way. Questions: 1) Do you already have an OpenAI or Claude API key, or should I set up the relay to support both so you can compare? 2) For the lightweight UI transcript view, do you have any design references or is a clean default layout fine for the MVP? Ready to start whenever you are. Kamran
₹42,251 INR in 13 days
7.5
7.5

Hi, Your project closely matches the kind of AI voice solutions we've been building recently. We've developed a conversational AI voice platform that combines Speech-to-Text, LLM-powered response generation, and Text-to-Speech into a real-time conversation flow. The solution supports natural voice interactions, knowledge-based responses, call handling, secure APIs, and is designed for low-latency communication. For your application, our focus would be on creating conversations that feel natural rather than scripted, while keeping the architecture scalable for future enhancements such as multilingual support, CRM integrations, analytics, and conversation history. Low response latency and reliable voice quality are critical for a good user experience, so we'd design the solution around those priorities. Before we begin, I'd like to understand: Is this a phone-based voice assistant or an in-app voice experience? Do you already have conversation flows and a knowledge base, or should we help design them? Are there any telephony or CRM integrations required? I'd be happy to discuss the architecture and recommend the best implementation approach. Regards, Avinash Jain Procode IT Services
₹145,000 INR in 30 days
7.0
7.0

Hi I will be able to help you. Please message me so that we will have detail technical discussion. I have 9+ years of combined experience in Mobile Application development, Website development, Desktop application development, 3rd party Artificial Intelligence api, AR/ VR, Chatbot, Blockchain- Cryptocurrency, CRM & ERP, Game Development and any other Software development. I am having expertise in Native on Android Java, kotlin and IOS Swift, and For Hybrid Cross platform on Flutter Dart & React- Native and for web and backend on react js and node js, Python Django. Please consider me and initiate a chat for further detailed discussion. Regards, Anju Logical Soft Tech Pvt Ltd, Indore(M.P)
₹40,000 INR in 20 days
6.6
6.6

I’ve built similar voice-first apps where users tap to speak, and speech is accurately transcribed and instantly sent to AI models for replies. Reducing lag is key—I’ve worked with Whisper and Google Speech to keep the end-to-end conversation under three seconds on good networks. For your app, I suggest using on-device recognition where possible (like Apple’s Speech framework on iOS) to cut latency, then routing transcriptions securely to the LLM via a lightweight API. Is there a preferred LLM or are you open to recommendations with proven low response times? Also, would you prefer real-time transcript updates as the user speaks, or only after they finish? I’ll ensure robust error handling so users always know what’s happening, even if their connection drops or speech isn’t clear. The code will be clean and modular for future features like TTS or accounts. I’ll deliver working MVPs with SwiftUI and Jetpack Compose, plus clear build instructions and test/demo links. Ready to start bringing this to life.
₹37,500 INR in 7 days
5.3
5.3

As an experienced Full Stack Developer and Digital Solutions Expert with an expertise in Mobile App Development for both iOS and Android, I am well-equipped to handle your project. I have been in the industry for over 8 years now and have successfully completed more than 200 projects. This vast experience has not only sharpened my tactical execution but also bolstered my problem-solving skills which would be crucial to implementing the complex functionality required in your app. In addition to delivering a fully functioning MVP that satisfies your specific requirements, I also believe in long-term support. I am skilled at providing clear setup notes, build scripts and most importantly efficient error handling - ensuring effective communication of any limitations or dropped connections that might impact user experience. Lastly, your vision to create an application that feels like "talking to the future" aligns perfectly with my passion for enhancing user experiences through cutting-edge technologies like AI. Combined with your speaking-optimized flow and my solid track record of building voice-driven apps - we can surely turn this vision into reality! Thanks once again for considering my profile and should you have any queries, do not hesitate to reach out. Let's convert our potential synergy into concrete solutions that astonish users with its seamlessness!
₹57,500 INR in 7 days
4.8
4.8

With over 5 years in the industry as a talented Full Stack Web and Mobile App Developer, I am confident that I can fulfill your need for a seamless and polished voice-first mobile application. I have a broad spectrum of skills that align directly with your project requirements, especially in React Native and Flutter for mobile development, which will be beneficial for both iOS and Android aspects of this project. Not only have I successfully built SaaS applications, e-commerce platforms, but also other conversational AI interfaces from databases management to API integration. I have a proven track record of delivering clean, optimized code, on time and offering clear communication throughout to ensure complete satisfaction. Optimizing performance is one of my key strengths; this skill will be valuable in ensuring the end-to-end round trip is within the 3-second threshold you've set for optimal user experience. Additionally, django_gender_code's clean code was mentioned as an important technique to master in order to improve productivity while developing software.
₹56,250 INR in 7 days
4.3
4.3

You want push-to-talk STT → LLM → on-screen response under three seconds, built native in Swift/SwiftUI and Kotlin/Jetpack Compose — that's exactly the voice AI stack we've shipped. We've built conversational apps using Whisper for STT and OpenAI API for LLM responses, integrated natively into iOS and Android. We know the latency traps in that pipeline and how to avoid them — audio capture, chunking, and streaming responses. No questions on the Clarification Board yet. Ping us and we'll walk through our STT → LLM approach and share benchmark latency data.
₹62,000 INR in 30 days
4.4
4.4

Hi there, we are a team of Full Stack Web and Mobile App developers and we can do this project in no time. Thanks Ashish Kumar.
₹56,250 INR in 7 days
3.8
3.8

Hi, I’m Karthik (15+ yrs experience). I specialize in low-latency voice AI architectures and mobile development. I can build your cross-platform conversational voice MVP for iOS (SwiftUI) and Android (Jetpack Compose) meeting your sub-3-second round-trip latency requirement. Technical Approach Audio Capture: Native AVAudioEngine (iOS) & AudioRecord (Android) with reliable push-to-talk state management. Sub-300ms STT: WebSocket streaming via Deepgram Nova-2 or Whisper (Groq API) for ultra-fast, accurate transcription. (Fallbacks: Native SFSpeech / Android SpeechRecognizer). LLM Integration: Streaming connection (OpenAI GPT-4o-mini / Claude 3.5 Sonnet) to stream response tokens directly to the UI in real time. Clean UI & Architecture: Reactive UI built in SwiftUI & Jetpack Compose using Clean MVVM architecture—making it simple to plug in TTS or user accounts later. Expected Latency (<1.2s Total) STT Stream: ~300ms | LLM First Token: ~400ms | Total Round-Trip: ~1.0–1.2s Deliverables (10 Days) Native iOS & Android MVP source code. Setup guide & local build scripts. TestFlight link, Android APK, and video demo. Ready to start immediately!
₹74,350 INR in 7 days
3.1
3.1

Hello there, We will build your voice AI app for iOS and Android: mic capture, speech-to-text, LLM integration, and a real-time transcript UI. For latency, we will stream Whisper transcription in chunks rather than waiting for full utterance completion. This shaves roughly a second off the round-trip. We will also open a persistent connection to the LLM endpoint so responses begin rendering as tokens arrive. A couple of quick things to confirm: 1) Do you prefer native builds (Swift + Kotlin) or would a single React Native codebase work for the MVP stage? 2) Which LLM provider do you want us to wire up first (OpenAI, Claude, or both)? The number quoted here is a starting estimate. The exact cost and timeline will be confirmed after we go through the full scope together. Looking forward to discussing further. Best regards, Faizan
₹42,365 INR in 13 days
2.9
2.9

Hi, I have 7+ years of experience in React Native and AI integrations. I've built real-time voice apps using Whisper/OpenAI, LiveKit, WebRTC, and streaming APIs. I can deliver a fast, cross-platform MVP with speech-to-text, LLM integration, and a scalable architecture.
₹56,250 INR in 7 days
3.5
3.5

At our company, we don't just tout our experience - we back it up with tangible results and happy clients. With over 7 years in the industry, specializing in mobile app development for both iOS and Android platforms, we've created impressive, robust applications that are intuitive, efficient, and tailored to the target audience- like your project demands. Designing an app that converts speech to text with minimal latency and delivers model-based responses immediately is exactly the kind of challenge that excites us. We're familiar with the Whisper API, Google Speech, Apple's Speech framework, and more- perfect for delivering a high-accuracy speech-to-text experience for your app users. In line with your quality targets, we aim to create an end-to-end round-trip that's below three seconds - even accounting for any hiccups or connectivity drops. Lastly, not only can we deliver your MVP on time with full functionality (and easy setup notes!), but we can provide a short video demo for user testing as well. Our attention to detail extends beyond coding - we want users to be able to use your app effortlessly while their conversation avec l'AI is seamlessly achieved. Add in our consistent on-time delivery and our proven track record of long-term support and investing in us becomes not just a smart decision but one worthy of the future outlook you have for your project.
₹56,250 INR in 7 days
2.6
2.6

As an AI specialist and mobile app developer for over six years, I possess the knowledge and experience necessary to propel your vision of a conversational voice AI app forward. My expertise in building autonomous agents, mobile applications, and full-stack platforms has consistently been applied to deliver production-ready solutions with a high degree of polish and efficiency. I am no stranger to managing large, complex projects that rely heavily on AI integration, making me well-suited for this task. Key competencies such as my experience with OpenAI's GPT-4o or Claude endpoint - in addition to languages like Kotlin, Swift, SwiftUI and libraries like Google Speech or Apple's Speech framework either on-device or cloud implementation - align closely with your project's requirements. Furthermore, I'm familiar with delivering quality within tight timeframes, which encourages me to proactively build solutions for potential challenges such as transient connectivity issues or inaccurate speech recognition. This commitment guarantees you'll have an app with minimal latency, robust error handling system, and excellent round-trip time.
₹37,500 INR in 1 day
2.7
2.7

As a seasoned mobile app developer with significant expertise in both iOS and Android platforms, I am well-prepared to take your voice AI app from concept to reality. Having built apps that operate smoothly and natively on both platforms, I can guarantee a streamlined and reliable microphone capture feature that is consistent across multiple versions of iOS and Android. Furthermore, my experience with frameworks such as React Native and Kotlin, along with my competence in utilizing the PostgreSQL, MongoDB and Firebase database systems aligns perfectly with this project's requirements. Not only can I build an efficient speech-to-text system for your app, but my strong backend skills will also ensure secure and seamless hand-off between the transcribed text and your desired LLM endpoint.
₹38,000 INR in 7 days
2.3
2.3

Namaste I’m Bojan and I love building apps that let people talk naturally You need a iOS and Android app that records speech, instantly transcribes it, sends the text to an LLM and shows the reply I will use SwiftUI with Apple Speech on iOS and Kotlin with Google Speech on Android, latency under one second A lightweight HTTPS client will forward the text to your LLM and update the UI instantly with prompts for network or errors Deliverables are a MVP for iOS and Android, build scripts, a demo video and TestFlight and APK links I delivered a similar prototype for ₹58 000 Do you prefer Apple Speech or Google Speech for transcription and do you already have an LLM API key Let’s have a quick call to start the sprint
₹50,000 INR in 2 days
2.5
2.5

Greetings, I have reviewed your project description and recently worked on a similar project. I believe I can help you deliver this successfully. Let’s open a chat to discuss your requirements in detail and determine the best approach for your project. Demo can me provided. Regards
₹56,250 INR in 7 days
2.1
2.1

Hi, I can build a production-ready cross-platform voice AI MVP with low-latency speech-to-text, seamless LLM integration, and a clean, scalable architecture that delivers a fast, natural conversational experience on both iOS and Android.
₹56,250 INR in 7 days
2.2
2.2

Getting that speak-to-reply loop under three seconds is the real engineering challenge here, and it's very doable with the right architecture choices. I'll build this in React Native so both platforms share one codebase while still giving you native mic access and consumer-grade polish. For STT, I'd go with Whisper's API over Google Speech for this use case. It handles accents and background noise better, and streaming the audio in chunks keeps latency tight. The transcribed text hits the LLM endpoint, and I'll stream the response token-by-token so the user sees the reply forming in real time rather than waiting for the full completion. One thing worth thinking about early: if you plan to add text-to-speech later, structuring the response handler as a stream now saves a full refactor. Best regards, Shayan
₹41,250 INR in 12 days
1.8
1.8

The three‑second round‑trip often fails when the speech‑to‑text step falls back to a slower network path. I'll wire the mic to the Speech framework on iOS and SpeechRecognizer on Android, then forward the text to the LLM via an HTTP call. A short in‑memory buffer will keep the UI responsive while the model streams its answer back. A common mistake is assuming the OS will auto‑retry failed speech requests, which often leaves the user staring at a frozen screen. I'll add explicit retry logic and clear UI cues so the app recovers gracefully from network hiccups. You can expect a smooth push‑to‑talk flow that stays under three seconds on a typical 4G link.
₹45,000 INR in 4 days
2.0
2.0

India
Member since Dec 19, 2017
₹1500-12500 INR
$250-750 USD
$250-750 AUD
$3000-5000 USD
$250-750 USD
₹750-1250 INR / hour
$250-750 USD
min $50 USD / hour
₹1500-12500 INR
$250-750 USD
₹100-400 INR / hour
$30-250 USD
$15-25 USD / hour
$30-250 USD
$150-450 USD
₹100-400 INR / hour
$25-50 AUD / hour
$1500-3000 USD
₹150000-250000 INR
$8-15 USD / hour
₹1500-12500 INR