Devs Hunting for Good Local STT/TTS Models
Developers and language learners are struggling to find optimal, locally-run speech-to-text and text-to-speech models. They face challenges with model size, accuracy, features (like diarization), language support, and licensing restrictions, often requiring significant experimentation and fine-tuning. The desire for self-hosting and avoiding cloud APIs is a key driver.
SOURCES (60)
“The best I’ve tried so far is vibevoice, but it is a little trickier to set up than most and you have to find an unofficial one on huggingface because they took down the larger parameter versions from GitHub.”
“Before Submitting [x] I searched open and closed issues and discussions for an existing report. [x] I reproduced this bug on the latest release AND on the current dev branch, right before submitting this report. I did not just check an old version or rely on a check from days ago. [x] I understand that maintainers want a well written issue before any code pull request. [x] This is not a security vulnerability. Installation Method Git Clone Open WebUI Version dev at 015dbc8 Operating System Linux”
“Bug Report: Inconsistent speech_started / speech_ended Behavior with Semantic VAD (Realtime API)”
“See https://choosealicense.com/no permission/. I want to reference the decoding logic for my own implementation, and from my understanding, I can't (legally) do that right now. I suggest MIT.”
“Hello community, I have been thinking of building an app for the company that just gets the calls from the automated answering machine to text, but I have huge problems with the quality of the translation. That's why I think I can use this model. My idea is to have a summary table every 20-30 calls about the content or important recalls I need to do. I run a really busy office, we get about 100 cals a day if not more. I want to get that but the devils are in the details, what do you think? W”
“Tiel Coder 35B A3B from Peculiar Ragdoll is just Ornith 1.5 35B A3B with their own "coding focused" imatrix quantization and the sharp chat template built in. But how good really is that quantization? Other quantizers also focus on coding. Perhaps it would be better to just get Ornith quantized from Bartowski or AtomicChat, proven quantizers, and then add the chat template yourself rather than get the Tiel Coder weights? The end result would be the same, just the quantization is differ”
“Could anyone recommend any free open source projects for voice changers that allow you to generate your own models? I don't want real-time conversion or anything, just something to let me record and process short audio clips. I've had some success with GPT SoVITS but it's just text to speech so it misses a lot of pronunciations. I hear ElevenLabs is a good industry standard, but I'd rather not spend any money on AI if it could be done manually. <Delete this if you have flared”
“I'm part of the Lokutor team that built this. Model: NVIDIA Conformer-CTC Small (13M params, int8). It runs on an ESP32-S3 with 8 MB PSRAM, no GPU or NPU. LibriSpeech WER is 3.7 / 8.2, versus 6.3 / 15.9 for Whisper tiny.en on a laptop. Under real noise (DEMAND: car, kitchen, cafeteria) plus babble and reverb, mean WER is 8.4 vs 12.1 for Whisper tiny.en. You can try the exact chip arithmetic on your laptop mic with live_demo.py. https://github.com/lokutor-ai/oido submitted by /u/S”
“Please add Parakeet TDT 0.6b v3 , fast model with good quality to transcribe text. Faster that Whisper. Model: Parakeet TDT 0.6b v3 Here a app using this voice to test how sound and how fast is. https://github.com/notune/android transcribe app”
“I see that its a very out-of-distribution audio and indeed it does sound robotic. We are currently improving the model and this will help us very much! Thank you! I will get back to you with more news soon.”
“#2 is odd -- can you explain how Swift's tokenizer differs? I wouldn't think they would attempt tokenizer surgery on the model for a fine-tune. What error was ds4 throwing?”
“Hey guys! Today, we are releasing SupraTTS-0.1-Beta , a tiny ~29.6M parameters Text-To-Speech model. The audio quality is a bit better than the original Glow-TTS (the architecture our model is using!) while it's keeping the same size. Here are some samples: https://huggingface.co/SupraLabs/SupraTTS-0.1-Beta#samples Link to the model on HF: https://huggingface.co/SupraLabs/SupraTTS-0.1-Beta I hope you can do something useful with it, e.g. on small edge devices and on CPU. Feel free to give us”
“In my previous post (Integration tuned for running Assist on a local LLM in 8 GB VRAM) I described my fully local-LLM Assist setup and mentioned that the Seeed reSpeaker gave me the best speech recognition of the three s…”
“Little side project I’ve been doing outside of ebook2audiobook but anyway idk I guess it’s ready to share at this point Web GUI or CLI Has a docker cause yall like that So far the tts engines I’ve got supported are XTTS v2, XTTS v1, Piper TTS, VITS, MMS / Fairseq VITS, FastPitch, FastSpeech 2, Glow-TTS, DelightfulTTS, Tacotron2 DCA, Tacotron2 DDC, Tacotron2 Capacitron, Overflow, SpeedySpeech, FastSpeech, Align TTS, NeuralHMM-TTS You can even use ebook2audiobook to generate synthetic data of 1 vo”
“I use the orcarouter uncensored version of Qwen3.8:27b. I think that is as much "fine tune" needed... The base version is actually really good as it is!”
Trying it out with audio.cpp but man is it garbage for audiobooks.
“Phonetic feedback for accent training is a tough technical challenge to get right. no pressure, but https://peerpush.com is a decent home for language learning stuff the crowd there pokes at.”
“Please fix basic voice input and Read Aloud stability before constantly changing models”
“Seconded. I really liked his old Gemma 2b model. Hope he makes one with e2b someday”
“It has multi language/region/localization per library and title, but we'll need to do a lot of testing for all of that.”
“I’ve been playing with Nemotron 3 Diarization , and it fills a gap I’ve had with local voice agents: keeping track of who is speaking. It’s a diarization model, so it gives you speaker labels rather than transcriptions or people’s names. It can stream its output and track up to eight speakers. I’ve been trying it with one-second streaming chunks, and the quality has been really good in my tests. I plugged it into my speech-to-speech setup on a DGX Spark and connected it to a Reachy Mini. The fun”
“This is a port of NVIDIA's Parakeet model family that runs end-to-end on the JVM. And it's fast , competitive with the native engines on CPUs. Steve Jobs' 15 minutes Stanford speech transcribes in ~15s in my laptop, using the 110M Q4_K model. # Transcribe anything jbang jinfer@qxoticai \ --model mudler/parakeet-cpp-gguf/tdt_ctc-110m-q4_k.gguf \ --transcribe %{https://qxotic.ai/snippets/jfk.wav} Also supports streaming transcription (as in the video above), where it continually update”
“I'm putting together a video editing mcp but the problem is, e.g with parakeet v3, is it's trying to be too smart. It seems to guess at what makes a sentence, removes um's and such. But we need those. What better models are there for this usecase? are there any models that potentially also do tagging, like, (laughs), so it gives more context for speaking cadence and such? submitted by /u/Zeeplankton [link] [comments]”
“Before Submitting [x] I searched open and closed issues and discussions for an existing report. [x] I reproduced this bug on the latest release AND on the current dev branch, right before submitting this report. I did not just check an old version or rely on a check from days ago. [x] I understand that maintainers want a well written issue before any code pull request. [x] This is not a security vulnerability. Installation Method Docker Open WebUI Version v0.11.4 Operating System Debian 13 Brows”
“MOSS-TTS Nano is still my favorite TTS model for cloning speakers. It's reasonably high quality at a very fast speed and supports 20 different languages. I'm surprised more people aren't talking about or using it. It was updated just last July to very little fanfare and no one talks about it despite it being what I would consider the "best" voice cloning TTS out there.”
“Dude, come on! There might be some use cases for having it. Maybe doing some deep research and summarization work.”
“Searching for an ASR model I came around them and they claim some crazy numbers in not only asr streaming: https://moondream.ai/blog/photon-2-1-speech-recognition but also for vision. Really curious if you guys also tried it or is it another engine with exaggerated numbers? submitted by /u/FerLuisxd [link] [comments]”
“I’ve been working on a local voice assistant setup: VAD → faster-whisper STT → local agent/LLM → TTS. Host: Debian 13 desktop GPU: NVIDIA RTX 3060 LLM runtime: Ollama Agent layer: OpenClaw / local voice agent Speech-to-text: faster-whisper TTS: local inference TTS Naturally, instead of giving it a normal assistant personality, I made the tts model train on Madea's voice. This was already a questionable engineering decision. I noticed the STT occasionally picking up speech when nobody was tal”
“I flashed a echo dot gen2 with echomuse and I’m now looking for a local pipeline that runs completely on a cpu. 2 years ago I started using vosk but stopped when my speaker died. Is vosk still the way to go when pre-de…”
“Please add Parakeet TDT 0.6b v3 , fast model with good quality. Model: Parakeet TDT 0.6b v3 Here a app using this voice to test how sound and how fast is. https://github.com/notune/android transcribe app”
“nanosamur.ai is a speech-to-tech platform that supports different models for realtime, semi-realtime, and batch transcription and provides a unified stack for robust speech processing, agentic workflows, and webhooks.You can run it locally on your computer via electron app + docker compose, but you can take the same stack and run it in cloud/k8s and/or on prem as it is built with scale in mind (it also has an observability stack built in). So in that sense it is more like Ollama + Ollama Cloud :”
“Nice comparison—the conversational part is what makes voice agents useful, and they usually don’t need polished typed prompts; colloquial speech is enough. For desktop work I’ve been trying a third path: speak on the phone and send the transcript to the PC cursor, rather than needing a mic or building a full assistant. FlowMic ( https://flowmic.app ) is one option; I help with its public launch. How are others handling input—keyboard, desktop STT, or phone-to-cursor?”
“Introduction My audiobook pipeline has to decide who speaks each line of dialogue in a novel, so the TTS can cast a voice per character. It runs on Gemma 4 26B-A4B (QAT Q4) with the experts parked in system RAM. Ling 3.0 tiny looked like the opposite bet: 7.9B total, 1.3B active, hybrid KDA/MLA attention, 4.8 GB at Q4_K_M, so the whole model sits in VRAM on an 8 GB card and the prompt gets processed on the GPU instead of the CPU. I gave it the same three chapters as the incumbent, plus a smoke c”
“I was looking to retire my google minis for some time with fast local replacement, unfortunately Onju Voice wasn’t “resposnive” I’m expecting voice assistant to be able to act as fast as human would, e.g. no delay after …”
“I’d like to share a custom integration that adds Meta’s Muse Voice Transcribe as a native speech-to-text engine for your Assist pipelines. Why you might like it: Truly streaming: audio is forwarded in 80 ms frames whi…”
“Hi everyone, Disclaimer: I’m not an expert on XMOS or audio DSP, but I did try hard to run experiments and debug this issue, with help from Claude Code. If I’ve misunderstood something, please correct me! I’m working o…”
“Hi, I have been looking into the HA TTS functionality and like that very much. I have been using it a bit now but as its just now streaming to my Nokia8000 TV box its no satisfying. The most annoying thing is that the N…”
“Summary Please add a TTS backend that can send text to a user configured HTTP endpoint and play back the returned audio. This would let people use self hosted TTS servers such as Kokoro FastAPI or Chatterbox TTS Server instead of being limited to Windows voices or Azure. Motivation Windows voices are limited in quality and variety, and Azure needs an account and can cost money. Local open source TTS models have improved a lot and can run on consumer hardware. Many of them ship with a local serve”
“I tried google and on almost each mode.people complains it’s not good enough? is there any good tts right now that support English and can express emotions for audio and not read it in monotone voice? submitted by /u/Alarmed_Wind_4035 [link] [comments]”
“This surprised me yesterday idk if a new feature or something especially with GPT6 or just a “Work project” feature. I use long voice dictation to go over a research paper and update with new ideas. But long context I often lose data from limited context and memory. So it’s a Here’s what it did: I was voice a long few hours session as I said. This time I opened a new project in the work mode and told it to keep a running tally of all the important parts and make that tally into a actual MD file”
“Description Description: When using the "OpenAI Compatible" addon to connect to strict local ASR servers (like FastFlowLM), dictation fails silently with a "No speech detected" error. I traced this to two separate limitations in how the client formats its requests: Missing File Extension on Audio Uploads: When the client POSTs to /v1/audio/transcriptions, the multipart form data passes the audio file without a standard extension (like .wav or .webm) in the filename. Strict local LLM/ASR backends”
“Clear and concise description of the problem Feature Request: Add Fish Audio as a TTS Provider It would be great if AIRI could support Fish Audio directly as a TTS/voice provider. Fish Audio has very good voice quality and would be especially useful for users who want more natural or expressive voices. Ideally, the provider configuration could include: Fish Audio API key Model selection Voice selection / voice ID Base URL if needed Test voice button This would also make it much easier than havin”
“Introduction My audiobook pipeline that has to decide who speaks each line of dialogue in a novel, so the TTS can cast voices per character. It's been running on Gemma 4 26B-A4B (QAT Q4). Bonsai 2 27B looked like it should win: a 27B-class model in 5.9 GB means a stronger base model fully resident on my 8 GB card, instead of a 4B-active MoE streaming experts over DDR5. I gave it four chapters. Here's what happened. Setup Hardware: NVIDIA T1000 8GB — Turing TU117, ~6.6 GB free. Ryzen 5 76”
I was watching Andrej Karpathy's introduction to neural networks https://www.youtube.com/watch?v=VMj-3S1tku0 and at some point it clicked that I was hearing distracting noises in it, typing and mouse clicks. I knew macOS has…
“what do you mean by ngram split, ie how do you picture it without impacting the quality if it makes sence?”
“If you ended up with one of the cheap unofficial ESP32-S3 “xiaozhi”-style AI toy boxes (sold on Alibaba/AliExpress under SKU ostb-xiaozhi-3st) and want to run it as a native ESPHome voice assistant instead of the stock x…”
“Thanks for shipping direction based emotion support for IndexTTS 2.5 so quickly. It's working well for automated per line emotion switching in my tests. After testing it for a while, I noticed two gaps versus IndexTTS 2's full capability: 1. No manual per dimension emotion vector control. Review Chunks currently exposes a single "emotion strength" scalar plus the natural language direction text. IndexTTS 2 supports 8 independent emotion dimensions (happy/angry/sad/afraid/disgusted/melancholic/su”
“Hi, thanks for feedback. So the new gemini is indeed quite good with consuming content, say a pdf. You can smoothly navigate a document, you can switch between reading and talking about it. But if you want to, say, highlight some passages or take notes, I didn't manage to get it to produce an annotated pdf (or any kind of file). If you tell it to remember a list of things, it will just randomly read that list to you at various places. Your list of notes will be present in the conversation, b”
“One practical footnote for anyone shipping this rather than benchmarking it: Latvian is already in the language list of NVIDIA's Parakeet TDT v3, along with Lithuanian, Estonian and the rest of the 25 European languages it covers. No fine tune needed, and it runs several times faster than Whisper large on the same hardware because it is a transducer rather than an encoder decoder. Worth putting in the same table before deciding that a 273 hour fine tune is the only route. Where the Whisper r”
Granola (meeting recorder) and Willow Voice (Speech to Text)
“Problem Multilingual/bilingual users who switch between languages in voice input (e.g. English + Hungarian) currently get inconsistent local Whisper STT results. Short voice clips often don't carry enough signal for Whisper's automatic language detection to confidently rank the intended language against its full 99 language vocabulary, so it sometimes locks onto and decodes as a completely unrelated language. Concretely, in testing on a self hosted instance: a plain English clip scored highest f”
“I've been mostly translating as a way to learn simple phrases, rather than long texts. I've been mostly using Russian to English and vice versa. In the Hy-MT2 huggingface, they have a "Hy-MT2 Translation Task Instruction Examples." And I've been using mostly using the one below, and it has worked fine for me: Translate the following text into Russian . Note that you should only output the translated result without any additional explanation : {source_text}”
“Update to my previous post a couple of weeks ago (https://news.ycombinator.com/item?id=49661638). Feedback for that run was brutal and constructive, so thought it would be good to follow up.The Viva is live – that was the "coming soon" feature from before. Once you have completed a concept's diagnostic, you will have a live voice conversation with the AI where it will ask follow-up questions and make you defend your reasoning orally as opposed to selecting a response. Finishing the conversation”
“First of all, thank you for adding this feature to VS Code; now a proper ASR model is used instead of the poor transcription from the VS Code Speech plugin. However, as shipped in 1.131 it is designed only for English speaking users. This is a regression: in 1.130 the multilingual model was the default . From the test plan ( 326388): Nemotron (default) nemotron 3.5 asr streaming 0.6b . Multilingual (35+ languages, auto detected). [...] Verify it transcribes in that language without any language”
“<img width="771" height="86" alt="Image" src="https://github.com/user attachments/assets/c399f971 7c67 4b07 89d1 8be670a80233" /”
“Check Existing Issues [x] I have searched for all existing open AND closed issues and discussions for similar requests. I have found none that is comparable to my request. Verify Feature Scope [x] I have read through and understood the scope definition for feature requests in the Issues section. I believe my feature request meets the definition and belongs in the Issues section instead of the Discussions. Problem Description gemma4 just released with combined audio and vision support Im sure the”
“🚀 The feature, motivation and pitch will qwen 3 asr 0.6/1.7B get the support for realtime endpoint similar mistral's voxtral? Alternatives No response Additional context No response Before submitting a new issue... [x] Make sure you already searched for relevant issues, and asked the chatbot living at the bottom right corner of the documentation page, which can answer lots of frequently asked questions.”
“Your current environment NA 🐛 Describe the bug Segment level timestamps are incrementally offset by 0.5s per segment. When transcribing long audio with vLLM Whisper, segment level timestamps are increasingly inaccurate. Chunking uses by default a 1s window to find low energy (quiet) regions for splitting, so the next chunk may start up to 1s earlier than the nominal chunk length. This offset is not compensated for when generating segment timestamps, resulting in an average delay of 0.5s per seg”
