PANE

Creators Want Cheaper, Simpler Text-to-Speech

Independent developers and users are struggling with expensive, complex, or privacy-invasive text-to-speech (TTS) tools. They desire simpler, more affordable, and often locally-run alternatives that respect data privacy and offer a more natural dictation experience, leading many to build their own solutions. The current market lacks accessible and user-friendly options.

creator-economyproductivityaiaudiodevtools
FIT
0%
SIGNAL
64%
SOURCES60
FRESHEST POST7H AGO
TRACKED SINCE107D AGO

SOURCES (60)

Audacity has an implementation of Whisper that you can install locally. I’ve been trying it out, and it seems to work pretty well, although larger models will take too much time to be useful for me.

r/podcasting7h ago

I'm 100% with you. Been using the new Notion AI app to dictate stuff into my database, but that's a little slow. I'd be interested in testing your app.

r/Notion7h ago
Source preview · reddit.com

Cool I am doing something similar. It is a live conversation over a set of premade questions though instead of just talking to it. Is yours conversational AI or just transcription? Would…

reddit.com7h ago

The 'calls you out' idea is useful, but I would make the tone extremely adjustable. A lot of people who fail at journaling already associate it with guilt, so if the app sounds disappointed it could backfire.\n\nA better framing might be: 'notice the pattern and restart gently' rather than accountability coach. For example: 'You skipped yesterday. Want to do a 60-second catch-up instead of a full debrief?' That keeps the anti-ghosting loop without making the user feel pun

r/SideProject7h ago

FWIW I do recall a post from a while back, where people did configure a conversation with Siri that would ask a question for certain properties in dictated page. The juice to me just didn't feel worth the squeeze.

r/Notion7h ago

local whisper is basically the exact answer to your whole list. since the audio never leaves your machine there's no training/privacy question at all, no length cap, and no subscription. on mac macwhisper's free tier is the easy path like the other commenter said. on pc grab Buzz or faster-whisper (subtitle edit has a whisper gui too if you don't want to touch a terminal). one heads up: it runs on your own cpu/gpu so a 2 hour ep isn't instant, and model size is a tradeoff, the &q

r/podcasting9h ago

Thanks, this is exactly the kind of feedback I’m looking for. The first version is focused on making text and voice capture into a Notion inbox as fast as possible, so iOS widgets and Siri integration aren’t included yet. But I completely understand why those would make a big difference compared to opening another app. Being able to dictate something and capture properties in the same action is especially interesting. Also, thanks for letting me know about my PMs — I didn’t realize they were tur

r/Notion10h ago
Source preview · reddit.com

The thing you have actually put your finger on is that the structure of a thought is relationships, and relationships are the first thing a flat transcript drops. So the fix that…

reddit.com12h ago

The map only reorganises when you pause and the raw transcript is all there underneath, so worst case you dump everything by voice, ignore the canvas, and look at the structure after. You’re making me think a capture mode that hides the canvas until you’re done might be worth building though.

r/ObsidianMD15h ago

My role of thumb is first focus on the information, later on the presentation. I can always change how it is displayed, but I can't always remember what I was thinking... Once the information is there, I organize it. Usually in a textual format, someone's using excalidraw or mermaid diagrams.

r/ObsidianMD15h ago

The transcription in ChatGPT probably only picks up and prints out recognizable word tokens, not sounds, but the underlying voice model might still be able to recognize and respond to sounds as well. It could be that your breathing noises just came through loud on the mic and that was all it had to respond to. It's probably fine, but maybe try out one of those apps that listens to and records your breathing at night anyway. Sleep apnea is super common. I've got it, and can confirm it&#39

r/ChatGPT1d ago
Source preview · reddit.com

Yeah thats the whole point

reddit.com1d ago

When I first started playing with Descript one of the first things I did was to see how well it would emulate my voice. What surprised me was its insistence on "poshifying" my voice. It consistently softened my strong Australian accent to give it a distinctly British middle class tone which just isn't me.

r/podcasting1d ago

Ouch! I really like Wispr Flow. It's AI enabled, so it's more accurate than other transcription software. I have fibromyalgia and use it when I'm too achy to type. They do charge if you go over a certain amount of words per week, but it's well worth it IMO. Hopefully you can still use a mouse.

r/Accounting2d ago

The game is a short, story-driven experience with Papers, Please-style gameplay. We originally planned to voice the characters, but we don’t have the voice talent for it and the realistic approach wouldn’t fit the style we’re going for. submitted by /u/VECTOR3Studio [link] [comments]

r/IndieDev2d ago
Source preview · reddit.com

Handy with Cohere Transcribe model.

reddit.com2d ago
Source preview · reddit.com

I text to speech email myself and then copy paste it.

reddit.com2d ago

It’s live on the App Store (iOS only for now), and it’s early — a handful of real recordings so far, not a ghost town but not crowded either. What I’d love feedback on specifically: **•** First 60 seconds in the app — is the “record → pin to location” flow obvious, or does it need explaining? **•** Does the privacy-pin-randomization idea actually feel reassuring, or does it just feel like friction? **•** Would *you* record something on your own life, or does it feel like a stranger’s app you’re

r/alphaandbetausers2d ago

Handy is an opensource speech to text app which runs locally. Tried it for a bit, and it seems to work pretty well.

r/ObsidianMD3d ago
Source preview · community.home-assistant.io

Starting to use this locally with the PE and also Echo, so mostly in exploration mode but want to shift entirely soon. One thing I’d like to do (and frankly should do…

community.home-assistant.io3d ago

It seems that the more crucial aspect here is not whether a tool can generate a transcript, but whether you can retrieve the exact client’s words a month later without reconstructing the entire story from scratch. While bot-free capture is beneficial, the true value lies in a searchable record that includes speaker attribution and a straightforward method to access the original source when verification is required. I have specifically developed Loreo for this purpose. Please let me know if you n

r/legaltech3d ago

The only actually good speech2text functionality I've ever seen is ChatGPT's, but without a subscription it doesn't become a reliable regular solution. Have any of you guys found a good method? submitted by /u/SuppaDumDum [link] [comments]

r/ObsidianMD3d ago

For some reason the original post wouldn't let me add a screenshot, but it works in the replies https://preview.redd.it/2xitlbkv0odh1.png?width=2584&format=png&auto=webp&s=2b3731e5b70c4564e063c03dc5eedf0c79e3e160

r/SideProject4d ago

Also just as an information piece for you, the ai said it can now take the audio and rebuild it and add chorus anywhere it wants sooo...Idk...You look more knowledgeable about the program it could probably be using. It mentioned Python soo maybe its using that library you mentioned.

r/ChatGPT4d ago

Idk, it literally rebuilt the audio and changed it soo, and I know from a few weeks ago it misunderstood what I said and created a song audio file. And no I disagree, its a pretty good method since it fixed loudness, it realized the s sound was being pushed to the left and side so its emphasized in the song. I never would have found that and using audacity etc, always has these issues never resolved. Even the D-Esser in Bandlab not working despite using it right. I would not have been able to up

r/ChatGPT4d ago

When feeding voice-to-text transcripts into an LLM for editing, the model often tries to "correct" NATO phonetic spelling or letter-by-letter spelling to match a misspelled proper noun in the transcript (e.g., changing the spelling of a spelled-out name to match what the transcriber guessed). To solve this, I added a "Phonetic Priority" rule to my system prompt framework. It forces the LLM to treat phonetic spelling as the absolute cryptographic anchor of truth. ### Phonetic

r/PromptEngineering5d ago

https://youtu.be/oxpGq5FITgA?si=nkHWLReGCDYe7QfL I got Hermes running in the native Debian Terminal in Graphene OS and its really slick. Voice dictation works amazingly. Im using a remote Hermes gateway running on my laptop as the backend, with Llama.cpp and Qwen 3.6 35b. Paired my mobile 5070ti with an RTX 3090 eGPU on TB5 for about 34GB of VRAM. 2500 tps prompt processing 80-150 tps generation w/mtp Incredible setup imo. Light and fast. 100% local. Just wanted to show off whats possible and ho

r/LocalLLaMA5d ago
Source preview · reddit.com

OH MY GOD... THANKS!!!

reddit.com5d ago

It is still possible to get it, but you have to use a custom GPT, e.g. like this one. You can remove all the extraneous commands from its definition, but keep the three under "Follow-up commands".

r/ChatGPT5d ago
Source preview · reddit.com

Appreciate the recommendation!

reddit.com5d ago

audio.cpp again. Hopefully you are not sick of it yet :) Release 0.3 adds five new models: Supertonic 3, MOSS-TTS-Local, MOSS-TTS-Nano, IndexTTS2, and Irodori-TTS. The highlight is Supertonic 3. It can hit 200 ×+ real time on CUDA (RTX 5090), 6×+ on CPU, and around 47 ms TTFT in CUDA streaming mode. In the demo (sorry for the rough demo), I used The Adventures of Sherlock Holmes as the input and generated around 10 hours of audio in about 3 minutes on an RTX 5090. Supertonic 3 was also pretty fu

r/LocalLLaMA6d ago

Hello I have recently started a channel , so far I got 100 subscribers that too organic. My biggest challenge is Voice Over, I am doing VO from google TTS, speaker I am selecting is Iapetus, temperature 0.74 However the issue is if I hear my VO in phone(Without head phones) 1 ft away(and with or without surrounding noise), the sound seems too low. I have tried increasing amplifier in audacity but no luck, tried increasing volume in capcut ,Edit the audio in auhphonic but still I feels the sound

r/NewTubers6d ago
Source preview · reddit.com

thanks you’re right

reddit.com6d ago
Source preview · reddit.com

thanks!

reddit.com6d ago

Just one transcriptions. I have a script running on the note generation LM that strips out the jargon and noises and only fetches relevant medical information one tge whisper model created the transcriptions. So far it captures about 80-90% accuracy. Hence why once done, the doctor validates it before save. And to keep it compliance to the whole legal stuff and not storing patient data, i delete the recording once note is generated and once doctor exports or start next recording the previous not

r/LocalLLaMA7d ago

Hi, thank you for actually stress-testing this, seriously. Both bugs you hit are fixed and shipped: Chinese: worldtyping.com/chinese-typing — try "nihao" + space again Voice: same page, try the mic Also added a bug/feedback form at the bottom of every tool page while I was in there, so future stuff goes straight to me instead of getting buried in Reddit. If you find anything else broken, please break it more. Cheers 🙏

r/SideProject8d ago
Source preview · reddit.com

AI transcription

reddit.com8d ago

Non-native speaker here. I could always read and write English, but when I spoke, people kept asking me to repeat myself. No app could tell me WHICH sound was wrong, so I built one. How it works: you say a word or sentence, and it breaks your speech into individual sounds and scores each one (green/yellow/red) in under a second. If you say "t" instead of "th", it literally tells you "sounded like t" and shows you how to fix it (tongue between your teeth, soft air).

r/indiehackers9d ago
Source preview · reddit.com

Currently using Arbor

reddit.com10d ago

honestly i was drowning in back to back interviews until a founder buddy put me onto lamponi. its just a little japanese app that records everything locally on my mac so no weid bot joins the call, and i can ask it questions mid interview instead of srolling through hours of text later. way less noisy than the big name tools imo.

r/ObsidianMD10d ago

Hey guys, I’m not sure how to post this here, but I’m looking for some advice on something I’ve been building. I’ve been trying to improve my English lately and needed a solid transcription tool to get feedback on my speech, basically to check if what I’m saying actually makes sense or is easy to understand. I looked around online, but everything I found were web subscription services that not only charge a monthly fee but also harvest your data. The closest thing to self-hosted was Whisper, but

r/selfhosted10d ago

Awwww bless your lil ole heart (joking people don’t admonish me)… hello fellow southern US located human. I have a trick for twangs and southern draws. When using text to speech, especially with words above the fifth grade level, I use a technique I coined “ventrilotalk”. Since I voice transcribe over ninety percent of my notes and our facility uses a mangled proprietary in house app that uses whatever Apple uses for text to speech I found that I was constantly having to manually back edit shit

r/nursing11d ago

Too bad this isn't Windows and/or Android - would love to test and provide feedback.

r/alphaandbetausers11d ago

The voices were updated. They don’t sound bad at all on live mode. The advanced mode ones sounded like they were holding back a really bad ass rip sometimes, I agree, but these are like way way better.

r/ChatGPT12d ago
Source preview · reddit.com

Better not timeout after five seconds. Biggest gripe

reddit.com12d ago

i use tactiq.io . chrome extension, nothing joins the call, transcript and summary are there when it ends. searchable history works well for going back to find what was said in a specific meeting

r/legaltech12d ago

Smart idea, and you're right to hedge, today it's a tiny slice. Almost no blogs ship their own spoken audio, so a "detect and prefer the human version" feature would sit idle 99% of the time right now. But the logic is exactly right as a fallback hierarchy: use the blog's own human narration if it exists → use a publisher-provided feed if there is one → generate TTS only when there's nothing else. That way quality degrades gracefully instead of defaulting to synthetic w

r/SideProject12d ago
Source preview · reddit.com

which one is best for tagalog?

reddit.com12d ago

they did that because they didnt see any way to make money from it. always fuckin money bro 🙃😭🥀

r/webdev13d ago

u/op - I like this; can you change the voices though? I'd like to be able to pick. Also what AI are you using? Is this all free tools? Or Deepseek + Qwen (TTS). Curious to understand the tech stack used to create this.

r/SideProject13d ago

Really appreciate this, the trust point especially. That's exactly why I built it fully on-device, nothing typed ever leaves the phone, and I'm trying to make that obvious the moment someone enables it rather than burying it. You're right that the permission wording carries as much weight as the feature itself for a keyboard. Good call on the weird contexts too, I've been testing in Messages and Notes but hadn't stress-tested forms and password fields properly yet. Adding tha

r/alphaandbetausers13d ago

Haha! It doesn’t matter what it sounds like to other people, it’s about being able to tweak it and synthesize it so it sounds like your inside head voice to you… then you can save the settings and use it like a “filter” in photos… each time you record your voice from then on it sounds just like YOU.

r/SideProject14d ago

SH|RA runs almost entirely on your machine, no cloud, no subscriptions to get started. It handles voice commands for apps, media, timers, notes, Discord, and more. It also builds a local memory of you over time that persists between sessions. I'm always looking to improve the app so feedback here or in the video comments is both welcome and appreciated. I am currently doing a beta phase and am looking for more testers. Here's a showcase of some of its features. https://youtu.be/4Ua8sU-9n

r/SideProject14d ago

Bonjour à tous, Je développe actuellement Mon VocalChef, une application web conçue pour cuisiner les mains libres. Le principe est simple : on dicte ses ingrédients à la voix pour obtenir des recettes et on suit les étapes de préparation pas à pas, uniquement grâce aux commandes vocales. Sur Chrome et Android, tout fonctionne parfaitement. En revanche, je rencontre un vrai mur technique avec iOS et Safari mobile. En raison des restrictions de sécurité très strictes d'Apple concernant l'

r/alphaandbetausers14d ago

I wish OpenAI would give it the ability to act on requests like “Please wait 3 seconds after I finish speaking before replying” A kind of hack that I use is, I tell it to only give me one-word responses until I tell it otherwise. That way its interruptions are at least brief. Or at other times instead of using Advanced Voice Mode I’ll use the feature to just transcribe my speaking.

r/ChatGPT14d ago

I've run an earlier less feature rich version on a 16GB GPU, using Qwen3.5 122B as the LLM. See the v1 branch on my GitHub. But I just don't think 12 will cut it, sorry.

r/LocalLLaMA15d ago

Qwen3.6 35B in my testing just didnt have the world knowledge to speak as intelligently. I have run this with Qwen3.5 122B and it was basically realtime, but significantly less erudite than 397B. Anything below 122B will win you realtime conversations, but the quality of responses and contextual connection finding won't be as good.

r/LocalLLaMA15d ago

Interesting! Thoughts on running on a Mac 64 gb Mac? I wonder if Qwen3.5-397B-A17B could be replaced with Qwen 35b MoE…

r/LocalLLaMA15d ago

I've posted up earlier versions of this project before, promising a GitHub link, but never got around to pushing the code from my local system. Sorry all, I have a very busy life :P Anyway, without further ado: GitHub: https://github.com/igorbarshteyn/athena Athena is a fully offline, privacy-first voice assistant that runs entirely on local hardware. Athena combines a large mixture-of-experts language model (Qwen3.5-397B), neural text-to-speech (Orpheus 3B), real-time speech recognition (Wh

r/LocalLLaMA15d ago

Ah that makes sense, the partial transcript appearing as you talk is a nice trick. And good call on battery being screen/mic not the connection.

r/selfhosted15d ago

SOLUTION LANDSCAPE

Brought to you byTop Sectors

A Player feature.See how many ways this pain can be solved, who's already building, and where the gaps are.