PANE

Local AI Rigs Cost More Than Most People Expected

Several posts express frustration and financial strain related to the high costs of running local AI setups. Users discuss expensive hardware, power consumption, and the difficulty of justifying the investment, highlighting a disconnect between aspiration and affordability.

aicosthardwareinvestmentlocal-llm
FIT
0%
SIGNAL
50%
SOURCES60
FRESHEST POST15H AGO
TRACKED SINCE55D AGO

SOURCES (60)

Yes! Except that I bought 1 R9700 instead of 2. Thought that I could grab a second one later. Well, they got more expensive now. Other than that, it runs smooth and you can do interesting stuff with 32GB VRAM alone.

r/LocalLLaMA15h ago

OpenAI just pushed some Jalapeno hardware that's supposed to be awesome. I think these hyperscalers will find it worth their while to invest tens of billions in a better approach to break Nvidias stranglehold

r/LocalLLaMA23h ago

The thing they SHOULD be worrying about is how good Privately run LLM's on your own (albeit high end) graphics card and computer is getting. I'm daily driving it at home now, and it produces more token value output (since it's close to or equivalent to Opus 4.6 today). And I haven't bothered to use the online models for other than casual chat for a year now. I think Nvidia kinda understands this, and their initial investment into big AI corporates (which is their main customers)

r/ChatGPT1d ago

As in actual physical heat. How are people coping with the heat that running a decent inference rig pumps out? I've got dual 5060 Ti GPUs and a relatively modest CPU (Intel 14 Core Ultra 5 245KF Desktop on an Msi MPG Z890 motherboard) and if I use it for a coding session or similar with Qwen 3.8 27B then it heats up the room terribly. I know the energy consumed has to go somewhere, but I didn't realise it would be this bad. Would throttling the GPUs' power consumption a bit make any

r/LocalLLaMA1d ago
Source preview · reddit.com

Everything

reddit.com2d ago

Just for context, I got curious after reading your comment and I also checked this text for AI. Out of seven different AI checkers I tried every single one said 100% human or almost 100% human. Pangram is far from perfect so maybe try not to throw accusations around based on limited evidence.

r/LocalLLaMA2d ago
Source preview · reddit.com

Doesn’t opengpu do this already?

reddit.com2d ago

Why would Nvidia not want open models? This is right up their alley, more people running their own local models, more chips they can sell.

r/LocalLLaMA2d ago

If you can get a GPU with at least 24GB of VRAM, then you can run many open weight models fairly well. Some fast, some really slow. Training and fine-tuning, never touched it. Why bother, when you can just feed the model the context it needs, give it tools, and it can figure it out, eventually. Less tools, the better. Power and cooling is the same as if someone is playing a graphically intense game. on-prem vs cloud? When you know how you will be using an LLM and figure out that you don't ne

r/sysadmin2d ago
Source preview · reddit.com

Supposedly solved on MacOs

reddit.com3d ago

This kind of thing is fun for toy chats. To do real work requires data that is often (almost always) private or proprietary. No way am I (or anyone else serious) giving any random node that data without a clear (cryptographic) assurance that it is not being spied on or stored. That’s is: “secure inference” … and I don think that is a solved problem on home GPUs yet.

r/CryptoCurrency3d ago
Source preview · reddit.com

Yes, that was the inspiration

reddit.com3d ago
Source preview · reddit.com

I’ll look into it, thanks.

reddit.com3d ago

If you have a 4 GPU rig, that’s 32gb and that would serve Qwen 3.8 just fine. And the models are only getting better

r/CryptoCurrency3d ago

could be, but I have no idea if this is economically or technically viable. Not all the workloads are infinitely parallelizable, and high latency can make it inviable even if you have the raw power.

r/CryptoCurrency3d ago
Source preview · reddit.com

This is already a thing. Bittensor TAO and it's Chutes subnet

reddit.com3d ago

The compute futures angle from ICE is wild never thought of gpu power like corn or oil but it makes sense

r/CryptoCurrency3d ago

Over the last few years, this topic has come up in light of Ai. Renting to Vast or Runpod isn’t necessarily easy. If you look at my comment history I have a 9y old post about mining on a gtx960 2 Gb, so no-slop. Just discussion, I don’t even have a link for you. Ai is just text, and doesn’t necessarily take a ton of bandwidth, more than blockchain for sure but it isn’t video. It is known as distributed inference. Would you guys use this? submitted by /u/desexmachina [link] [com

r/CryptoCurrency3d ago

I work with Sky Forge Compute — we rent GPU capacity, so read this with that in mind. The conclusion below points at buying more often than it points at us, which is why I think it's worth posting. Every "should we buy or rent" thread I see argues from vibes. It's arithmetic, and the answer turns on one variable almost nobody measures honestly. The formula break-even hours = purchase price ÷ hourly rental rate break-even years = break-even hours ÷ (hours per day × days per week

r/mlops3d ago

They can also win a lot of community love by just releasing an old but still beloved model like gpt4o. That's not really going to impact their revenue in any way, but I promise half the people on reddit is going to be singing SamA's praises if they ever do that lol.

r/LocalLLaMA3d ago

I am so right there with you. GPT-OSS-120B (and 20B, to a lesser extent) remains an incredible model. It’s a year old now but I still use it for a wide variety of tasks. OpenAI’s work on various kinds of classification stuff using Safeguard models makes it crazy useful. Basically give it a policy, and a document, and it will judge whether the document follows the policy. It does this better than anything else, with justification in the reasoning traces in a structured way. I’m using it for a mul

r/LocalLLaMA3d ago
Source preview · reddit.com

Or an unsloth clone

reddit.com3d ago

Yeah I've been pointing that out for a while now. I think they're hedging their bets. At the pace local is advancing, in model intelligence, harness quality, and hardware performance, the value of online frontier models could begin to fizzle out for many types of task. I'm trying to tell people we shouldn't treat this like some kind of victory though. If most/all non-Chinese AI goes away, the Chinese AI companies will stop open-sourcing their models overnight. They're not giv

r/LocalLLaMA4d ago

Isn't perplexity a bit shit?

r/LocalLLaMA4d ago

Anyone notice that nearly all the recent nVidia announcements are about local AI? They have sucked the frontier labs/CSPs dry, time to move to the next host!

r/LocalLLaMA4d ago

Looks like they have a post trained version of Qwen they're calling PPXL 27B. I think you can also use other models so value might just be in being a (possibly) better harness.

r/LocalLLaMA4d ago
Source preview · reddit.com

submitted by /u/dd32x [link] [comments]

reddit.com4d ago

It’s still just a single gaming gpu being used for AI. It’s fairly low end for AI

r/LocalLLaMA4d ago
Source preview · reddit.com

Valid, just curious. Good perspective lol. thank you

reddit.com5d ago

I'd sell some of that ram.. then use that money to buy and hook up some gpus to the server and run some local AI stuff.

r/selfhosted5d ago

--ubatch-size 2048 did the trick. Though obviously this takes some VRAM -> less expert layers in VRAM -> slower generation speeds. For my agentic coding it was worth it though.

r/LocalLLaMA5d ago

What about the energy cost of running the local setup, would be interesting to know how that compares vs. a subscription.

r/LocalLLaMA5d ago

Thank you. I think that's a brilliant use-case for local AI or for AI in general. I'm starting to think people are using the wrong settings for their AI. Idk about other software, but for Open WebUI, you have to "tune" the settings of your model first that fits your liking. Even as simple as disabling "thinking" by ollama significantly makes their responses faster.

r/LocalLLaMA7d ago

https://preview.redd.it/66n9s9oaickh1.png?width=403&format=png&auto=webp&s=94117c07ca1e218326313c0d289364c3c36741e2

r/LocalLLaMA10d ago

>dedicated chips will be created for these architectures, and you'll just load the weights. This has already proven to work, with labs running Gemma or whatever it was at 14000 tps. With 2k context window, lol. Thats nowhere near usefull. Only think you need is tensor cores and a lot of memory bandwidth\interconnect, and software support of course. And boom, you got npu that can run almost any llm. Thats the reason why nvidia gpu price are skyrocket, and gpu like 3090 is still really well

r/LocalLLaMA10d ago

I've been wanting to get into small scale local ai for a while, but I'm on a pretty tight budget. Don't mind thinkering a bit to get things working Are these old datacenter cards still worth buying? submitted by /u/Far-Classic-9963 [link] [comments]

r/LocalLLaMA10d ago

Yes, that's my point, to have a pc that runs a good AI and runs it well you will need to spend an enormous amount of money, which I don't think any of us has.

r/ObsidianMD10d ago

Yup. It's cheap. So use API credits. And what you run in the cloud will be faster and higher quality (full precision) compared to what most people run at home. Can you run 27b with significant context on a 3090? Seems a bit too big at the normal quants, but I might be wrong. I wonder if 2 5060 TI 16gb would work out better than 1 3090. I imagine some folks out there have done benchmarks. I know the 3090 has higher memory bandwidth and there are issues with 2 GPUs , but 5x series has more per

r/LocalLLaMA10d ago

There's going to be a tipping point when intelligence becomes too cheap to serve and GPU's utilization starts to decrease.

r/LocalLLaMA10d ago

very cool but, model closed, dataset not really open, it's the input they used to make the data, this would require a lot of time and llm input to test 15 attempts * 6000 samples, compile and benchmark each, then 3 billion tokens of RL training

r/LocalLLaMA11d ago

The physical cap on frontier model companies will make nah force AI local. For instance OpenAI, anthropomorphicologists take too much resources. Cities and construction, natural resources- power grids can only output so much. Billions invested whatever, the $ can’t make nature grow for them. /rant

r/LocalLLaMA13d ago

There’s actually a whole field of study for this, look up LLM / model densing laws. Also a recent talk at AIE Worlds Fair called “The Desktop Frontier”. Lots of great data. In my research, late 2027 is a safer bet for Mythos level intelligence on a 5090 GPU. Even if it’s 12+ months off it’s insane that this will be possible at all. What is a 5090 worth in a world like that.

r/LocalLLaMA13d ago
Source preview · reddit.com

Underpowered hardware due to high costs.

reddit.com13d ago

There are few gamers who want to mess with AI. Also, there are few programmers in the world, mathematically speaking. Another funny coincidence. And, of those only a set of them want to mess with AI. And, many of those who do are on the frontier models. Nerd is a specific subset.

r/LocalLLaMA13d ago

Supposedly next gen top of the line Mac will have 2tb of memory, but tensor processors may not even need all that memory and there is some company embedding models right in silicon on old processes and getting 1000s of token throughput. The models are fixed but it could be possibly be built like old game cartridges where you can switch them out when better ones come along. So many cool possibilities coming.

r/selfhosted13d ago

I'm way ahead of you. I got a Mac studio m4 and plan on pile up another with exo. Not selling and buying.

r/selfhosted13d ago

make it clear to model makers that there's demand for it, not drown out gpu-poor posts with super high end tier hardware stuff, and perhaps some adjustments to the subreddit (like flairs or sidebar/wiki information) to define a clearer definition of the different tiers, so that people can flag a post with a flair referring to such a tier, making it easier to find the stuff that applies to them

r/LocalLLaMA14d ago

What do you mean by how viable it is? From a compute resource standpoint? Quality? Computer resource it runs now, to do it at scale I need to put the llms onto a decent gpu, with that I can support around 93000 users on 1 host and remove all my external llm costs. Quality wise I’ve put a lot of work into this and removed some sources which poison the well and built a system to detect it. Having better news sources will always be the big thing to make it better though.

r/Entrepreneur14d ago

It honestly feels like it drains at least as fast as terra, sometimes it feels like it drains as fast as sol/opus, considering it's significantly cheaper than those two, it never makes any sense to me. Honestly terra/luna-maxfeel like the most balanced options now.

r/Notion15d ago

Usually the argument is that they aren’t thinking, but reproducing things they’ve seen. It’s like training a chimp to do a human task. If you train it well and the chimp does the task, is the chimp intelligent or just copying the human to get reward points? Some would say the chimp shows intelligence, so the conversation really stops there. But for people who don’t agree the chimp shows intelligence, it’s hard to justify how AI is “intelligent” and not just a system working to get the biggest re

r/selfhosted15d ago

Local AI will sit with Apple. Their proprietary hardware acceleration engine and Unified RAM chips can run the same models an NVIDIA GPU at a fraction of the price. Might not be as fast, but it can still run them. If I were to consult SMB and Medium-sized owners, I would tell them to go Mac if on a budget, NVIDIA if not.

r/Entrepreneur15d ago

Hello! I have recently built my AI rig (3x RTX 5060 Ti 16gb, with possibly a 4th on the way if I can fit it). I love it, it runs great, and I am getting between 70t/s - 110 t/s (according to the pi agent web UI, have not confirmed it yet). While it is fast, I struggle to put it to use in the way I was hoping. My dream has been to be able to put it to work writing code autonomously so that I can have it sketch out my ideas before I commit to developing them, however, every attempt I make just see

r/LocalLLaMA16d ago

The companies / countries pushing out free open weights / open source AI are the ones playing catch up. Period. It isn't cheap to create these models and if you are giving them away, you need a longer term strategy on how this will make you money. Regarding open weights models, yes I think having those models available will create a need for more inexpensive 'consumer grade' AI HW. But I doubt that the companies making the consumer HW are the exact same companies releasing the open w

r/LocalLLaMA17d ago

The future of AI is models on a chip. A processor that IS the model will be created each year or other iteration. This will be able to run the model at 1000x current speeds with less RAM usage and much less power usage. Frontier models will likely still own their own data centers. But open weight or other models may produce chips for desktop use.

r/LocalLLaMA17d ago

For years the honest answer to "why self host the model" was that you don't, not really. You run something small, it's worse at everything, and you keep paying for the hosted one for the work that actually matters. Hobby answer. I've given that answer myself plenty of times. But. Someone published a decode run this month off one 128GB desktop box, full size open weights model, 40.2 tok/s single stream. One measurement, no independent second run that I've found. I don&#3

r/selfhosted17d ago

Fair enough, and honestly I'm deliberately trying not to price it on my time. If I did it'd be 80 hours times whatever rate, call it 3k, and that feels like the wrong frame for something a software house would quote 15 to 40k to build. But I'm also not really asking anyone to put a value on my effort. The stuff I'm stuck on is structural, and I don't think it's subjective: retainer vs licence plus yearly maintenance, how much support to include before it becomes billable,

r/webdev18d ago
Source preview · reddit.com

Time to man the generator bicycles...

reddit.com18d ago

I'm looking to build a serious local AI workstation, with an absolute maximum budget of £3,500. My benchmark is ChatGPT Plus. I use it heavily for professional work: deep research, analysing PDFs/images, producing reports, PowerPoints, strategy documents and generally turning rough briefs into polished deliverables. I'm happy for local inference to be significantly slower. Quality matters far more than tokens/sec. What I'm ultimately trying to build is a self-hosted AI work assistant

r/LocalLLaMA18d ago
Source preview · reddit.com

True but trying harder to prevent that

reddit.com19d ago

laptops as servers always bite me on the fan. if that thing's running 24/7 with docker containers i'd keep an eye on the hinge temp, had one melt the display cable after a few months

r/selfhosted19d ago

SOLUTION LANDSCAPE

Brought to you byTop Sectors

A Player feature.See how many ways this pain can be solved, who's already building, and where the gaps are.