AI Devs Priced Out of GPU Hardware
AI developers are facing challenges acquiring and utilizing adequate hardware for local LLM development and training. High costs, limited availability, and concerns about future-proofing are driving a search for optimal solutions, often involving expensive Apple products or alternative hardware options. The shrinking availability of high-memory configurations from Apple is exacerbating the problem.
SOURCES (60)
“Hello all Just following up on a post I made last week about my experiment to try minmax my Mac Studio. In particular, I've had quite a lot of success with pushing things even further. Across a 20 minute test with three concurrent sessions, my…”
“Is it worth running local models on this old beast? Dell PowerEdge R710 (2009-12 era) Dual Xenon 5500 (I'm pretty sure) 48Gb DDR3-1066 submitted by /u/Motor-Independent572 [link] [comments]”
“What is the best local model for 48 gb mac m5 pro? I will be using Lmstudio/Bionic. submitted by /u/Jedirite [link] [comments]”
“I’ve been seeing a lot of people talking about small models (~7B), and I discovered them when I started running models locally on my MacBook Air M1 (8GB). I’ve realized that 99% of the time, we don't actually need massive 1.3T models like Kimi k3, nor do we need to pay a ton of dollars for a Claude subscription. We just need to learn which models excel at specific tasks and run them locally. For example, the Gemma 3 1B has been surprisingly good at writing. Also, the Mistral 3B is excellent”
“Having played with localllm for a while now I can confidently tell you there are better ways to spend your money. The Mac Studio will not match the speed of a good nvidia pro card, especially if multiple people use it. And even then, it won’t be able to match the inference quality of Claude’s Haiku. You’ll need four of them connected to get enough memory to run a reasonable model. But even if you could get suitable performance on one, say you’re running a good 31b model that’s GPT 4 class, you’r”
“if they release the weight, it would be 4bit native 1.4T, so a single or couple of Mac Studio in 2027 should be able to run it”
“Could probably get away with using the Q8 for practical use. Better than most alternatives at that price point”
most likely not anymore coz' of the price hike and memory reduction
“My pc is limiting me a lot ! I wanna upgrade but I feel like I won’t be satisfied, my goal is smooth playback of course export time matters but smooth playback while working with long 4K footage doing masking and effects. Does anyone has or had a max m2 with 64 g of rams how does it perform ? On YouTube most reviews check render time ! I’m ok with rendering while sleeping lol but I want smooth editing Thank you in advance submitted by /u/karym7 [link] [comments]”
“It would fit in my old (circa 2019) laptop with an 8GB 2070. Going to give that a try.”
Maybe a nice large monitor and layout multiple remote desktop sessions.
“they've increased the price of high-RAM machines by over 50% That's only because of RAM cost. Apple is activity looking to address that by trying to source RAM from China. It follows that a new 512GB Studio would be $15K As I said above, that's an exigent circumstance. It's like when companies charge fuel or tariff surcharges. Look at history. What Apple does is charge about the same for the new model as they did for the old model. Even if it has more capability. Like more RAM. T”
“And you're totally right that the ebay cost is irrelevant. What matters for comparison is what they would sell the machines for today. We know they sold them new for $10K, and that since then, they've increased the price of high-RAM machines by over 50%. It follows that a new 512GB Studio would be $15K, which puts a 1.5TB Studio easily at $45K, especially since it would also have a more advanced processor.”
“I have no idea why anyone would want to spend that kind of money on a unified memory system when they could spend the same or less on an absolutely insane GPU setup that completely blows it out of the water for speed. Less RAM, sure, but still plenty of it to run killer models.”
“I do run full fat GLM-5.2 locally! (At the blazing speed of 7 tok/s gen and 25 tok/s prefill)”
“With high bandwidth flash that may be within reach for consumers if the powers that be deign to grace us with such a blessing”
“But for many use cases we don't need SOTA. That's why local is important, so we can now and forever run these models.”
“Nah, they see an opportunity here. If they can establish themselves at the server for small/medium shops for LLMs, they can take that to the bank.”
“Can you imagine running full DeepSeel or GLM locally. That will be endgame. By the time that happens, the models will be unimaginably more performant too.”
They’ll probably acquire one of the other labs after the bust.
“If they deliver on 60-80 cores and ~160 gpu cores, I think it would much better than a current day Turin.”
“That's a great price for a system with 1.5TB of RAM. Last quote I got for a system in that memory range was well over $100k per node.”
“I’m looking for some opinions from people who have already gone down this road. Right now I run everything on my Mac mini M2 Pro (16GB RAM), which is also my primary workstation for design and development. Current stack: jellyfin (native) docker (orbstack) sonarr, radarr, lidarr, prowlarr, qbittorrent (through gluetun), nzbget, slskd, jellystat + postgresql, file browser, cloudflare tunnel, unpackerr, jellyseerr Media lives on an external 20TB USB drive. Planned additions: immich paperless and .”
“hey, haven't been up to date for longer about what currently are the best models. so whats the best recommended model for this? can be either vlm or llm. i saw qwen 27b in another thread? and any way to look in a ranking reliably from time to time for best models or is it really more based on what people try out and experience here? submitted by /u/Emotional_Thanks_22 [link] [comments]”
“tl:dr; RTX5090 for 3400€ or Bosgame M5 AI for 2500€? I have a fairly new computer with a 5080 and 64Gb of RAM. I've been having loads of fun with local LLMs. In the end I find myself using thu free Claude and Deepseek V4 Pro through their API because it's so fucking cheap and way better than anything I can run on my machine. For some applications though, I have to use local models: sometimes for uncensored models, sometimes for privacy, sometimes I just want to have full control of the s”
“Hey everyone, I just setup my new HP Prodesk 600 G6 i5-10500T 8gb (soon to be 16) RAM, 256gb storage, with Ubuntu/Docker/Komodo/Homepage, as well as many private instances of things I use daily (I use my NAS as a DVR server for my tv service). What are somethings I can add to my new setup? I'm new to this but want to maximize and explore this new world I've unlocked. submitted by /u/Upset-Award-1685 [link] [comments]”
“My use case: - LLM for coding (chat and agents, be prepared for future hybrid inference). - Code compilation Many CPU recommendations state that a Ryzen X3D is not worth the premium (for LLMs), it could be even slower due to (slightly) lower boost frequencies. Thus, a Ryzen 9950X would be the best pick on a normal consumer mainboard, targeting one or two Radeon AI Pro R9700. However, looking around in this sub, I see far more 9950X3D CPUs than 9950X. Why? Did the CPU recommendations change recen”
“(Repost as I removed the Video Intro, Im guessing people dont like it. Added some animated Benchmarks too) Hi yall, I benchmarked my 128GB M5 Max Macbook Pro and here are the results. I also made it into a Video if you want it a deeper dive with more of my methods and takes, but Im sharing all my findings below regardless of whether you watch it. I am new to Local AI, so do let me Know if I did anything Wrong, or if theres any room for improvement! I would love all feedback, Im just excited to t”
“Hi yall, I benchmarked my 128GB M5 Max Macbook Pro and here are the results. I also made it into a Video if you want it a deeper dive with more of my methods and takes, but Im sharing all my findings below regardless of whether you watch it. I am new to Local AI, so do let me Know if I did anything Wrong, or if theres any room for improvement! I would love all feedback, Im just excited to try these out - I did buy this Mac for other purposes and just wanted to mess with Local AI for fun, but thi”
“it's usually active b parameters = GB VRAM. But if you trim it down let's say use quantization Q4 you might be able to run larger models e. g. 20b on a 16 GB VRAM GPU / unified memory (Mac). Also context runs in VRAM as well so let's I can run Gemma 4 12b model on my 16GB VRAM and still have around 32k context to fit in my VRAM. Which is not a lot btw, you easily need 1 mil tokens context window if you are serious about building a business. On top of everything there is the harness a”
“I run a machine with the i5-10500T for couple year now. No issues its my main container and vm host with 64gb of ram”
“I am running a 12 core with 16gb and a 512gb drive and it runs a lot with out any problems but I could certainly use a bigger drive and 32gb or ram would be nice but the 16 works well.”
“If I have a 128GB M4 mac, and want to run the Qwen 27b model in like 8-bit or 16-bit, ideally in LM Studio, although I guess I could try to learn how to set up llama.cpp (I'm a noob), am I better off using the unsloth MTP version of the 8-bit or BF-16 GGUFs, or better off without the MTP just using MLX 8-bit or MLX 16 bit? So far I've never used the MLX method before and have only ever run GGUFs in LMStudio so far, so I'm not really sure how all that works or which is better for diff”
“I am using dell xps 16 which is ultra cool and powerful but of course it also in the expensive side”
“I would recommend buying refurbished equipment, as even i5 or i7 from 5 years ago offers a significant cost savings over buying new. The equipment is very capable especially for business users.”
“Ok but you build the management workflow (at least the basics), and then you start deploying macbooks. I did that as well when Macbooks started becoming popular, I dogfooded a Macbook M1, and then nobody wanted them anyway. So it was a bit of a waste. I loved the hardware though. If only my personal M3 had a way to run Linux it would be the best laptop I could hope for.”
“Off loading to system ram really slows things down. But thats is a really good deal”
“Maybe, when Apple finally releases the new Mac Studio, they will offer a 2TB RAM option ...”
“We anticipated the issues and bought as much as we could last year. I got a quote for a laptop just for curiosity, and it was 2600 now for 1600 laptop.”
“Your college's business school will have the pc requirements listed on their website. Check with that before you buy anything.”
“Alternatively get something with more form factor and get a cheap usb 10 key for when you need it?”
“You can get a Vert refurbished Lenovo Thinkpad Tseries on Amazon w warranty still for $300-400 Ssd, 16 GB ram (upgradable etc) They are workhorses, easy to repair, widely available parts etc”
“My RTX 3070 Linux gaming laptop stays around ~75-80°C under sustained interference load”
“No idea why this got downvoted, you asked a sincere question and got “more is better duh””
“Ok I now started testing different LLM on my MacStudio and I am just curious. I use Menubar to see the teperatures of the components - and yeah the CPU and Grafic are all quite hot. Average CPU cores are at 62 celsius and all Graphic cores are at over 80 celsius. Ok my mac is doing this for just 2 days now, and I could go "Yeah a mac will survive that for sure". But still what do you guys experience with your Studios or Minis for a longer period of time? Don like having to get my mac r”
“Waiting for Apple to release things resulted in me waiting while all the 512 and 256 and even 128 studios (that I could have afforded at full sticker) disappeared only to triple in price from scalpers. I guess I’m “lucky” to have grabbed a 128gb M5 Max MacBook just before the price increase - and on microcenter sale to keep me busy while I also wait to see the astronomical pricing on the next gen of studios. Personally I’m guessing they top out at 256gb and push $20k usd price wise.”
“I kept answering the same question for friends ("I've got a 16GB MacBook / a 3060, what can I actually run?") and got tired of guessing, so I started a spreadsheet. It grew into a real dataset, so I put it on GitHub under CC BY for anyone to use or fix. Rule of thumb I landed on: at Q4_K_M a model needs roughly 0.6GB of memory per billion params, and you want to size to about 70% of your RAM/VRAM so the OS, context and KV cache still have room. From that, the comfortable ceiling pe”
