AI Devs Priced Out of GPU Hardware
AI developers are facing challenges acquiring and utilizing adequate hardware for local LLM development and training. High costs, limited availability, and concerns about future-proofing are driving a search for optimal solutions, often involving expensive Apple products or alternative hardware options. The shrinking availability of high-memory configurations from Apple is exacerbating the problem.
SOURCES (60)
“Spark coz even though the decode is okay perf nothing in the same price bracket can do compute bound stuff for prefill, image & video gen, other model archs. If that's why you're getting these accelerators”
“I was running a setup that's basically similar to the Strix Halo with additional gpu attached to help speed it up. Even with the memory and extra vram, moving to dedicated 2x 7900xtx is so much better. I can load bigger models on my strix machine. Yes, and the one gpu i had attached helped. But I bought it for two additional egpu ports that I can't use, thus replacing with a desktop and really... having experienced both sides, unless you are going to buy two sparks+ to be able to load to”
“I’ve been running qwen on my personal Mac but I’m getting to the point where’d I’d like to have something always on, running various jobs, and some more ability to experiment and earn about fine tuning. I’d like to keep things <$5k if possible. To anyone with any of these 3 platforms, what’s your experience been like? I’m drawn towards the DGX spark for concurrency and CUDA (which I have very little experience with) but I’m a little turned off by it’s memory bandwidth. I have the most experie”
“I have a need for a 'research assistant'. I'm looking for advice on model choice and set-up, as well as level of hardware needed. There are two scenarios actually but they have overlap. Scenario 1: Plough through large amounts of semi-structured natural language to both search for specific types of information or specific topics, and extract that info. The amount of data is way over any feasible context, but it breaks down easily enough (usually) so some sort of looping set-up that I”
“One thing these prices have totally fucked beyond all recognition is my sense of what something is worth. I have a 5090FE i got from best buy. And like everyone else salivating over beefy mac studios and more of these beefy GPUs. But i just realized probably a vastly larger impact on my life can be made with the updated apple vision pro. Somehow $3700 isn't outrageously expensive any more.”
“I recently moved to a 14" MacBook Pro M5 Pro for personal use. I mainly got it for the efficiency/power of the Apple silicon, to experiment with local AI, and because my AMD Surface Laptop doesn't output 120Hz on my 5k2k monitor. I previously shunned macOS, but after using it, it's not bad, but it's also not mind blowing like some others have told me. In fact, there are some features/settings on Windows that aren't the default or available out the box, such as volume control”
“Look on eBay and see what they sell for. I have everyone here have a ready spare at their location.”
“I upgraded the DDR4 on my gaming laptop with RTX 2060 6GB a long time before the RAM price increase. And even today, I'm still struggled to find a use case for this 64GB of RAM that my 32+16 desktop cannot do better. MoE is the answer, but it does not solve everything. KV cache and attention layers still need to sit in GPU, and 6GB could be quite tight. Not to mention the situation of the models. I got the old 80B-A3B that I wished to run but did not have the RAM back then running on the lap”
“So this is a great idea, but even cheap laptop memory is pricey these days. It doesn't feel like it's cost effective against my time and effort to do it. (No one else on my team would know how to do the swaps.)”
“Obviously more memory is good, more context, bigger models, but some jumps don't actually unlock a meaningful difference in ability to run different or better models. For example, I don't currently view jumping from 32+16 to 64+16 as a particularly worthwhile upgrade as compared to going to 32+32, though correct me if I'm wrong. I'd like to build a DDR4 + HBM2 based inference machine to complement my main, 32 GB DDR5 + 16GB GDDR7, computer. The idea is that even if the hardware i”
“Yeah, I think they're a good bang for your buck with larger models. Once you creep over 48-64GB of VRAM needed to run models graphics cards become the more expensive and more power hungry option.”
“That's like saying "why run qwen 3.8 flash when GLM 5.3 exists" Yeah if you want to run models larger the ~64GB then buy an M5. I'm just saying depending on your goals a couple 3090s could be a great fit if you want to still use a capable model with solid speed and keep costs reasonable. I get 130-140 t/s on 3.8 27b with it topping out around 1000t/s for 8 concurrent requests on 2x 3090s. But if they decided they want to run 3.8 flash then it starts to make more sense to go wit”
“This may be somewhat specific to Qwen3.8 27b and the Apple M5 series, perhaps, but enough of us are running this combo that it's worth tossing out there. GGUF models with MTP have been the fastest way to go for some time for token generation, except possibly for a few tweaked MTPLX models running on alpha-stage MLX forks. Prefill, however, was still much faster for M5s under MLX. This has caused me to switch models depending on the expected generation/prefill mix, which is annoying. While I”
I wanted to try out Omarchy before doing a full install, but I only have an M series MacBook so I would have to find a laptop and there's no official M…
“For AI inference on an M-series Mac you want to run MacOS. Linux on M chips is not mature for inference yet, you will lose a ton of performance and create extra headaches. Totally understandable if MacOS is a deal breaker. It's not for everyone.”
“Is anyone actually managing to use LLMs for serious agentic coding without needing a charger with them at all times? And side note; do your laptops not get super hot and noisy? lol submitted by /u/maddie-lovelace [link] [comments]”
“Honestly, Mac Studio may be the better bang-for-the-buck choice if the goal is a quiet personal box for smaller or mid-sized local models and you don’t need CUDA. From the Dell side, the Dell Pro Max with GB10 is aimed at a different problem. It’s a compact Linux/DGX OS AI appliance with 128GB of unified memory, support for models up to 200B parameters on one unit, and up to 400B when two are connected. That makes more sense when students need larger models, NVIDIA/CUDA compatibility, private or”
“We are planning to host four M5 ultra Macs so that 100 users can use them as Openclaw. There will be no other burden, only inferences will be applied here. Can this handle 100 users? I'm considering either Qwen3.8 27b or Qwen 3.8 Next Flash, and I'm curious about the range of realistic models. Realistically, we should probably consider up to 100 users when there are 30 to 40 users stationed there and occasionally 80 to 90 users request at once submitted by /u/Interesting-Prin”
“Initially I wanted to buy W7800 due to its 48GB variant(which's good for 30B range models with context, MTP, Vision, etc.,), but price went up suddenly ~3X. Not worth buying RDNA3 card at that high price. AMD is infamous for dropping support for old cards. So I chose R9700 which's RDNA4. Wish R9700 came with better bandwidth.”
“The CMP 170hx is by far the best bang for buck if you're willing to deal with custom drivers. It's an A100 with about 2/3 disabled. $2k, 64GB VRAM, 1.8TB/s bandwidth, 200ish BF16 TFLOPs. Pipeline parallel it will blow the head off a DGX Spark, and with 2 cards tensor parallel is very feasible (even though they're limited to PCIe 2.0).”
“Got my V100s from China today. Testing on my bench now with my trusty arctic S8038-7k. It's much quieter and more compact than those blower fans and can comfortably cool two V100s running Qwen 3.8 Q8_K_XL in llama.cpp with -am tensor. Temps peak in the high 60s.”
“There is one way to make the Mac studio look like a bargain or at least competitive”
“Symmetrical data is all that matters here, so yeah maxes out at 80Gbps. 120 asymmetrical for monitors is a moot point.”
“It is crazy how fast prices are increasing. I'm pulling my hair out to keep ahead of this for students. Servers aren't even an option any more. submitted by /u/geekender [link] [comments]”
“Hi all, I've got an opportunity to use grant money to buy a pretty fancy work laptop. I do mostly computational work, and I have consistent access to a huge GPU cluster. I know people like Macbooks but I really don't like the interface or keyboard. Does anyone have any other recs? I've used Lenovo Thinkpads my whole life but now I'm second guessing myself submitted by /u/pavelysnotekapret [link] [comments]”
“"They'll scrap 16Gigs cards from each one of them, and combine them to make a bigger one" wow dude, you're genius”
“openAI won't use datacentres but mac mini for inference. Good, fair point well done”
“Because they're run by completely tech illiterate people with no sense of how finances work. Small wonder the chinese unicorns create killer models and run them on domestically produced cheap hardware. OpenAI and Anthropic won't be around for long.”
“I don't know if it's token/watt. Nvidia GPUs are definitely faster but not sure about energy efficiency.”
“Normally, whenever a new model dropped, I always chose the non-vision version just to save VRAM; I though that only use case was when you were the one sending the picture. However, with the release of QWEN 3.8 27B I decided to give it a shot, and it has been one of the best decisions I have made, as this makes the model way more capable for autonomous coding. With no vision, the model will try to complete the task and get back to you once it thinks that it is done with no problem. But there are”
“I'm thinking of buying a 3XS from Scan (UK based): https://www.scan.co.uk/products/3xs-dbp-g1-32r-amd-ryzen-9-9950x-64gb-ddr5-32gb-nvidia-rtx-5090-2tb-m2-ssd-ubuntu It's for my general software development work and for running models locally (e.g. Qwen3.8-27B, Minimax H3). I can't justify spending a fortune, I just wanted to open a few more doors. The specs: Scan 3XS DBP G1-32R, £5,299.99 (approx $7,200) : - RTX 5090 32GB - Ryzen 9 9950X - 64GB DDR5 - 2TB 990 Pro SSD - 1000W PSU - Ub”
“There's a 2 VM cap to MacOS virtualization per the documentation you just cited... It's enforced by their EULA and at the kernel-level. They need to buy more hardware if they want more MacOS sessions.”
“I've been experimenting with local models on an M5 Max MBP w/ 128GB of RAM since March of this year. Generally I've had very good results. Where things were lacking initially was with tool calling and the need to rely on tool calling for functionality like web search, which is otherwise well integrated in the cloud models. There is also a lot more work required on the harness side, however at this point (August 2026) there is not only much better tool calling in local models, but community su”
“Just ordered a TB5 enclosure plus 4TB SSD to prepare for M5 Ultra 512GB. I am planning to download the following models in advance. Would like to ask the M3 Ultra 512GB owners whether the quants are the "best" performing ones given the 512GB constraint. What's the max size of context length you can get for GLM-5.3 and Kimi-K3? If possible, can you post your pp and tg for them? Does it make sense to run 8-bit models for glm-5.3-flash, qwen3.8-flash-next and DSV4F? Are there other bi”
“This makes absolutely no sense. Anyone who has ever done even the tiniest amount of fine-tuning knows that Macs are terrible for fine-tuning. They're relatively okay for inference because they're the cheapest option, or used to be. Even there, you have performance issues. This is a nonsensical post.”
“Mac Studio, 96GB unified memory ( M3 Ultra ). I want the largest usable Qwen3.8-Flash-Next setup, and I'd rather not burn 100GB of bandwidth on the wrong download. Here's my math. Please tell me which parts are wrong. What makes this model weird 177B params total: a 125B MoE trunk + a 51B n-gram embedding table + a 4B MTP head. The table isn't compute, it's a lookup — 20M hashed bigram/trigram entries, 320M rows of 160 values, and each token touches exactly 16 rows, roughly 2.7 K”
“sounds like 48gb refurb is the play here. as my 5080 +5060ti pc will always be faster”
“Bought a 14" MacBook M5 Pro 48G in May for $2299. Did try some llms but the 14" apparently doesn't have better thermal. Plus you would endure slow prefill + decode given relative low compute & memory bandwidth. I later got a 5090 prebuild and also built a RTX Pro 6000 box”
“There’s reasonably small models that fit within 50~ gb with context at high tps locally. There’s really not ones that fit in less that this guy can’t use remotely. But yes, I don’t spec less than 128gb unified now on my purchases now.”
“A Laptop will do I think. But it should be powerful enough. The MacBook Air does the job like I mentioned, but you‘ll hit a wall pretty soon. I find that annoying. If I had to buy something new for the purpose of video editing with the setup you mentioned, and I wanted it to be a laptop, I‘d probably buy a MacBook Pro with a M5 Pro with the 20 GPU Cores, and 48-64 GB of RAM. The M5 Max would of course be better but it‘s also very much more expensive. I think there are probably non-Apple devices”
“i’m looking to upgrade from my 14inch M1 pro 16g/512 to a 16 M5 pro. i’m deciding between 48gb and 64gb ram i also have a desktop with - 5700x3d - 64gb ddr4 ram - rtx 5080 + rtx 5060 ti 16gb (31gb usable combined vram) currently running qwen3.8/3.6 27b q6 and qwen3.6 a35b i use the Mac for Docker/K8s, development, and Moonlight streaming from my desktop. I have also connected to my desktop’s LLM server over Tailscale and it works great my main issue with the M1 Pro is that 16gb is too limiting e”
“Studio M5 Max 128GB makes running Qwen 3.8 Flash Next (Q4) at okay speed (on M4 Max, PP 480 t/s, PG 38 t/s) to smooth sailing while leaving enough RAM for a larger context. The Flash model is faster and slightly better than the 27B dense model. The ideal Mac to get is the 256GB M5U but the M5 Max 128GB provides a really good middle ground.”
“Trying to decide between three Mac Studio configs and I keep going in circles. German pricing, so: M5 Max, 64GB — €4,099 M5 Max, 128GB — €5,849 M5 Ultra, 96GB — €6,589 Please do not tell me to get the 256GB Ultra. I know it's the answer for GLM-5.3-Flash at Q4. It's ~€11k here and it is not happening. I'd rather hear why one of the three above is enough (or isn't). What I actually do with it Processing private/personal data locally — that's the main reason I want this on my d”
“no such thing as unusable. just gotta apply it to a different task. Some models are large, and always suitable - others are smaller, focused, faster, and powerful when harnessed. Perhaps 15x as many mini-thoughts could be just as bountiful as one large one. Depends on the use-case...”
“Yes, but you will have to sweep cache size to find the best value for you. It should be around 64gb, less if you have lots of apps running”
“Apple is rumoured to make a base M7 with 235-240GB/s. That’s close to Strix Halo. M7 Pro should double and M7 max double M7 pro bandwidth.”
“ohh that makes sense, hilarous a random guy with a 128gb ram mac could do the same though lol...”
