PANE

VRAM Runs Out Before the Model Fits

Users are consistently frustrated by the limitations of their GPUs when running large AI models. They're grappling with insufficient VRAM, seeking workarounds like using multiple GPUs or offloading display tasks, and desiring better tools to estimate VRAM requirements. This impacts productivity and experimentation with cutting-edge AI applications.

aiproductivitygpumodelsvram
FIT
0%
SIGNAL
64%
SOURCES60
FRESHEST POST11H AGO
TRACKED SINCE101D AGO

SOURCES (60)

exact same issue, but on windows and with a 7900xtx + 9800x3d :/

r/selfhosted11h ago
Source preview · reddit.com

Thank you will takea look

reddit.com11h ago

Yes unfortunately this GPU is too old :( It's because of software limitations and AMD doesn't really support it anymore either

r/selfhosted11h ago
Source preview · reddit.com

Strix Halo 💪

reddit.com11h ago

> If they were ever going to do it, it would've been when they switched to Apple Silicon chips, and...they didn't!They did, though;- macOS lost support for eGPU and dGPU drivers- OpenGL support was frozen to focus on Apple's mobile GPU API- Apple promoted Mac Catalyst apps instead of pushing Mac-native ones- The iOS and iPadOS App Stores became available on macOS alongside the native store- The macOS design language was reworked to take cues from overly padded touch-first design- iBoot became th

HN1d ago
Source preview · reddit.com

Free central heating too

reddit.com1d ago
Source preview · reddit.com

least sus files

reddit.com2d ago
Source preview · reddit.com

I dont need love i need Vram and bandwidth

reddit.com2d ago

25000$? I'll pay ~5000 usd incl. taxes MAX, new. Not worth it to me, I've had enough ROCm pain in my life. Now that hotaisle is 3$/hr not 2$/hr I can't imagine touching that disaster of an architecture unless Lisa Su pays me

r/LocalLLaMA2d ago

I wouldn't worry to much about the 16x to 8x drop. For inference it's mostly about delay not data transfer rates. For instance you can set up rdma over two computers using a 4x connectx4 networking card and because it removes the latency aspect and also has more bandwidth than is needed for the cross talk between the cards during inference it allows for full tensor parallel. I'm setting this up on my two computers with 2x3090 each. The cards pcie are bifurcated so each card has an 8x

r/LocalLLaMA3d ago

If we assume we get a doubling of performance every 18 months then: GB200 compute: ~40 PFLOPS FP4 5090 : 3.3 PFLOPS FP4 For 12x improvement need about 5.5 years. So approximately 2032. If you do it for memory size or cost drop you get a similar result. If you do it from a raw model capability perspective then the time span is shorter because of software improvements. One estimate could be a doubling in capability every year, then need about 4 years or 2030 time frame. I've seen other estimat

r/LocalLLaMA4d ago

I've tried a few small models with mixed results. Anybody have a good go-to for this? submitted by /u/mil_phickelson [link] [comments]

r/LocalLLaMA4d ago

Now hold up, hear me out. This is purely from a technical perspective and not an economical one. Just an engineer having a thought. When SLI originally came out, one major issue with it was that every single game had to manually optimize itself so that offloading work to 2 GPUs would be possible. And gamedevs simply do not have time. Before we address why that was an issue. First let's take the naive approach of dual GPU offload. What if we rendered half of the screen on one GPU and half on

r/gamedev4d ago

Can you explain this? I have a Macbook M5 Max with 36 GB of Ram and I can barely get Qwen 3.6 27B to run with a 40k context window. No idea what I'm doing wrong, I'm running it through Cline and when I try to have it analyze some code or something it just... sits there. I was under the impression this was supposed to be blazing fast, but it certainly doesn't feel that way.

r/LocalLLaMA4d ago

I figure I can run 30b or even 72b models on it, but this is my first time running my own local LLM. I want to play with prompt engineering and see the most complex things I can get it to do. Any hot tips? I know it all moves fast and this sub is wired in I have 32GB ram also edit: omg this sub is so helpful lol, thank you everyone submitted by /u/ThomasAger [link] [comments]

r/LocalLLaMA4d ago

I've been casually following the progress of this card for a while and it seems a bit off. From the comments on github, it seems like that group has privately figured out to unlock the VRAM around a month ago back when the cards were still ~$200. Then it gradually rose to ~300 until a few days ago when it started spiking to (as of the time I posted this) $500+ due to the release of that paper and repo listed in OP. If every CMP170HX can be unlocked to 80GB (or even 40GB), then they would&#39

r/LocalLLaMA4d ago

They're almost always hooked up with pcie x1 per card, that's way too slow a link for meaningful tensor parallel.

r/LocalLLaMA4d ago

You can only run them on Vulcan as they're not supported by ROCm. I had an RX580 running a llama.cpp container and it was brutally slow even on Qwen 3.5 9B with a heavy quant. I'm talking under 10 tok/s. I found a deal on an open box 5060ti 16GB and switched it out. The difference is massive. I can run Gemma 4 12b with a more reasonable quant and 64k context and it is conversational. It was fun setting it up just to show I could do it, but the webui server and Nginx would time out waitin

r/LocalLLaMA4d ago
Source preview · reddit.com

256k is not full. The model will do 1M

reddit.com4d ago

If you're not running a datacenter and you don't have to worry about permits for MW of power, how is less power so important? You can lower power limit on RTX 6000 Pro 600W too if you want and get a similar effect, but with a better cooler. I think eh floor is 400W there tho but I'm not sure. As far as I've heard the 300W Max-Q is also louder in operation from the 600W variant, it's a blower cooler and they are always loud.

r/LocalLLaMA4d ago

Just noticed that miners wants to get rid of their old riggs for cheap , is it a bargain for local AI or is it trash? submitted by /u/StandardLovers [link] [comments]

r/LocalLLaMA4d ago

How much context do you and the super users need per session? How much context do regular users need? What does a session look like? Is it a zero-shot (maybe with a follow-up) or an extended heartfelt dialogue between man and machine? Maybe you need 32k context per session, maybe you need 256k. Makes a big difference in VRAM footprint at runtime. If any use cases involve OCR or vision, offloading those tasks to a sidecar model helps reduce context use in the user session. If you put 4xRTX Pro 60

r/LocalLLaMA5d ago

If you don't know, the CMP 170HX is essentially a A100 that has had it's compute and memory crippled so it can only mine crypto. It was a product of the crypto craze, and was released shortly before the crypto crash. Well, I was scrolling around and found out, apparently, it can be reverted back to an A100 using an exploit in Nvidia's Falcon security processor. https://www.researchgate.net/publication/408132536_A_Canary_in_the_Crypto_Mine_Defeating_Stack_Protection_in_a_GPU_Secure_Co

r/LocalLLaMA5d ago

100%. Btw, even if someone only has 128GB, they should try 4.7 at Q1 (100GB). It will struggle a lot as you approach 30k tokens, and it will probably mix up some things, but it will give you Opus-tier prose.

r/LocalLLaMA5d ago

Two DGX Spark and a Connect-X7 cable give you about 250GB of usable memory for $7000 8000 USD. This allows using some interesting models at 4-bit. For what seemed like an eternity, the only serious models in that size were GLM 4.5/4.6/4.7 (194GB), Qwen 3.5 397B (210GB), and older MiniMax. Then 3 months ago we got MiniMax M2.7 (131GB), Deepseek V4 Flash (160GB) and Xiaomi MiMo 2.5 (181GB). Shortly later we got StepFun 3.7 Flash (129GB), and now we just got Tencent's Hy3 (182GB). I know this s

r/LocalLLaMA5d ago
Source preview · reddit.com

Nemotron Ultra would fit in that setup

reddit.com5d ago

I've started building my rig and it turned out 24GB VRAM is not enough for my tasks (inference only without offload to RAM). What will be the best way for an upgrade for dual 3060 - add used 3090 for 700-800 euro or add 2 more barely used 5060Ti 16GB for the same price. Slower newer but 56GB vs 48GB total. Don't know how nvfp4 is usable actually. submitted by /u/esw123 [link] [comments]

r/LocalLLaMA5d ago

Ah seems like the bar size is too small. Which mobo im curious? I like the spacing betwee the gpus. Also what is your p2p bandwidth using the official test? Mine is slower with p2p, but lower latency.

r/LocalLLaMA5d ago

My desktop has a Ryzen 5700X with 32GB of DDR4 and a 5070Ti (16GB). I also recently acquired an RTX 3060 12GB. While it's nice being able to run the 27B Q5_K_XL model entirely on VRAM, I'm limited to a context length of around 120k (if I quantize V to Q8, leaving K at BF16). Before I got the 3060, my daily driver was 35B Q6_K_XL (no MTP) with `-ncmoe 30` and V quantized to Q8. I was getting around 750pp and 55tg, which is fast enough for my use. I'd hoped to use the 3060's additi

r/LocalLLaMA6d ago

32 GB GDDR6 each, 512 GB/s bandwidth, compute seems solid enough. They're not that expensive. What's the catch? There has to be a catch, right? I'd rather get some V100's but they're double the price. submitted by /u/_TheWolfOfWalmart_ [link] [comments]

r/LocalLLaMA6d ago

The PCIe loss isn't great, you might indeed need a new PSU but most of the time a motherbord with PCIe 5.0 4x port works just fine. PCIe 5.0 x8 on dual slots is a bit faster, but this applies only to model loading and unloading, so it's not THAT much of speed difference. You could also buy 2 3090s and connect them with a NVLink, then you have 48GB unified vram ;)

r/LocalLLaMA6d ago

They really don't, no. Vulkan: 50 lines to allocate device memory. Cuda: One single line. What kind of extensive documentation stack do you want for functionality that is trivial in Cuda? And that exact issue continues through every little step of the way to your first usable application. I know there is VMA, it is a very poor solution to a problem that shouldn't even exist, and it only poorly addresses one of 100 parts of the API where Cuda is vastly simpler than Vulkan. Cuda also doesnt force

HN6d ago

Ran into this recently and figured others might find the approach useful. A 7B model in FP16 uses about 14GB of VRAM. K8s gives it an entire A100 with 80GB. Device plugin says fully allocated. Next pod stuck in Pending. 66GB sitting idle. Classic problem. There are three ways to approach this, and the way I think about choosing one comes down to a single question: what is the trust boundary between workloads? If everything is on the same team, dev/test, and you just want pods to stop getting stu

r/mlops6d ago
Source preview · reddit.com

2 more than you came up with

reddit.com7d ago

Search for https://huggingface.co/SamPurkis/gpt-oss-puzzle-88B-GGUF It works with the PR.

r/LocalLLaMA7d ago

You might get those 4g decoding problems on gpus with no video output but if it has a video output it will almost always just work.

r/LocalLLaMA7d ago

Second slot is slower but should be okay to handle a 5060 Ti or you can get a used workstation card like the RTX A4500

r/LocalLLaMA7d ago

I only spent 2 hours making the bios accept the dual gpus, only 5 hours configuring VLLM to run deepseek v4 flash dspark, but totally worth it. I truly believe in the near future we will have to rely on ourselves. submitted by /u/BitXorBit [link] [comments]

r/LocalLLaMA7d ago

Here are my fingers typing again a question I’ve researched without real people. The gpt and Google responses are always the same and I don’t like the answers. I have a 5080 rtx. Yeah, mistake. I’m a software engineer but also plan on running several llms locally for product development. Inference speed isn’t too important, gaming isn’t important, but model size is. I’ve had a dgx spark and returned it - I liked it but am testing different setups. I’m considering some apple development but I wou

r/LocalLLaMA7d ago

Whatever fits from my list https://preview.redd.it/3zp7je9jm1dh1.png?width=1714&format=png&auto=webp&s=bf831168682ff70215c911af431a73dec560bd2b

r/LocalLLaMA7d ago

HBM doesn't matter on the P100. The card is so horribly compute choked due to the lack of DP4A instructions and tensor cores that it can't really use all of it. Source: ran 2 P100s for a few months, ended up selling them for this exact reason.

r/LocalLLaMA7d ago

Hi, I'm a long time lurker and this is my first post so please be gentle. I'm running a Ryzen 9 5900x 64gb paired with a 5080 RTX. It's been great so far. My daily driver is qwen 35b a3b, while it doesn't break any speed records, it is usable (Q6, about 30-ish decode with 256k context). The thing is, I'd like to dabble with dense models (e.g., Qwen3.6:27b, and Gemma4: 31b) as well with bigger MoE models. I was thinking about the following upgrade paths, and wanted you guys to

r/LocalLLaMA7d ago
Source preview · reddit.com

Bro is rich now (512GB ram)

reddit.com7d ago

yeah, that could be the most straightforward approach and you might wanna test it first if you have a big model with satisfying performance. If your dataset is not big enough then by fine tuning bigger model you may just overfit it (same risk goes to training the smaller model if your dataset is small or not representative)

r/LocalLLaMA7d ago
Source preview · reddit.com

ooh, might try after this. Thanks

reddit.com7d ago

7900XT prices start at about $625. That wouldn't be a bad price for what I'd be getting, but it's still double the budget. And I've been burned too often when trying to sell used electronics and also musical instruments, so I just don't any more.

r/LocalLLaMA7d ago

Scheduled reboots might help as a temporary workaround, but I’d avoid treating them as the fix. Since performance comes back after a restart and then degrades again, a memory leak or resource handle leak in the PMS client sounds quite possible. I’d monitor working set, private bytes, handle count, thread count, CPU, disk latency and network response times over several hours, ideally comparing a healthy terminal with one that has started slowing down. Process Explorer and PerfMon should help show

r/sysadmin7d ago

Hey everyone. Trying to figure out the best setup for my hardware. I've got a 64GB DDR5 RAM laptop and A 20GB 7900xt egpu. Usecase is only pi-coding-agent locally. So far, Qwen3.6 35B A3B has been the only model I've found that would work well with CPU offload (MoE) at higher quants - I've managed to fit Q8 at 100k context or ud-q6kl at full 262k context. It works okay, on the slower side but for me anything with 20+tks is very useable. However, I'm interested in what you guys ar

r/LocalLLaMA7d ago

Hey guys. Just a curious and maybe been feeling a little burn out on this whole experiencing local AI and stuff. I own a rig with quad RTX3090s and 96GB for ddr5. But its on am5 platform so the pcie lanes aren't max but I have it running in pairs of nvlinks as well so cross gpu pair is fine. Pair 1 runs at gen 4 x8 and pair 2 at gen 4 x4. I wanted to know was , what would you do with this setup? Specs: 9900x 48GB ×2 Ddr5 5600 mhz cl 30 ram Asus X870e Proart RTX 3090 ×4 nvlink ×2 Gen 5 nvme 2

r/LocalLLaMA8d ago

so it heavily depends on a model In my case, they are exactly the same brand and model. Both are running the same VBIOS. The only difference is when they were made. Based on the serial number, the high wattage one was made earlier than the low wattage one.

r/LocalLLaMA8d ago
Source preview · reddit.com

Is the gap between q4 and q6 really as big as people say?

reddit.com8d ago

how much vram do you need and what model do you think is the next major upgrade from the good old qwen 3.6 27b as of today? submitted by /u/WhatTheFlukz [link] [comments]

r/LocalLLaMA9d ago

Saw this on a random social media post. I believe what we are seeing here is two sxm2 gpus, right? V100 or a100? Assuming two v100 with 32gb each, can you get nvlink working with these weird aftermarket base boards? Costs, experience, regrets? Not trying to build a similar case, but i am curious to compare the costs against a similar consumer/gaming setup. submitted by /u/SnowyOwl72 [link] [comments]

r/LocalLLaMA9d ago

Here is the upper limit of what can be done with $100 bucks worth of video cards. You can have 3 concurrent users with plenty of context, better speeds or close enough speeds than a bunch of cards that provide less VRAM and cost 4+ times. 0.00.008.388 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg) 0.00.008.391 I device_info: 0.00.089.439 I - CUDA0 : NVIDIA P102-100 (10144 MiB, 10013 MiB free) 0.00.197.645 I - CUDA1 : NVIDIA P102-100 (10144 MiB, 10013 MiB free) 0.00.197.656 I - CPU :

r/LocalLLaMA9d ago

I would totally do that if my budget was enough to stretch to 4x48GB! The big difference between 96 and 128 is quite a lot of bigger dense models can be run at 128 than would fit on 96 without offloading

r/LocalLLaMA9d ago

Even more reasons to remove the 2x 3090 entirely. There isn't that big of a difference between 96 and 128gb. Remove the 2 cards and add 4x 48gb instead of keeping 2x 24 + 2x 48

r/LocalLLaMA9d ago

I'm in a position to upgrade from the classic dual 3090 setup to something that can run the next step up, eg 100-110GB+ model files. I've got an asus x299 ws pro and a 1600w PSU, with 128GB ram, and I want to add to my 3090s, rather than replace them. I'm trying to weigh up modded 4090s (2x48GB), 2x A6000s, or 2x5090s. Anybody made similar hybrid setups with these? Any thoughts on what's going to get me to where I want to be (which is basically a useable quant of DeepSeek V4 flas

r/LocalLLaMA9d ago

I had one RTX3060 and was able to use all models like Qwen3.6 27B, Qwen3.6 35B and Qwen3.5 122B with CPU offload. I've added second 3060 and now only dense model can be loaded as usual, 35B loaded only once and 122B refuses to load at all. What setting should I change or where to look for potential problem solving solutions? Update: Seems 12 layers in VRAM is the limit for dual 3060 with qwen3.5-122b-a10b. Reduced to 12 and model loaded. 35b-a3b loaded as well after PC restart and when I cha

r/LocalLLaMA9d ago

Hi there, got a 7950x,128GB DDR5, RTX 4090 and RTX 3090TI. I'm currently running Qwen3.6 27B Q8 with 262k Context at Q8 with llama.cpp. It's not touching the DDR5 RAM at all but at the same time I couldn't get 122B A10B or the likes to run. Is my FOMO justified or isn't there anything better than this model to run currently? submitted by /u/Common_Warthog_G [link] [comments]

r/LocalLLaMA9d ago

I am really excited to share this one. On the X99-E-WS motherboard.. while old and PCIE 3.0 - I think it's still pretty capable for what I'm trying to do (1TB VRAM across 3 machines). The board has seven physical PCIe x16 slots shared through the onboard PEX8747 PCIe 3.0 switches and with a 40 lane CPU + all seven slots populated, the board supports an x16/x8/x8/x8/x8/x8/x8. What I tested was putting a PEX8749 card on the x16 slot so that 4x MI50's ran on the switch thus freeing up 3

r/LocalLLaMA9d ago

SOLUTION LANDSCAPE

Brought to you byTop Sectors

A Player feature.See how many ways this pain can be solved, who's already building, and where the gaps are.