Running Multiple GPUs on Oddball Hardware Is a Mess
These users are grappling with the limitations of their existing hardware (motherboards, power supplies, PCIe slots) when trying to leverage multiple GPUs or higher-end GPUs for demanding AI workloads, particularly large language models (LLMs). They're weighing the performance trade-offs of dual-GPU setups versus single, more powerful cards, and struggling with VRAM constraints and PCIe bandwidth limitations.
SOURCES (60)
“My 5090 is running at gen 5 x8 My 3090 only has a gen 4 connector so it can run at most at gen 4 x8. You cant run a 3090 at gen 5 x8 (gen 4 x16) on a taichi because the physical wiring…”
“A few months ago, I was lucky enough to purchase two RTX 3090 GPUs for a total of $1,000. I originally built the machine for ComfyUI, but I also use it as a remote GPU server for other computers on my network. One thing I still haven’t fully figured out is how to choose the right local LLM for my hardware. There are so many models, parameter sizes, quantization levels, formats, and inference engines that it can be difficult to know what will actually perform best. My current setup includes: 2× R”
“Hello everyone, Recently sold my EVGA RTX 3090 FTW3 Ultra Hybrid after not being able to find a second GPU in a good shape and a reasonable price. My new setup: Motherboard: MSI MPG Z890 Carbon WiFi GPUs: 2× NVIDIA GeForce RTX 5060 Ti 16GB tried both 595 and 610 drivers GPU slots: CPU-connected PCIe5 slots configured x8/x8 Memory: 48 GB DDR5 8800MHZ CUDIMM CPU: Core Ultra 7 270k plus M2_2: NVMe SSD I am testing PCIe P2P on a dual RTX 5060 Ti workstation using the patched NVIDIA open kernel modul”
“Thank you for the information. I just wanted to buy me a new Intel CPU to run my RTX 5070ti set up when I build my PC like that so that that I can always add two GPUs. I am only 16 so that's all I have.”
“hmm.. things are fine on a taichi z890 for me, 285k. if ur board supports it, you gotta have x8 x8 enabled in bios, and make sure you populate your nvme and pci slots correctly.. use the wrong pci or m2 slot and you get nerfed to x4 or x2. but otherwise no issues”
“I was so annoyed when I realized the X870E ProArt couldn't use 2 Gen. 5 NVMe drives when the PCIe is in 8x 8x mode.”
“Since a lot more people are trying to build their own multi-GPU machines, I thought I should help to prevent a common mistake people make with building multi-GPU machines. Which is using an Intel consumer platform like Z890 for multi-GPU setups. Although the CPU provides 24 PCIe 5.0 lanes with 16x available to bifurcate to 8x8x on two PCIe x16 slots on the higher end boards, this is completely useless for AI inference/training workloads that require P2P between the GPUs. In my testing I used an”
“Recently old data center GPUs skyrocketed (like the mi50 going from ~100$ to 300-400) but I've seen that site recommend a few times and I see they have the same card for 50-100 USD. Is that a scam? submitted by /u/Far-Classic-9963 [link] [comments]”
“If it weren't for power consumption, I'd go for a 4x or even 8x setup, but as the adage goes, If my grandma had wheels, she'd be a streetcar. Now, 2x V100 seemingly strikes a balance between, on one hand: a non-negligible amount of VRAM, namely 64 GB (capable of running such models as the Qwen 3.6 family at GGUF Q8, for instance); a fair 900 GB/s of bandwidth; a helpful 300 GB/s GPU interconnect thanks to NVLink; affordable upfront capital outlay; non-inordinate power consumption--ad”
“If you want to run tensor parallelism and have no nvlink bridge, PCIe bandwidth matters a lot. x4 tanks your performance.”
“I’m troubleshooting my dual 3090 rig and would appreciate ideas on what to do next Specs: EVGA X299 FTW K motherboard Intel i9-7940X 128 GB RAM 2 × RTX 3090 Founders Edition 1600 W PSU Asrock Each GPU is installed in its own x16 PCIe slot There is roughly one slot of space between the cards Both GPUs run long AI workloads without issue when tested individually. The problem only occurs when I run the workload on both GPUs at the same time. After several hours, the entire computer hard-locks. The”
“Hello, What are the consumer motherboards that can physically fit 2 RTX 3090 FE, with enough room between them for airflow ? Which motherboard are you guys using for 2×3090, and how is it ? I already manage to build a list of motherboard than can run 8x/8x by filtering "SLI compatible" cards on PCPartPicker (read the tip here, ty btw) I already found multiple posts for this issue but can't really find information on the remaining space between the cards. submitted by /u”
“Currently I have 6x3090s but was considering getting a pcie splitter and using the last free slot of my motherboard to add two more. Would that be worth it? Tbh it would be satisfying to achieve such a build since it maxes out the motherboard but not sure it will offer much value other than satisfaction looking at it. I was thinking, Deepseek flash V4 Q8 would then become possible, and with enough extra memory for a big context. submitted by /u/dazzou5ouh [link] [comments]”
“Oh good, I don't have nvlink myself and it works just fine. You may be able to fit one in the mobo you have with a m.2 to 16x pcie. It has 4x transfer in one of those slots which should work for the extra card. Another way could be one of those plex cards that have a bunch of m.2 slots, they are like $100 on Amazon. They split the slot into four m.2 slots and if you put it in your 16x GPU slot, if it's your only one it would split into 8x with two cards plugged in. May want to talk to De”
“I think it's worth noting that with the motherboard you have you will also be compromising the performance of one of your cards. That B550 board has this PCIE topology with the CPU you have: 1 x PCI Express x16 slot (PCIEX16), integrated in the CPU: - AMD Ryzen™ 5000/3000 Series Processors support PCIe 4.0 x16 mode Chipset: - 4 x PCI Express x16 slots, supporting PCIe 3.0 and running at x1(PCIEX1_1, PCIEX1_2, PCIEX1_3, PCIEX1_4) So your running your second card at PCIE3 x1! Edit: The Asus B5”
“I think the problem is not the model it is the kinds of GPUs. The 3060 is a lot slower, than the 5070 Ti. So the faster GPU has to wait for the one. When we use GPUs at the same time the slow one is usually what decides how fast we can go. The 3060 and the 5070 Ti are just not the same. That can cause problems. The 5070 Ti is waiting for the 3060. That is not good.”
“...I have no explanation why I did this, I don't even use my Local AI 🤷🏽♂️. But I did a lot of 3d printing to get it set up. Yes it runs off one 1400W PSU. The external card is using Oculink via m.2. The CPU is a 9850x3d, RAM is 64 GB 6000 at CAS36 (🥲) I honestly wish I could afford better cards, but I'm in Canada and it's usually over a grand for anything with over 16gb. submitted by /u/Vaguswarrior [link] [comments]”
“For AMD there is a Linux kernel arg gtt-something that sets the max iGPU share. Not sure if same for Intel. I just saw this but it seems like a bios/driver setting maybe: https://www.intel.com/content/www/us/en/support/articles/000101789/graphics.html”
“Hi there I got a bd790ix3d Mainboard (only one pcie) with a rtx5070 ti 16g. Vram, 96gb ddr5 RAM and a free m.2 slot (4x pcie). I want a little more LLM power... Should I get a m.2 to oculink adapter and a external gpu? Everything is in a sff case, so there is no room inside the server for another GPU. I've read it is possible to pair Nvidia with ati for interference? What's the GPU of choice for tight budgets right now? submitted by /u/Designer_Elephant227 [link] [comme”
“the extra cost for shroud and fan which makes them a lot more You don't need a shroud. A standard PC slot exhaust fan works just fine. That costs like $10. Take off the mounting bracket and cut a couple of slots to clear a bracket, or just take the bracket off the V340, and it slots into the end. So it's $120 for the pair all in. also that they are 60% slower at PP. What are you basing that on? Remember, there's two GPUs per card so you can take advantage of TP even with one card.”
“Right now I have Gemma 4 26B-A4B (Q4_K_M) running reasonably well (12-15 t/s) on my hardware, which is an i5-8500, 48 GB of DDR4, and an RTX 3060 (12 GB)—PCIe is Gen 3. However, it's just not very smart. It's good at paraphrasing everything I say, which helps if I'm trying to make sure I've covered every aspect of a topic I'm writing up, but it doesn't generally chip in with anything insightful. I've also tried running the 31B model, and it really is a meaningful step”
“Hello all, sharing some data points: In this setup there are 2 cards connected direct to motherboard via PCIE 3 16x slots. Running llama.cpp. Ubuntu 24.04, single xeon motherboard The cards are 1x RTX 3090 24GB power limit 250W and 1x Titan RTX 24GB power limit 225W nvidia-smi with Qwen3.6-27B-UD-Q4_K_XL.gguf loaded at 180k context: https://preview.redd.it/t28n6nxiooch1.png?width=764&format=png&auto=webp&s=c8f9507a39c09a6c05a5d401a6f6aeab1a5af94b My test today was comparing this setu”
“I have a GTX 4070 12GB card in my computer and have 2 1050 Ti's collecting dust. Would there be any benefit to adding them to the machine? submitted by /u/chuckbeasley02 [link] [comments]”
“I have a dual Xeon system with 512gb memory. I was running two 5090s, then sold them and upgraded to a pro 6000 Blackwell. The 6000 was repurposed to another system and I don’t have the budget to buy another one at the moment. I was thinking of getting a 5080 and v100 or dual 5080 for the time being. Thoughts? submitted by /u/jsconiers [link] [comments]”
“I hate to burst bubbles (AI bubble notwithstanding) but I don't expect these prices to be very friendly with memory increases. I don't know if 5070ti+5080 will get price increases but I would assume so before launch and put 5080 super at $1699 and maybe 5070ti super at $1299? Then 5090 to $2499?”
“It only has like 200GB/s bandwidth lol the RTX 5050 has more bandwidth than this .”
“Current multi GPU setups seem to be bottleneck'd by the overhead associated with keeping the GPU's in sync, (not the bandwidth of the PCIe bus). i.e. TTFT improves when the latency between cards is decreased, but the decode speed seems flat despite doubling the compute, all this while the PCIe bandwidth is nowhere near being saturated. This suggests it is the overhead of making the sync's (not the bandwidth) that is bottle-necking decode speeds. I am looking into some way to increase”
“I have a similar build. I've heard it doesn't make a huge difference. It does depend on what you plan on doing with it though. I've heard PCIe 5 x8 is equivalent to Pcie 4 x16 Following.”
“the time it would take to earn/save a little more to get something better would be less than just firing on what you can afford right now then to soon find you paid for hard limits and lost time dicking around with budget gear.”
“So, as we dual strix halo peasants suffering through our atrociously slow prompt processing, I also thought about using the disk space on both nodes efficiently. As we all know we have our usb4 capable providing us with a 10Gbps link which is basically useless during inference, except the model load time via RPC. And it's also nice to transfer once downloaded model between the nodes faster than the measly internet download speed of 1Gbps. I paid for two usb4 — I'm gonna have my two usb4.”
“I have an RTX 4070Super (12gb) + 64GB ram Ryzen 7700x Motherboard has 2 pcie x16 slots (MAG-B650M-MORTAR-WIFI) I have a budget of around 1-2k $ should i get 2 rtx 4060 ti 16gb and split 1 x16 into 2 x8 with some converter? or are there any better card for locak LLMs in this budget or am i stuck here? submitted by /u/Beautiful_Egg6188 [link] [comments]”
“I’m replacing a 3060 dialed to my 5060. I run together and separate. That setup I was happy with but wanted the consistent parts etc. I added a second 5060 ti 16gb today. Well I got it in and I can’t say it’s better. Prompts that would come out a nice JSON formatted string now goes into a malformed prompt. I’m not using any flags on llama-server, or at least fiddling. The problem seems to be coming from the prompt processing. I will receive the results I want but then will continue into the temp”
“I'm building a local LLM workstation and would appreciate some advice from people already running 2×3090s. Current hardware: ASUS Crosshair VIII Hero (X570) One Gainward Phoenix RTX 3090 Looking for a second used 3090 (not necessarily the same model) Both GPUs will be power-limited to ~250W I'm trying to keep the case budget under 200 euros SEK (including any extra fans), but might stretch if neccesary... So far I've been looking at: Fractal North XL Mesh (looks nice, but worried abo”
“Highlights: 4 x 48GB modded 4090s - 192GB VRAM 128GB DDR5 Pro WS WRX90E-SAGE SE 3000w PSU 240V/30A dryer line Q. Is putting a server on a dryer line a good idea? A. No, or emphatically yes. Splitters on this line are not code compliant, so I have to turn off the server to use the dryer, OR buy a smaller dryer that can go on the 20A. Also I've had two nuisance trips while idle in the past month due to laundry GFCI. A dual conversion pure sine wave UPS is on the way. This room is my only optio”
“I added: 1x RTX 3090 - 610 USD 1x Arc A770 - 222 USD 1x PCIe x1 to 4x USB 3.0 PCIe riser New cpu cooler Specs: Modified Zalman Z9 Plus Case 2x Zotac RTX 3090 24 GB 1x Intel Arc A770 16 GB 48 GB DDR4 RAM AMD Ryzen 5 1600X MSI X370 SLI Plus All parts were purchased second hand except the RAM sticks (before the crisis) and the case. I bought the first RTX 3090 for 540 USD to build this server over a year ago. Findings after 2 hours of testing: I thought the Vulkan backend would work well for multi-”
“https://preview.redd.it/h40uz1bvhn9h1.png?width=808&format=png&auto=webp&s=f68d2640255989fdefa3c6e5e4a5b0e1690731f6 Motherboard is a Asus Proart Creator B850 Neo Slot 1 & Slot 2 (PCIe 5.0): These are the two main physical x16 slots. If you occupy both slots simultaneously, the motherboard automatically splits the CPU's primary 16 lanes into PCIe 5.0 x8 / x8 mode. M.2_1 (PCIe 5.0 x4): This slot has 4 dedicated lanes wired straight to the CPU, meaning it runs at full speed with”
“My end goal is to have multippe small to medium models running locally for data parsing and extraction tasks, working with logs, and many data inputs with slight reasoning capabilities. Then it will be nice to also generate images with it and computer use. I will still have big models like Opus to handle huge design and difficult bug hunting tasks, but as a "junior developer," I want to have the local models. Lastly, if it would allow me to build loras and distilling medium-sized model”
“Tell me if it's a good idea or not, I have zotac solid 5090 with 128gb RAM, thinking of selling only 5090 and getting 5 x 5060ti 16gb also use these PCIE 4.0 x16 Extender Riser Cable, planning open rig for AI, is it good idea? submitted by /u/Specialist_Pea_4711 [link] [comments]”
“I bought the Biostar Z890 Valkyrie because it was on sale and had three PCIe 5.0 slots connected to the CPU (x16 or x8/x8 or x8/x4/x4), which I thought would be great for running dual GPUs for LLM inference. The problem is that now I want to add a SATA expansion card to the bottom PCIe slot, but this will drop the middle slot to x4 speeds. Would I see a performance hit for inference if I run the two GPUs in x8/x4 mode, both when the model if fully loaded into VRAM and when I have to use partial”
“A few months ago, I bought a RX 7900 XTX 24g to start toying with local LLM, at 900€ new. Little I knew that now I want to add a second card to my rig, but prices have gone insane! Adding a new 7900 XTX would cost me 1200€ as new now, used price is around 900€ now, and the last budget option would be going for RX 7900 XT 20g at 700€ best, which is still quite expensive for old cards. Nvidia cards are through the roof, but as my setup is RNDA 3 I should go 7900 XTX or XT to keep llama.cpp happy w”
“Hi, I did a lot of reading online and was hoping you guys could help me out a bit more. It's technical stuff and I'm still learning, and a lot that I found is also outdated already. Perplexity says the upgrade is mainly worth it for the extra vram, but I was curious about your thoughts. I have an Aorus Elite X570 (1x PCIe 4.0 16x and 1x PCIe 4.0 4x), R9 5950X, RTX 5090, 64 GB DDR4-3200, 2TB NM990 SSD, 1200W PSU I use the pc as - ollama server with Qwen 3.6 to control an OpenClaw agent on”
“Hey guys Thanks in advance for your help and knowledge! My setup is born out of the parts I had at hand. Wanting to maximise VRAM with an RTX 4070 that I had in another system that I only used once or twice a year. So right now my system is 14600kf, 32gb RAM and rtx 5070ti 16gb on pcie 5.0 16x, and a rtx 4070 on a pcie 4.0 16x slot that runs over the chipset of my z790 motherboard. I know that my cpu has only 20 pcie lanes and that this is not great, for what I'm doing, which is why I'm”
“I am looking for a nice looking multi gpu case but can't find any good once.. Only this, any thoughts? It is a 6 gpu tower dual chamber https://www.alibaba.com/x/1lAq6Gz?ck=pdp submitted by /u/Timziito [link] [comments]”
“I wanted to do some local AI in a VM, so I bought an RTX 3090 and thought it would be possible to make a PCI passthrough. I have done that some years ago with an RTX 3060 and got it to pass through with full speed, so I thought that would be possible. So, the setup is an Alpine hypervisor with some VM's. I made a PCI passthrough from the hypervisor to a VM with Nobara Linux, which works, but only with gen 1 PCIe speeds. Hypervisor: Alpine Linux 6.18.2-lts, libvirt 11.10.0, QEMU 10.1.3 Guest:”
“There isn't much information around about multi-GPU setups with the R9700, so I'm writing this up in case it helps anyone in the same situation. Here's my setup, the tests I ran, and the numbers from the server logs. Setup ThinkStation P7, Xeon w7-3455, 128 GB RDIMM 2× Gigabyte Radeon AI PRO R9700 32 GB (64 GB VRAM total) Ubuntu 24.04 LTS, Docker 29.5.3, containers managed with Komodo (komo.do) ROCm 7.2.1 Image: llamacpp-rocm:gfx1201 Model: unsloth/Qwen3.6-27B-MTP-GGUF/Qwen3.6-27B-Q8”
“For context, I have a modest 32Gb rig running Nvidia GPUs (5070 Ti + 5060 Ti, the latter over an adapted x4 NVME slot so not as fast as if I had a motherboard with multiple proper CPU connected PCIe lanes). I can run the 27B models on it nicely enough, but the bottleneck is context. I’m a software engineer so I work on very large code bases and my sessions are often long, touching many components. I use Opus 4.8 almost exclusively, and that 1m context window means I can work efficiently. The rec”
“Hi, I'm having weird issues with my 3090 on inferencerence via lmstudio , it just: unloads the model/ model crashes + nvidia driver resets freezes the pc gives blue/black screen and the computer restarts or straight up restarts everything. I tried running it regularly, undervolted with afterburner, limited with nvidia-msi -pl (thought the issue was some power spike going beyond my PSU). It reduced the crashes, and their level, but still happens. During benchmarking, I see no issues (even tri”
“My employer has a GPU node that is mostly sitting idle. It contains 8 NVIDIA Quadro RTX 6000 GPUs with a total of 192 GB VRAM, and 512 GB RAM, and approximately 112 CPU threads to play with. I want to suggest we repurpose it for local inference. I need to make a case for this with my boss. What kind of models could I run on here that I couldn't do on a single card machine that would be worthwhile to use? submitted by /u/thehardsphere [link] [comments]”
“So i've got a 5090 and RTX PRO 6000 on a newer PCIE 5.0 motherboard with dual x8 splits and i want to add on another 5090 so i can hit 160gb total. I'm just running inference and not training, using a basic layer split and no tensor paralellism etc. I'd like to avoid investing in a server board if possible and i have two spare NVME slots and a huge power supply. I hear ADT Link makes pretty decent NVME to PCIE x16 adapters, for example: https://www.adt.link/product/F43V5.html I can&#”
“I have a gigabyte AI Top with 2 x8 slots but find myself wanting more PCI slots, so i'm speccing out a workstation board with a lot more bandwidth. Intel W790's stand out because the processors are dirt cheap compared to threadrippers and i plan to do zero CPU offloading. But it's a quad channel memory setup. DDR5 ECC 16GB is going for about $400 on ebay per stick.. just nuts. Is it possible to just run two sticks? if so, what price am i paying? ( desktop-like memory bandwidth? ) ”
“9 months ago, I set up an RTX 6000 + Threadripper Pro 9975WX, 256GB DDR5 RDIMM 6400, 3000W PSU (I planned ahead of time for today), all in a 9000D airflow chassis. The CPU is liquid-cooled using an AIO. I know, I know, the prices are crazy and so will my workload. I'm ready to get the next three RTX 6000's to load up. Plenty of airflow with eight(8) front ROG intake fans, the top has the CPU radiator and five(5) ROG fans, and two at the rear. Looking for community feedback on adding all”
“Hey guys, I need some advice on my current setup. I'm currently running an AMD 9900x, 64gb DDR5, and a 5070ti 16gb. I want to expand my VRAM for open-source LLMs and am thinking about adding another 16gb card (options: 5060ti, 9070, or 9070xt). My Gigabyte X870 EAGLE WIFI7 has one PCIe 5.0 x16 slot (already occupied) and two PCIe 3.0 x1 slots. Is it worth putting the second GPU in an x1 slot, or will it be a major bottleneck? Do I need to upgrade my motherboard to make this setup work effect”
“I want to build a headless compute machine to run a RTX Ada 4000 (20GB) with a RTX Pro 5000 (48GB) or RTX PRO 4500 (32GB) in parallel for inference. The goal is not running one large model using 2x GPUs, but rather running separate models on each GPU. Why these GPU config? because I already had a RTX Ada 4000 and don't want to sell it for now, but it's not enough to run larger models. This is going to be 95% time for inference and 5% occasional fine tuning LoRA / QLoRA type. This machine”
“Hello everyone, how are you all doing? I'm seriously thinking about building a Dual Xeon system in an open case, with 04 of these cards. Is it possible to do this without much trouble using llama.cpp? What are your opinions on these video cards? submitted by /u/Intelligent-Taste-36 [link] [comments]”
“I have the following old consumer GPUs in my house: 9070 XT 16gb 5700 XT 8gb GTX 1080 TI 12gb GTX 970 3.5gb R290 4GB (It's an older code GPU but it still checks out) I have access to the following PCIe slots: 2x PCIe 5.0 16x 4x M.2 to Oculink eGPU at PCIe 4.0 Technically also a 1x PCIe 3.0 I'm running an 9850x3d on a x870e AM5 board with 64gb ddr5 @ 6000 What's the best way to leverage the old hardware I have? submitted by /u/Vaguswarrior [link] [comments]”
