Running Multiple GPUs on Oddball Hardware Is a Mess
These users are grappling with the limitations of their existing hardware (motherboards, power supplies, PCIe slots) when trying to leverage multiple GPUs or higher-end GPUs for demanding AI workloads, particularly large language models (LLMs). They're weighing the performance trade-offs of dual-GPU setups versus single, more powerful cards, and struggling with VRAM constraints and PCIe bandwidth limitations.
SOURCES (60)
“I am currently using a Dell Precision 7730 I got second hand from my work, it runs hot and occasionally crashes. It's an i9 and 64gb of ram. I am thinking of putting my old 2600x in a spare motherboard I have and giving it…”
“Hello! I'm thinking of upgrading my board to ASUS ProArt X870E-CREATOR WiFi so I can effectively use my dual R9700 cards. Right now, my board does not properly support PCI-E 8x split and it causes ROCm to fail when using both cards (single GPU it's just fine). I know it's a long shot.. but anyone out there have this board and dual AMD cards? can be any type, so long as you're using ROCm for tensor split or tensor parallelism. Thanks! submitted by /u/Jorlen [link]  ”
“Hey all! I dont really know where to ask I thought this might be the place. I am an electrical engineer by trade, but I worked with FPGAs. What I learned there I was easy to adapt to work in CUDA. My FPGA work also allowed me to understand CPU architectures on a deep level, with perf and vtune in my toolbelt I am able to perf profile everything. But! This is the 3rd time that on an interview I got grilled on BigO. I understand the concept but I lack formal education/practice. Is there any reputa”
“X570 Taichi Specs AM4 X870 Taichi Specs AM5 I'm considering these boards as it seems that with the correct CPU, you can run 2 GPUs in x8/x8 PCIe. I use Linux full-time, I'm comfortable "trouble-shooting" or configuring. The largest models I'd consider running (for coding) are: - Qwen3.x-27b (Dense) - Qwen3.x-3xb-a3b (MoE) I'm sure I could use smaller models for other tasks and pleasant t/s speeds. Is it effective to have either MoBo and a combo like this? - 2 x 7900 XT”
“Prerequisites [x] I am running the latest code. Mention the version if possible as well. [x] I carefully followed the README.md. [x] I searched using keywords relevant to my issue to make sure that I am creating a new issue that is not already open (or closed). [x] I reviewed the Discussions, and have a new and useful enhancement to share. Feature Description On Intel Macs with more than one physical GPU (e.g. internal Radeon Pro plus an AMD eGPU via Thunderbolt), the Metal backend currently exp”
“Hey peeps, I am looking to make the jump to an open air case and would appreciate any recommendations on bifurcation adapters to convert x16 to x8/x8 and reliable riser cables please. For riser cables I have seen these which look decent. https://www.amazon.com/gp/aw/d/B0C415JCHX/ref=ox\_sc\_act\_title\_1?psc=1&th=1 For bifurcation, I've seen the c-payne stuff but they are rather expensive. I am looking to run 4 cards at x8 each so looking for 2x bifurcation adapters. Not sure if I need a”
“Im using a second PSU to power my watt hungry GPU. Haven't figured it yet how to synchronise both power units so that they start and stop at the same time. Im just doing manual startup turning the workstation and GPU about the same instant. So far it works OK but I do need to devise synchronisation tweak because it screws sleep - GPU wont turn off and starts spinning its fans at 120% when running alone”
“A 3090, not to mention the A6000, are on completely different price levels than 2x 3060. I'd suggest to increase the budget to have a 3090 only if OP wanted to use a diffusion model, but for qwen 2x 3060 is fine and it cost less than half of a 3090, at least in Europe.”
“I have a 5070 Ti and two 5060 Ti (all 16Gb cards). I planned to add another 5070 Ti giving me two pairs of 32Gb each but the NVidia prices have just jumped by 25% where I am and I've found a 3090 Founders Edition for a good chunk cheaper than the 5070 would cost. It's 8Gb more VRAM and even slightly higher memory bandwidth, but I've read that mixing Ampere with Blackwell comes with a performance hit in llama.cpp using tensor parallelism. I believe TP isn't possible at all in vLLM”
“Need some advice and guiding. I have a Dell 5820 workstation with an Intel Xeon w2245 and two RTX 5060 Ti (16GB) cards; this is a PCIe Gen3 system, but three lanes can work in 8x mode. The system runs Linux and mostly llama.cpp. My main work is done via Hermes Desktop. I mostly use NVFP models: esatapedico/Qwen3.8-27B-NVFP4-MTP-GGUF (Very High) and lately cdiamond/Qwen3.8-27B-iMatrix-NVFP4-MTP-GGUF. With this setup, I can use the models at 210k context, Q8 KV cache, and vision in VRAM. Ubatch do”
“Two RTX 3080 20G are about the price of one RTX 3090. Working well for me. 7900 XTX is also cheaper than 3090, but slower and not CUDA.”
“I see prices going up for any GPU even remotely related to AI at Microcenter (there are 5 in my general area). A R9700 that was $1300 a couple of months ago is now minimum $1700. An Arc B70 that was $1000 is now $1500.”
“Power delivery is the least challenging part(for GPUs), you have just a bunch of 12V lines with various connectors, you could power it even with an ordinary high current 12V power supply. You would just need to buy connectors with leads in bulk and make your own wires or a power rail(a single pair of thick copper conductor branching out close to each GPU. 3.3 and 5.0 V lines, often needed for various small things on the adapter electronics, should be generated on the board with buck converters,”
“I'm considering adding another A40 to my setup but my case doesn't have the space (or power) for it. It's a server chassis so there's no easy way to just pick the MB and put it in a new case, it's a bunch of custom connections off the PSU to the MB. What would be perfect? A dual PCI external box that would allow me to put in 2 A40's using nvlink with 1 or 2 connections coming out of that box running back to the server into the existing x16 slot. It's easy to find a si”
“I see a lot of people buy DGX Sparks, and turn them in to clusters to run large models. Wouldn't it be better to invest $16k into an AMD Epyc server with 768GB or even 384GB of 6000Mhz DDR5 ram, and let's say 2x3090s or 5080s, instead of 4 DGX Sparks with 512GB of ram? Epyc's theoretical bandwidth is around 576GB/s, DGX Spark's is roughly 273GB/s. Based on a quick check, both systems are worth around $16k. Please help me to understand this logic, are there benefits to having DGX”
“I got the CMP170 cards and unlocked them. I wanted to share my set up for CUDA since maybe it would be useful to others. First off, I hate e-waste and we are in a special time for RAM. I wanted to have a DIY CUDA box, and I had started by adding additional cards to an old asus predator prebuilt I had around, which also had 64gb DDR5. To add the CMPs I needed more CPU lanes and newegg had some really good deals on CPU/MB/etc combos. Didn’t need a combo with RAM, otherwise I would have gotten it i”
“I currently have the following rig cobbled together: MSI mpg z890 carbon Intel Ultra 7 64Gb DDR5 6000 RTX 5070 Ti (16Gb) 2 x RTX 5060 Ti (16Gb each) The 5070 and one of the 5060s are in the main CPU connected PCI slots (running at x8). The second 5060 is on a CPU connected M.2 slot via an M.2 to PCIe 4.0 x4 riser. I have one remaining CPU connected M.2 slot that I could use for a fourth GPU (also at PCIe 4.0 x4). My max budget is around $3,500 (£2,500 as I'm based in the UK). So, the options”
“Oh the Nic’s aren’t noisy at all, but they run hot asf if you’re running 100gbe or above with fiber cables vs, DAC. I literally have Noctua fans zip tied to every Nic. I run multiple GPUs per node too and the heat is a little much. I have water-blocks just haven’t gotten to the point of blocking the cards yet. Once I block the GPUs I’m hoping the extra room will help alleviate some of the heat dissipation from the cards hitting the Nic’s. The Microtik doesn’t give a shit if you’re running fiber”
“Oh my! Well, from faint 30-years-on recollection, it was the end of '96 so I guess I needed to save up for it. Thanks for the correction!Ordinary relatively cheap consumer Matrox cards could be paired up. I couldn't afford a fancy dual-output card, but I could afford 2 second-hand Matrox Millennium cards, and that did the job fine. At that time, Matrox was about the only brand of (parallel PCI) graphics card where the drivers could handle multiple cards. You couldn't do it with ATI or S3 or any”
“My setup is really a shitshow lol but I like to tinker so it keeps me on my toes. I have a 1000 psu as well, I cal the 5070ti amd 4070 and let the r9700 have ots full 300”
“Are you using M2 to PCIe x4 adapters already? May be better than 2 lanes per GPU.”
“So i spent considerable time trying to figure out why one of my eGPUs has degraded from x4 to x1 permanently. Yesterday while cleaning i found the culprit. Lesson learned: Don't buy eGPU risers that have HDMI connectors. I was going for Oculink connectors but the seller ripped us off. Before anyone asks: yes that's tinfoil separated by duct tape on the back of the pcb - it greatly helps EMI problems. If you can't see it: check the right HDMI connector. At least we have the means to r”
“TP. 1GPU Is on PCIE gen 5 16x, the second GPU is on m.2 to oculink (equivalent of PCIE gen 4 4x)”
“I think it’s built for 1 gpu setups. Not 2. Not many people rock two 5090’s. 5060/5070 more so. Ps I love my dual 5060’s. Great bang for buck.”
“will work fine, the issue is mounting it in the case(zip ties work for this.. but is janky), lots of adapters on ali express, just make sure the ribbon cable comes out at the angle you need it to”
“https://www.microcenter.com/product/711417/asrock-intel-arc-pro-b65-creator-single-fan-graphics-card My RTX 5070 isn't cutting it at 3 t/s for Qwen 3.8 27B, I'm still stuck using 35B with offloading. Obviously there are better options like strix, dgx, 5090, etc. but I can't pay over $1k for hardware right now. $650 for the b60 or $900 for the b65 is a little more reasonable, I have extra am4 parts so that would be the only cost other than a cheap case. I know intel arc is way behind”
“Need some expert assistance and advice: I know AM5 has pretty big limitations on PCIE lanes etc. What I wanted: Slot 1 - AMD R9700 32GB Slot 2 - AMD 9070 XT 16GB Slot 2 would run at PCIE 4x. That's understood, but still ~8GB/s allegedly. What I got was 0.1GB/s. According to Claude (and I already know it can go off on a non-existent tangent) - ASMedia's Promontory 21 bridge doesn't implement AtomicOp routing, and you have two of them daisy-chained. Nothing above it can fix that. Windo”
“Forgive the typos and rambling - human actually wrote this post 😂 I've had a Strix Halo board for about a year and been playing around with it for various projects when it's not just being a beefy linux machine. I ordered the 128GB Framework Desktop board pre-panic and I'm very grateful for that. I also grabbed an R9700 Pro AI card late last year for another machine, thinking it would be fun to compare the two. I ended up parting out the machine the R9700 was in for something else a”
“How does a compute pool work under the hood to give the user/application the illusion of a unified system? And if we hypothetically applied that model for consumer gpus , what additional work would it create for game developers to make a game run across multiple GPUs? My core assumption is that such a setup could potentially provide improved performance by combining both compute capacity and VRAM. Thanks 🙌 submitted by /u/Worried-Newspaper-65 [link] [comments]”
“I have one slot, a single PCIe 1x. The rest it’s GPU and NVMe card. What performance difference would there be considering it is a Ubuntu server with only terminal? The two GPUs are 5060TIs and one is video right now and there is not on board video. submitted by /u/Frizzy-MacDrizzle [link] [comments]”
“Hi Guys! I want to build a small LLM Server for my Homelab and I'm currently unsure what GPUs to buy for it. Currently planned are 4 Cards on a Threadripper Board, so all cards will get full PCIe 4.0 x16 - 64GB VRAM But whats the better choice here? Looking at the number, the 5060 Ti has 448 GB/s Bandwidth and the RX 9070 640 GB/s. Prices in Europe are quite the same, up/down 30€ between those two. Whats your input on this? Are there better options in the ~2600€ territory? submitted by”
“I expect these are literally 2 B60s put together on the same card? Each run at PCIe 5.0 x8 and show up as 2 different GPUs?”
“TL;DR Personal AI needs real GPU headroom for interaction, memory, and adaptation — that need does not shrink; it is structural. The mobile device has to remain the center of the experience because the camera, microphone, files, display, sensors, and user interaction live there. The best GPU cannot live inside that device because its power and weight make it nearby infrastructure, not handheld hardware. So the GPU has to move nearby — and a nearby GPU box only works if existing applications stil”
“That topo output already answers it, and it's bad news for TP=2. Link type between every pair is PCIE, 2 hops, weight 40 - no direct peer link, so every GPU-to-GPU transfer goes up to the root complex and back. With XGMI you'd see XGMI as the link type and a much lower weight. Numa Affinity -1 on all three also means the runtime has no NUMA pinning hint, so buffers can land on the wrong side. In rocminfo, check the Agent blocks: both GPUs should report the same gfx1201 and the same CU co”
“*on certain setups. Based on TinyGrad's open-gpu-kernel-modules but forked for more GPUs using https://github.com/aikitoria/open-gpu-kernel-modules/ Worked on my 2x3090 gpu setup with a ProArt Z790-CREATOR motherboard submitted by /u/luedtek [link] [comments]”
“Has anyone successfully patched the drivers of a modded RTX 3080 20gb to get Rebar support? I already updated my 3090. Rebar is required for P2P which would give me roughly +10-15% inference speeds, at least for vLLM. Anyways, I'm fairly certain P2P isn't possible for this card. A mixed GPU setup certainly doesn't help either. I'm curious to hear what others have done. I might stick to pipeline parallelization with MTP in llama.cpp. I get about 50 t/s with that Any input is welco”
“Recently heard about how these 8gb mining ewaste cards are actually closeted 64gb monsters and im kinda bummed I missed out when you could get them for 200$ a piece. That being said though less than 2K is still an amazing deal for an ampere generation gpu with 64gb of vram. I have a machine with 4 3090s in it currently. I guess what Im asking is for someone to convince me that im not missing out and if there is any genuine reason why it would not be worth it to replace 2 of the cards with cmps.”
“I have a gigabyte ds3h v2 b450 motherboard. Currently hosting an RTX 3090 I have a spare 3070 and I wondered, can I run both? Got myself a riser cable and… Top slot card, bottom slot riser = card pushes the riser, doesn’t fit Top slot riser, bottom slot card = card pushes the sata cables doesn’t fit… So I either need another riser and put both gpus out of the case, or an atx motherboard with more space in between the slots or maybe I should look for another solution. I’ve read some of you are us”
“can u share me the supplier you got from? also when u say pcie2.0 unlcoked is this the x16 mod with capacitors soldered?”
“I just saw the price of RXT 6000 Pro and the Thor IGX. while the price looks similar but the IGX you will get a full setup not only the GPU. Anyone looking into this? submitted by /u/ahstanin [link] [comments]”
“For bifurcation look for expansion cards that are compatible with mobo - eg ASUS says their mobo is compatible with the hyper m.2 or whatever it’s called - and you can change slot mode from x16 to „ASUS hyper m.2 card” and it basically changes the slot from x16 to x4x4x4x4 and you can plug whatever into it that works in that mode - same for other vendors - consumer mobos often don’t call it bifurcation literally but some vendor specific name”
“good point about the bifurcation. the manual doesnt specifically state what options are availabale, but i have seen it in the BIOS. and that was enough info to find this thread: https://www.reddit.com/r/LocalLLaMA/comments/1qim4ip/msi_mag_b650_tomahawk_pcie_lane_bifurcation/ which is exactly what i am trying to do.. like literally the same setup. so there is hope without a new mobo! nice. thank you for the reply, it is very helpful.”
“totally physical logistics discussion.. if there is a better sub, please let me know and i will repost.. but i see many using dual cards and want the details on the mobos used. my current B650 doesnt have slots that are far enough apart to put in 2 of these XTXs. so i thought SFF and oculink on other PCIe slot, then i see its thru chipset and its 3.0 and x2, total horseshit.. i want 48GB of juicy VRAM on the same connection type and speed (i.e. PCIe 4.0 x16/x8), but i need to sort the logistics.”
“Speaking of -sm tensor. I have 2x5060ti and i get some 40-50tps form 3.8 27B q6. I used HWinfo to see how saturated the PCIes are during inference are and as expected both the PCIe 5x16 slot and PCIe 4x4 were fully saturated. I cant help but feel like my 2nd gpu slot is a bottleneck (4x4), im considering getting a riser for my spare NVMe 5x4 slot and plugging the card there. I wanted to hear your experiences with PCIe bottlenecks and risers before comitting to any purchase. submitted by &#”
“I have this same machine with a 950w psu. I have one white female power inlet still open and i think i can connect a 10 pin to dual 8 pin connector to it. Is that what you did or how did u rig the power up? Having a heck of a time finding the 10 to dual 8 pin in the US, but there are some overseas i see. Just get nervous about the cables frying or starting a fire. Any tips would be great. TIA!”
“I had similar idea back in march when the hype around ANE was big and everyone was tried stuff yet this seemed impossible to me, as even with skipping required synchronization (results in corrupted outputs) the maximum PP I got on a 4B dense was slightly behind baseline at best and with synchronization wasn't even close. In this case they shard just part of the MLP and GDN, with all optimizations up claim to get ~50% better prefill rate on Qwen3.8 27B q4 on M3 Ultra. Other people including m”
“My about modded rtx series cards are butchered soldering jobs and low quality work in general. For a card bound to get hot I'd rather not risk it breaking down I found some ~250€ mi50s on 1688/Alibaba”
“Added 3rd GPU and now getting VGA error LED on motherboard. 99% seems riser fault cause it isn't working with any GPU in any slot even when one GPU connected but maybe anybody had something similar with Gigabyte B850 motherboards? PCIe 3.0 x16 30cm riser. Everything was powered on, photo was done before connecting power cable. Reordered 4.0 20cm. submitted by /u/esw123 [link] [comments]”
“I believe this part should be compatible. https://ebay.io/m/c9vmMA I'd double check with the seller. You can always return it if it ain't.”
“We are managing 5 physical hosts, each has two L40s NVIDIA cards. Using Proxmox as hypervisor and each host has one Ubuntu VM with GPU cards are passthrough. There are several LLMs are running on each card with all vLLM over Docker. The problem I'm facing is, each GPU cards VRAM utilization is around %90. So there are 5 GB VRAMs are sitting there freely. I wonder if anyone has a any elegance solution to this kind of infrastructure to make use of the free VRAM across several cards ? Because t”
“I think I start to throttle pretty hard around 95c, I start getting less than 1t/s. I think you are onto something with the thermal paste/pads. These are stock and probably have never been replaced.”
“If its overheating, its highly likely you need to apply fresh thermal paste. Its like 15-20C difference in some cases. What are your hot spot temps?”
“Ive heard it said a few times on here that miss-matching cards are not a good idea, but I just have to ask because the idea of throwing a 3090 into my existing gaming pc just seems like such an easy win for getting 40gb of vram. I'm pretty much content with the idea of sticking with models that will run well in that 40-48gb range. I am very budget minded to the point where even springing for the 3090 is a significant investment and while I will eventually like to find another 3090 and build”
“Hello All, if I combine 4070 and 3060 will it affect local performance since cards are 1 gen apart? Also should I calculation full powerdraw for power consumption? I currently have 750W Psu. Second question- How do I integrate it with opencode, should I run LM studio as server? Third- for those who have similar 24GB vram, what models are you guys using for generic coding? submitted by /u/HsSekhon [link] [comments]”
“I got them for aliexpress (SFF 8654 8i cables), boards from there. Power I powered them on with PCIe 6pin cables.”
“Prerequisites [x] I am running the latest code. Mention the version if possible as well. [x] I carefully followed the README.md. [x] I searched using keywords relevant to my issue to make sure that I am creating a new issue that is not already open (or closed). [x] I reviewed the Discussions, and have a new and useful enhancement to share. Feature Description As anyone who is probably trying to get the most of their hardware, I am trying to do the following thing: I want to load a model on my po”
“What is the issue? From Reddit , I've noticed that people were able to get the Radeon 680M iGPU working on Ollama but not in Ubuntu. I added the HSA parameter but it seems to load slower than it runs on the CPU. I've noticed that a person using a 6700M which is a graphics card but older than my iGPU works https://github.com/ollama/ollama/issues/6152. I am not sure what to do and how to fix this because people have been saying that it doubles the speed which isn't the case for me. I have restarte”
