LLM Performance Tuning & Optimization
Enthusiasts and hobbyists are struggling to achieve optimal performance when running large language models (LLMs) on their custom hardware setups. They face challenges like slow token generation speeds, compatibility issues with specific models and tools (like llama.cpp), and difficulty scaling performance across multiple GPUs.
SOURCES (60)
“I think you will be disappointed with the results. I spent 3 days trying to get qwen 2.5 coder 3b to behave with a really rigid harness. It just isn’t smart enough, it was struggling to call tools. If it were me, I’d target 7b…”
“I acquired an older enterprise server a few years ago for a project that I never continued with. I'm wondering if it's worth investing in GPUs rather than building a brand-new AI machine. I was going to sell this, but I don't think it would go for very much, and thought adding a few GPUs would be a worthy investment. Hardware goes back to mid-2010s. Machine POSTs and memory has passed MEMTest. My primary goal is to run local LLMs for privacy and security. Ideally this machine would h”
Generation on an empty context is 23tps. Down to 14 tps at 30k.
“Yeah I realised about M3 and made an edit to my comment! StepFun sounds interesting for creative and conversational stuff, I’m gonna have to bump it up my interest list! I’ve never tried a fine tune of any model. Seems people generally don’t think much of them. I will look into those ones you mention though. Any idea of why they’re supposedly better?”
“Nice! I am definitely jealous of that GPU performance. I get about 23tps on 397B on an empty context. It’s good enough for my workflow generally, and it runs my OpenClaw. I tried Hermes but it’s soooo slow in comparison.”
M3 is 427B, so pretty close to what you're running right now.
“I'm not quite understanding but wouldn't the conversion also cost? I don't see how this is improvement to cost rather than just speed?”
“some proof: https://i.imgur.com/qqEjxmW.png https://github.com/amoghmunikote/cmpunlocker CMP 170HX 8gb — Perf + Memory + PCIe Gen2 Unlock NVIDIA Driver 610.43.03 (patched open kernel modules) GPU: CMP 170HX (0x20C2) Result: 64 GB HBM2e + 173 TFLOPS BF16 + PCIe Gen2 x4 (2 GB/s) OS: Linux x86_64 (Fedora, Ubuntu, Debian, Arch, openSUSE) Kernel: 5.x — 7.x WHAT'S INSIDE ───────────── cmp170hx-unlock/ ├── install.sh — installer (run this) ├── NVIDIA-Linux-x86_64-610.43.03.run — NVIDIA userspace (l”
“I have three Radeon AI PRO R9700s. I wanted to know if I could tune a quant for this specific hardware instead of just using the generic upstream formats. Came across the Github project for ROCmFPX by the wonderful Carlo Pasquale and started digging in. The fastest variant I've used is the MTP version. The format is Q8_0_ROCMFPX. My part was the gfx1201 decode tuning - VDR8 vector-dot width and a measured wave policy plus all the validaton. Each blok stores 32 signed int8 codes and one UE4M3”
“disabling mmap helps too - although it takes way longer to load. also make sure it's the same quantity version i shared”
“I have a Nimo Strix Halo system up and running, it's not the fastest, it can't run the biggest open weight models or high quants of 100-200b models, but it works really well for most of my use cases. I'm a hardware and signal/power integrity engineer. Mostly I need janky python scripts, some drivers for lab equipment, and some internal devops scripts for kicking off projects and managing my infra. It's good enough for that, and I'm able to develop nice internal tools for data”
“I benchmarked 4 OpenAI GPT Models: GPT 5.5, 5.4, 5.4 mini, and 5.3 codex spark, against each other in a Doom Benchmark built with Codex. (5.6 model results coming soon!) Models control doom players through MCP tools, observe the game state, plan, strategize, and fight each other across multiple rounds. GPT-5.5 placed first with a 67% score. It collected 4x as many health packs as the next highest model. Won 80% of rounds where it secured the shotgun. Used resources to retreat, recover, and re-en”
“Howdy folks, Long time reader here. This time, I would like to share my experience with this model that I have been cautiously optimistic about when I saw the announcement post on this thread. ----- My setup: AMD Ryzen 5 on AM5 chipset (I genuinely don't remember the exact CPU. Just pick whatever fit in the budget after the GPU) 32GB DDR5 running at 6000MT/s 4060Ti 16GB LLM Provider backend : PrismML fork of llamacpp running behind llama-swap. Harness: Pi agent with a set of custom made exte”
“Not fable level but closer but could be a lot better. That shouldn't be so hard to do but it requires a lot of time and some money to rent the GPU hours. Currently the best 30b class model is sonnet 4.5 level which seems archaic and very incapable compared to current sota.”
“So llama.cpp is faster despite all fuss. Expected. Yeah, the possibility excites me too. Being able to run it is always better not not able to, however slow it is”
“Hi, I'm in a desperate situation. I've built a dedicated PC for local Ai usage (mostly Vision OCR for RAG info extraction) Specs: Ryzen 3 4100 4c/8t 8gb 3200MT/s ram RTX 5060 TI It works fine, but somehow llamacpp always trashed the RAM and then the whole PC just crashes since it's OOM Is there a way to stop it from doing that? This is my starting script. (Please don't trash on me, I only had a windows installer stick with me) set LLAMA_SERVER=Llamacpp\llama-server.exe set CHAT_M”
“Nice to see gfx1030 still supported, hopefully AMD keeps supporting their hardware for longer”
“I managed to get this model working on 2x 3090s with full 262k ctx and N=4, if anyone is interested to try it, thanks to this quant: https://huggingface.co/danielrmay/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-W4A16 Nemotron-Labs-3-Puzzle-75B-A9B (W4A16) on 2× RTX 3090 — vLLM, no CPU offload, full 262K, N=4 Hardware: 2× RTX 3090 (SM 8.6), PCIe Gen4 ×8, no NVLink, aikitoria BAR1-P2P patched driver active. Engine: vllm/vllm-openai:cu129-nightly dev1060 (9e57de71). Model: danielrmay/NVIDIA-Nemotron-Labs”
Wanted to share a recent post I made regarding my success using OpenClaw with a local model. I am pretty much using 90% local, with my 5070 Ti card. I wanted to…
“FP8 kv cache. 0.6 gpu max util gets you cca 500k context on dcp 1 and 0.75 1M. On dcp 2 twice as much, on dcp 4 4x. Upside down to not have the fans towards monitor/me. I can move it on the left side of the monitor instead of the p620 workstation maybe, even if there is less free space. For now, upside down doesn't bother me.”
“Ctx Size: 48k Prompt processing: 27t/s Token/s: 4.31 t/s RAM Usage (on cold start): 6.2GB with model loaded (no-mmap) Yes it is not very fast but it is a good AI setup consuming roughly 25W and it can do complex tasks pretty well now. submitted by /u/BoogerheadCult [link] [comments]”
“input8192-c1 round 1/5: aggregate=91.20 tok/s, per-request=134.15 tok/s input8192-c1 round 2/5: aggregate=95.16 tok/s, per-request=142.64 tok/s input8192-c1 round 3/5: aggregate=95.67 tok/s, per-request=143.83 tok/s input8192-c1 round 4/5: aggregate=95.84 tok/s, per-request=144.21 tok/s input8192-c1 round 5/5: aggregate=92.94 tok/s, per-request=138.39 tok/s input8192-c2 round 1/5: aggregate=40.00 tok/s, per-request=72.11 tok/s input8192-c2 round 2/5: aggregate=131.67 tok/s, per-request=106.57 to”
“Disclosure: It is open source, Apache-2.0 licensed, and currently alpha. Repository: https://github.com/mireklzicar/gpuhedge I started working on it after benchmarking a 17 GB AI model across several serverless GPU providers. On the primary provider, requests usually either completed in roughly 6–8 seconds or took around 90–122 seconds after a fresh GPU cold start. Simply switching to another provider did not remove the problem because every provider had its own tail. GPUHedge treats this as a s”
“Xeons are better at just about all AI workloads, not just k_transformers and CPU inference. Anything the is heavily memory, I/O heavy bound (which AI is when running multiple GPU).”
“TL;DR: Added native nemotron_h_puzzle support to mlx-lm ( PR #1535 ), then compared 4-bit vs 5-bit expert quantization (both with 6-bit dense layers, BF16 output head, group size 64) on a 64GB M2 Max. Results (same prompts, 5 seeds per task, temp 1.0 / top_p 0.95): 4-bit experts 5-bit experts Dense paths 6-bit 6-bit Output head BF16 BF16 Checkpoint 42.03 GiB 49.88 GiB Peak MLX memory 49.68 GB 58.12 GB Average generation 14.27 tok/s 10.53 tok/s Local task checks 24/30 21/30 Long-context retrieval”
“Im very happy with q5 on 16 gigs of vram. of course, its only partial offload but every little bit helps! And its even at a full sized context window.”
“Recently I've found a deal to buy 9374f for cheap to replace my bottle-necked 9135. 8 CCD looked delicious. But first benchmarks showed me no decoding advantage. Until I used 48 threads. Non 64 or 32, which gave even worse performance than 9135 in some scenarios. Still not sure it was worth it as 9374f is much worse for gaming. Benchmarks (ik_llama.cpp latest version) with 4800 DDR5 for Unsloth GLM-5.2-UD-IQ4_XS: * 9135 PP TG N_KV T_PP s S_PP t/s T_TG s S_TG t/s 8192 128 0 31.835 257.33 14.7”
“What you using for bifurcation? What’s your mobo and are you using something like oculink?”
“Yeah but for me, in my initial testing, I was only seeing average 15tok/sec for DS4... So unless I had it setup badly wrong, this is still faster or equal overall.”
“Just added a second 5060 16gb to my server for 32gb total VRAM + 80GB of ECC DDR4 It's pcie 3.0 so I think tensor parallel is not going to run well in any config but I get around 3200 tok/s prompt processing and 100 tok/s generation with qwen 3.6 35B. 27b runs at about 600 PP / 23 generation in tensor split. What settings should I be looking at to optimize my server now that I'm splitting weights across two cards? Currently running ggufs with llama.cpp but I have vllm nvfp4 models too. &”
“but I don't see how moving to CPU instead of GPU changes anything... It changes the economics quite significantly, because CPUs cost much less. The trick is getting LLMs to perform well enough on such hardware so that training one feels more like tuning a large database rather than running a small city . Basically, the same way a child learns. This is an interesting approach and feels like the obvious way to learn (it works for us humans right?). Have you made any attempts to build a (neural”
“I agree with this premise of testing the model against the hardware, and doing it this way is a valid way of doing the model-for-the-hardware comparison, but it isn't exactly a valid way to disparage one model over another the way OP presented it. His title focuses on the model and not getting the best he can from his hardware...and worse he didn't actually say in his title he's comparing 4-bit to 8-bit so it feels a little...disingenuous in that regard. Additionally, OP either didn&”
“Debian 13 llama.cpp --n-gpu-layers 999 --split-mode tensor -fit off --threads 4 \ --threads-batch 8 --batch-size 2048 --ubatch-size 512 --flash-attn on \ --cache-type-k bf16 --cache-type-v bf16 \ --no-mmap --mlock --ctx-size 393216 --checkpoint-min-step 4096 \ --ctx-checkpoints 64 --cache-ram 16384 --cache-reuse 256 --parallel 2 \ --cont-batching --spec-type draft-mtp --spec-draft-n-max 3 --jinja \ --chat-template-file /models/chat_template.jinja --temp 0.6 --top-p 0.95 \ --top-k 20 --min-p 0.0”
“Wow... https://github.com/kacper-daftcode/vLLM-Moet Using this customized vllm provided as a docker, I'm able to run DS V4 Flash on a single RTX 6000 Pro (apparently it also works on a single 5090 - check his readme, but I haven't tried). Apparently this also works with GLM 5.2 (though you need at least two 6000 pros, which is still amazing). Setting 130K context, I needed around 150 GB of RAM to get past the safetensor sharding, but once it is fully loaded in VRAM I am able to fit it al”
“Yup - eight DGX Sparks (call it 40k including the switches & cables) would do it. Or an M3 Studio Ultra 512gb (so about 15k ish)”
“Focus not only on token generation speeds, but also on prompt processing speed. In agentic use prompt processing is 95-99% of tokens in session.”
“https://wccftech.com/intel-arc-pro-b70-beats-nvidia-rtx-5090d-in-deepseek-r1-ai-llm-over-2000-tokens-sec/”
“I am running a local AI server on my home server that I also play media on. I need to leave a few GigaBytes of vram free for my hardware encoder to transcode media. I am coming from llama-swap where I could just put --fit-target 3072 and it will automatically offload GPU layers to system ram once it reaches 3GB of vram free, however LocalAI does not have that option. I have looked through the documentation and online, but could not find any settings I could use except for an unreliable source th”
“I was curious whether any of the FA-3/4 optimizations transfer to RTX GPUs. vLLM/SGLang attention falls back to FA-2 on consumer cards (FA-3 and FA-4 are datacenter-only), so I wanted to know if there's any performance left on the table, and I rebuilt the attention kernels from scratch. The kernel reaches parity with FA-2 (206us on RTX5090 with batch=1, heads=8, seq_len=4096, head_dim=64), but unfortunately, FA-3/4 optimizations are either not applicable or not helpful on consumer cards. It”
“Sure but it's kind of a moving target: - you got people with 6,8,12,16,24 and then multiple cards: each takes different quants / ctx - you got people with GPU and people with unified arch / CPU only, I guess someone has even both! - on the spot I can figure 3-5 different usage patterns, from coding, conversation, image, speak, video -- each of those usage may have 2-3 config from one quick shot to long ctx, maybe fast, hi precision, hi concurrency...”
“Hello. So, first, ill probably explain the context, and why we (well, I kinda have a team, so ill say we) want to build something like that. We live in Russia. There are many preconditions, that point to the fact that pretty soon, we'll get our internet access (global internet, apparently) completely cut off. And we are preparing to go fully offline for indefinite period. I know, this sounds weird. No one believes that. But I prefer to be ready for the worst case scenario. ...and we kinda wa”
“Thanks for the detailed response. I am getting about 30 t/s single concurrent with Qwen3.6 27B FP16 262K context. Someone else on the sub seems to be able to reach 67 t/s, but I'm absolutely struggling to reproduce that on my system. I'll have to try out llama.cpp if the performance is so close.”
“If you use DCP2 or 4, you can fit a decent amount of ctx without having to resort to 4-bit KV cache, with about 20-50% prefill speed loss depending on which way you go. Judging by the benchmark results here, I think it's a good trade.”
“which cpu? shiit 8 channels ddr4 ain't making any difference over my poor dual channel ddr5 it seems. it should be twice as fast”
“Idk if this is the right place to ask but ill try anyway. I use a 24b q4ks at 12k ctx My specs are: Rtx 2070, 12gb of ram (8+4), i5 7400. I get around 1.70tks using this regex (blk.(?:[2-9]|[1-3][0-9]).ffn_up|blk.(?:[2-9]|[1-3][0-9]).ffn_gate|blk.(?:[1][2-9]).ffn_down)=CPU using around 7.5gb of vram I recently switched to q4xs instead and also dropped my kv cache to q8 and batch from 512 to 256. My vram usaged dropped to 6.7gb while getting 16k ctx I am unfamiliar with the tensor/regex stuff so”
“no, they literally said themselves they intentionally sacrifice absolute intelligence for computational and efficiency and deployment throughput”
“8K/1K is a pointless benchmark, only useful because it’s consistent. In the real world, you are not going to fit the KV cache for a sufficiently high concurrency on the leftover VRAM of a 4xB200.”
“I updated the stream but was UNABLE to find any B200 rigs to test. I believe one of my suppliers has a B200x8 rig available for me now, so I can test. I know Fester and Luke tested with our blended fork and Luke's NVFP4 was once again the best NVFP4 accuracy wise. I'll take a look tonight/tomorrow on running some KLDs. Side Note: GLM5.2 was not QAT'd yet, so as far as we know, what they released is best effort Quants (FP8 - W8A8) and BF16. If they released a QAT of it, would be amazi”
“So I've had one Gemma 4 E2B running through llama-server as the only model in a local tool that watches my screen and lets me search/chat over it later. Same model does all three jobs: - looks at the screen and turns it into structured info (what app, what I'm doing, rough layout) - audio — voice memos + meeting transcription using E2B's audio encoder, so I didn't have to bolt on Whisper - the actual chat/RAG over history(you can try chatting with a rag built over your screen his”
