PANE

LLM Performance Tuning & Optimization

Enthusiasts and hobbyists are struggling to achieve optimal performance when running large language models (LLMs) on their custom hardware setups. They face challenges like slow token generation speeds, compatibility issues with specific models and tools (like llama.cpp), and difficulty scaling performance across multiple GPUs.

devtoolsaillamaoptimizationperformance
FIT
0%
SIGNAL
82%
SOURCES60
FRESHEST POST7H AGO
TRACKED SINCE100D AGO

SOURCES (60)

I think you will be disappointed with the results. I spent 3 days trying to get qwen 2.5 coder 3b to behave with a really rigid harness. It just isn’t smart enough, it was struggling to call tools. If it were me, I’d target 7b…

r/LocalLLaMA7h ago

I acquired an older enterprise server a few years ago for a project that I never continued with. I'm wondering if it's worth investing in GPUs rather than building a brand-new AI machine. I was going to sell this, but I don't think it would go for very much, and thought adding a few GPUs would be a worthy investment. Hardware goes back to mid-2010s. Machine POSTs and memory has passed MEMTest. My primary goal is to run local LLMs for privacy and security. Ideally this machine would h

r/selfhosted1d ago
Source preview · reddit.com

Generation on an empty context is 23tps. Down to 14 tps at 30k.

reddit.com1d ago

Yeah I realised about M3 and made an edit to my comment! StepFun sounds interesting for creative and conversational stuff, I’m gonna have to bump it up my interest list! I’ve never tried a fine tune of any model. Seems people generally don’t think much of them. I will look into those ones you mention though. Any idea of why they’re supposedly better?

r/LocalLLaMA1d ago
Source preview · reddit.com

How's your tks

reddit.com1d ago
Source preview · reddit.com

MiniMax M3 is 426B from memory

reddit.com1d ago

Nice! I am definitely jealous of that GPU performance. I get about 23tps on 397B on an empty context. It’s good enough for my workflow generally, and it runs my OpenClaw. I tried Hermes but it’s soooo slow in comparison.

r/LocalLLaMA1d ago
Source preview · reddit.com

M3 is 427B, so pretty close to what you're running right now.

reddit.com1d ago

I'm not quite understanding but wouldn't the conversion also cost? I don't see how this is improvement to cost rather than just speed?

r/LocalLLaMA2d ago

some proof: https://i.imgur.com/qqEjxmW.png https://github.com/amoghmunikote/cmpunlocker CMP 170HX 8gb — Perf + Memory + PCIe Gen2 Unlock NVIDIA Driver 610.43.03 (patched open kernel modules) GPU: CMP 170HX (0x20C2) Result: 64 GB HBM2e + 173 TFLOPS BF16 + PCIe Gen2 x4 (2 GB/s) OS: Linux x86_64 (Fedora, Ubuntu, Debian, Arch, openSUSE) Kernel: 5.x — 7.x WHAT'S INSIDE ───────────── cmp170hx-unlock/ ├── install.sh — installer (run this) ├── NVIDIA-Linux-x86_64-610.43.03.run — NVIDIA userspace (l

r/LocalLLaMA2d ago

Running it is fine, it's deliverability that isn't

r/selfhosted2d ago

I have three Radeon AI PRO R9700s. I wanted to know if I could tune a quant for this specific hardware instead of just using the generic upstream formats. Came across the Github project for ROCmFPX by the wonderful Carlo Pasquale and started digging in. The fastest variant I've used is the MTP version. The format is Q8_0_ROCMFPX. My part was the gfx1201 decode tuning - VDR8 vector-dot width and a measured wave policy plus all the validaton. Each blok stores 32 signed int8 codes and one UE4M3

r/LocalLLaMA3d ago

disabling mmap helps too - although it takes way longer to load. also make sure it's the same quantity version i shared

r/LocalLLaMA3d ago

I have a Nimo Strix Halo system up and running, it's not the fastest, it can't run the biggest open weight models or high quants of 100-200b models, but it works really well for most of my use cases. I'm a hardware and signal/power integrity engineer. Mostly I need janky python scripts, some drivers for lab equipment, and some internal devops scripts for kicking off projects and managing my infra. It's good enough for that, and I'm able to develop nice internal tools for data

r/LocalLLaMA3d ago

I benchmarked 4 OpenAI GPT Models: GPT 5.5, 5.4, 5.4 mini, and 5.3 codex spark, against each other in a Doom Benchmark built with Codex. (5.6 model results coming soon!) Models control doom players through MCP tools, observe the game state, plan, strategize, and fight each other across multiple rounds. GPT-5.5 placed first with a 67% score. It collected 4x as many health packs as the next highest model. Won 80% of rounds where it secured the shotgun. Used resources to retreat, recover, and re-en

r/ChatGPT3d ago

Howdy folks, Long time reader here. This time, I would like to share my experience with this model that I have been cautiously optimistic about when I saw the announcement post on this thread. ----- My setup: AMD Ryzen 5 on AM5 chipset (I genuinely don't remember the exact CPU. Just pick whatever fit in the budget after the GPU) 32GB DDR5 running at 6000MT/s 4060Ti 16GB LLM Provider backend : PrismML fork of llamacpp running behind llama-swap. Harness: Pi agent with a set of custom made exte

r/LocalLLaMA3d ago

Not fable level but closer but could be a lot better. That shouldn't be so hard to do but it requires a lot of time and some money to rent the GPU hours. Currently the best 30b class model is sonnet 4.5 level which seems archaic and very incapable compared to current sota.

r/LocalLLaMA4d ago

So llama.cpp is faster despite all fuss. Expected. Yeah, the possibility excites me too. Being able to run it is always better not not able to, however slow it is

r/LocalLLaMA4d ago

Hi, I'm in a desperate situation. I've built a dedicated PC for local Ai usage (mostly Vision OCR for RAG info extraction) Specs: Ryzen 3 4100 4c/8t 8gb 3200MT/s ram RTX 5060 TI It works fine, but somehow llamacpp always trashed the RAM and then the whole PC just crashes since it's OOM Is there a way to stop it from doing that? This is my starting script. (Please don't trash on me, I only had a windows installer stick with me) set LLAMA_SERVER=Llamacpp\llama-server.exe set CHAT_M

r/LocalLLaMA4d ago

Nice to see gfx1030 still supported, hopefully AMD keeps supporting their hardware for longer

r/LocalLLaMA4d ago

I managed to get this model working on 2x 3090s with full 262k ctx and N=4, if anyone is interested to try it, thanks to this quant: https://huggingface.co/danielrmay/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-W4A16 Nemotron-Labs-3-Puzzle-75B-A9B (W4A16) on 2× RTX 3090 — vLLM, no CPU offload, full 262K, N=4 Hardware: 2× RTX 3090 (SM 8.6), PCIe Gen4 ×8, no NVLink, aikitoria BAR1-P2P patched driver active. Engine: vllm/vllm-openai:cu129-nightly dev1060 (9e57de71). Model: danielrmay/NVIDIA-Nemotron-Labs

r/LocalLLaMA4d ago
Source preview · reddit.com

Wanted to share a recent post I made regarding my success using OpenClaw with a local model. I am pretty much using 90% local, with my 5070 Ti card. I wanted to…

reddit.com5d ago

FP8 kv cache. 0.6 gpu max util gets you cca 500k context on dcp 1 and 0.75 1M. On dcp 2 twice as much, on dcp 4 4x. Upside down to not have the fans towards monitor/me. I can move it on the left side of the monitor instead of the p620 workstation maybe, even if there is less free space. For now, upside down doesn't bother me.

r/LocalLLaMA6d ago

Ctx Size: 48k Prompt processing: 27t/s Token/s: 4.31 t/s RAM Usage (on cold start): 6.2GB with model loaded (no-mmap) Yes it is not very fast but it is a good AI setup consuming roughly 25W and it can do complex tasks pretty well now. submitted by /u/BoogerheadCult [link] [comments]

r/LocalLLaMA6d ago
Source preview · reddit.com

submitted by /u/Dear-Economics-315 [link] [comments]

reddit.com6d ago
Source preview · reddit.com

What CPU are you running?

reddit.com7d ago

input8192-c1 round 1/5: aggregate=91.20 tok/s, per-request=134.15 tok/s input8192-c1 round 2/5: aggregate=95.16 tok/s, per-request=142.64 tok/s input8192-c1 round 3/5: aggregate=95.67 tok/s, per-request=143.83 tok/s input8192-c1 round 4/5: aggregate=95.84 tok/s, per-request=144.21 tok/s input8192-c1 round 5/5: aggregate=92.94 tok/s, per-request=138.39 tok/s input8192-c2 round 1/5: aggregate=40.00 tok/s, per-request=72.11 tok/s input8192-c2 round 2/5: aggregate=131.67 tok/s, per-request=106.57 to

r/LocalLLaMA7d ago

Disclosure: It is open source, Apache-2.0 licensed, and currently alpha. Repository: https://github.com/mireklzicar/gpuhedge I started working on it after benchmarking a 17 GB AI model across several serverless GPU providers. On the primary provider, requests usually either completed in roughly 6–8 seconds or took around 90–122 seconds after a fresh GPU cold start. Simply switching to another provider did not remove the problem because every provider had its own tail. GPUHedge treats this as a s

r/mlops7d ago
Source preview · reddit.com

The GPUs actually got really warm haha

reddit.com7d ago
Source preview · reddit.com

Whoa great work OP this is cool af

reddit.com7d ago

Oh wow almost double! What about tg?

r/LocalLLaMA7d ago

Xeons are better at just about all AI workloads, not just k_transformers and CPU inference. Anything the is heavily memory, I/O heavy bound (which AI is when running multiple GPU).

r/LocalLLaMA8d ago
Source preview · reddit.com

263kWh? How?

reddit.com8d ago

TL;DR: Added native nemotron_h_puzzle support to mlx-lm ( PR #1535 ), then compared 4-bit vs 5-bit expert quantization (both with 6-bit dense layers, BF16 output head, group size 64) on a 64GB M2 Max. Results (same prompts, 5 seeds per task, temp 1.0 / top_p 0.95): 4-bit experts 5-bit experts Dense paths 6-bit 6-bit Output head BF16 BF16 Checkpoint 42.03 GiB 49.88 GiB Peak MLX memory 49.68 GB 58.12 GB Average generation 14.27 tok/s 10.53 tok/s Local task checks 24/30 21/30 Long-context retrieval

r/LocalLLaMA8d ago

Im very happy with q5 on 16 gigs of vram. of course, its only partial offload but every little bit helps! And its even at a full sized context window.

r/LocalLLaMA9d ago

Recently I've found a deal to buy 9374f for cheap to replace my bottle-necked 9135. 8 CCD looked delicious. But first benchmarks showed me no decoding advantage. Until I used 48 threads. Non 64 or 32, which gave even worse performance than 9135 in some scenarios. Still not sure it was worth it as 9374f is much worse for gaming. Benchmarks (ik_llama.cpp latest version) with 4800 DDR5 for Unsloth GLM-5.2-UD-IQ4_XS: * 9135 PP TG N_KV T_PP s S_PP t/s T_TG s S_TG t/s 8192 128 0 31.835 257.33 14.7

r/LocalLLaMA9d ago

What you using for bifurcation? What’s your mobo and are you using something like oculink?

r/LocalLLaMA10d ago

Yeah but for me, in my initial testing, I was only seeing average 15tok/sec for DS4... So unless I had it setup badly wrong, this is still faster or equal overall.

r/LocalLLaMA10d ago

Just added a second 5060 16gb to my server for 32gb total VRAM + 80GB of ECC DDR4 It's pcie 3.0 so I think tensor parallel is not going to run well in any config but I get around 3200 tok/s prompt processing and 100 tok/s generation with qwen 3.6 35B. 27b runs at about 600 PP / 23 generation in tensor split. What settings should I be looking at to optimize my server now that I'm splitting weights across two cards? Currently running ggufs with llama.cpp but I have vllm nvfp4 models too. &

r/LocalLLaMA10d ago

but I don't see how moving to CPU instead of GPU changes anything... It changes the economics quite significantly, because CPUs cost much less. The trick is getting LLMs to perform well enough on such hardware so that training one feels more like tuning a large database rather than running a small city . Basically, the same way a child learns. This is an interesting approach and feels like the obvious way to learn (it works for us humans right?). Have you made any attempts to build a (neural

r/LocalLLaMA10d ago
Source preview · reddit.com

There's already CUDA support.

reddit.com10d ago

I agree with this premise of testing the model against the hardware, and doing it this way is a valid way of doing the model-for-the-hardware comparison, but it isn't exactly a valid way to disparage one model over another the way OP presented it. His title focuses on the model and not getting the best he can from his hardware...and worse he didn't actually say in his title he's comparing 4-bit to 8-bit so it feels a little...disingenuous in that regard. Additionally, OP either didn&

r/LocalLLaMA11d ago

Debian 13 llama.cpp --n-gpu-layers 999 --split-mode tensor -fit off --threads 4 \ --threads-batch 8 --batch-size 2048 --ubatch-size 512 --flash-attn on \ --cache-type-k bf16 --cache-type-v bf16 \ --no-mmap --mlock --ctx-size 393216 --checkpoint-min-step 4096 \ --ctx-checkpoints 64 --cache-ram 16384 --cache-reuse 256 --parallel 2 \ --cont-batching --spec-type draft-mtp --spec-draft-n-max 3 --jinja \ --chat-template-file /models/chat_template.jinja --temp 0.6 --top-p 0.95 \ --top-k 20 --min-p 0.0

r/LocalLLaMA11d ago

Wow... https://github.com/kacper-daftcode/vLLM-Moet Using this customized vllm provided as a docker, I'm able to run DS V4 Flash on a single RTX 6000 Pro (apparently it also works on a single 5090 - check his readme, but I haven't tried). Apparently this also works with GLM 5.2 (though you need at least two 6000 pros, which is still amazing). Setting 130K context, I needed around 150 GB of RAM to get past the safetensor sharding, but once it is fully loaded in VRAM I am able to fit it al

r/LocalLLaMA11d ago
Source preview · reddit.com

So a solid 5-8 t/s on MI50, very nice!

reddit.com11d ago

Yup - eight DGX Sparks (call it 40k including the switches & cables) would do it. Or an M3 Studio Ultra 512gb (so about 15k ish)

r/LocalLLaMA11d ago

Focus not only on token generation speeds, but also on prompt processing speed. In agentic use prompt processing is 95-99% of tokens in session.

r/LocalLLaMA11d ago

https://wccftech.com/intel-arc-pro-b70-beats-nvidia-rtx-5090d-in-deepseek-r1-ai-llm-over-2000-tokens-sec/

r/LocalLLaMA11d ago

I am running a local AI server on my home server that I also play media on. I need to leave a few GigaBytes of vram free for my hardware encoder to transcode media. I am coming from llama-swap where I could just put --fit-target 3072 and it will automatically offload GPU layers to system ram once it reaches 3GB of vram free, however LocalAI does not have that option. I have looked through the documentation and online, but could not find any settings I could use except for an unreliable source th

r/LocalLLaMA11d ago

I was curious whether any of the FA-3/4 optimizations transfer to RTX GPUs. vLLM/SGLang attention falls back to FA-2 on consumer cards (FA-3 and FA-4 are datacenter-only), so I wanted to know if there's any performance left on the table, and I rebuilt the attention kernels from scratch. The kernel reaches parity with FA-2 (206us on RTX5090 with batch=1, heads=8, seq_len=4096, head_dim=64), but unfortunately, FA-3/4 optimizations are either not applicable or not helpful on consumer cards. It

r/LocalLLaMA11d ago

Sure but it's kind of a moving target: - you got people with 6,8,12,16,24 and then multiple cards: each takes different quants / ctx - you got people with GPU and people with unified arch / CPU only, I guess someone has even both! - on the spot I can figure 3-5 different usage patterns, from coding, conversation, image, speak, video -- each of those usage may have 2-3 config from one quick shot to long ctx, maybe fast, hi precision, hi concurrency...

r/LocalLLaMA11d ago

Hello. So, first, ill probably explain the context, and why we (well, I kinda have a team, so ill say we) want to build something like that. We live in Russia. There are many preconditions, that point to the fact that pretty soon, we'll get our internet access (global internet, apparently) completely cut off. And we are preparing to go fully offline for indefinite period. I know, this sounds weird. No one believes that. But I prefer to be ready for the worst case scenario. ...and we kinda wa

r/LocalLLaMA11d ago

Thanks for the detailed response. I am getting about 30 t/s single concurrent with Qwen3.6 27B FP16 262K context. Someone else on the sub seems to be able to reach 67 t/s, but I'm absolutely struggling to reproduce that on my system. I'll have to try out llama.cpp if the performance is so close.

r/LocalLLaMA11d ago

If you use DCP2 or 4, you can fit a decent amount of ctx without having to resort to 4-bit KV cache, with about 20-50% prefill speed loss depending on which way you go. Judging by the benchmark results here, I think it's a good trade.

r/LocalLLaMA12d ago

which cpu? shiit 8 channels ddr4 ain't making any difference over my poor dual channel ddr5 it seems. it should be twice as fast

r/LocalLLaMA12d ago

Idk if this is the right place to ask but ill try anyway. I use a 24b q4ks at 12k ctx My specs are: Rtx 2070, 12gb of ram (8+4), i5 7400. I get around 1.70tks using this regex (blk.(?:[2-9]|[1-3][0-9]).ffn_up|blk.(?:[2-9]|[1-3][0-9]).ffn_gate|blk.(?:[1][2-9]).ffn_down)=CPU using around 7.5gb of vram I recently switched to q4xs instead and also dropped my kv cache to q8 and batch from 512 to 256. My vram usaged dropped to 6.7gb while getting 16k ctx I am unfamiliar with the tensor/regex stuff so

r/LocalLLaMA12d ago

no, they literally said themselves they intentionally sacrifice absolute intelligence for computational and efficiency and deployment throughput

r/LocalLLaMA13d ago

8K/1K is a pointless benchmark, only useful because it’s consistent. In the real world, you are not going to fit the KV cache for a sufficiently high concurrency on the leftover VRAM of a 4xB200.

r/LocalLLaMA13d ago

I updated the stream but was UNABLE to find any B200 rigs to test. I believe one of my suppliers has a B200x8 rig available for me now, so I can test. I know Fester and Luke tested with our blended fork and Luke's NVFP4 was once again the best NVFP4 accuracy wise. I'll take a look tonight/tomorrow on running some KLDs. Side Note: GLM5.2 was not QAT'd yet, so as far as we know, what they released is best effort Quants (FP8 - W8A8) and BF16. If they released a QAT of it, would be amazi

r/LocalLLaMA13d ago

So I've had one Gemma 4 E2B running through llama-server as the only model in a local tool that watches my screen and lets me search/chat over it later. Same model does all three jobs: - looks at the screen and turns it into structured info (what app, what I'm doing, rough layout) - audio — voice memos + meeting transcription using E2B's audio encoder, so I didn't have to bolt on Whisper - the actual chat/RAG over history(you can try chatting with a rag built over your screen his

r/LocalLLaMA13d ago

SOLUTION LANDSCAPE

Brought to you byTop Sectors

A Player feature.See how many ways this pain can be solved, who's already building, and where the gaps are.