AI Model Users Struggle with Practicality
Users are experiencing a disconnect between benchmark results and real-world performance of AI models, particularly Qwen. They face challenges with operational realities, model stability, bias mitigation, and suitability for specific tasks, despite promising initial scores. This highlights a gap between theoretical performance and practical usability.
SOURCES (60)
“No it's possible because of html/css have a ton of training data. If you look at the thinking trace of the model do you see it checking and correcting its work? Or is it just generating html/css output tokens first try”
“Until such time that Qwen3.8 weights are published, this is off-topic for LocalLLaMA.”
“Until such time that Qwen3.8 weights are published, this is off-topic for LocalLLaMA.”
“this is good for them, not bad. they can train and run those models internally on top of data only they have access to.”
“There's very little that Harvey will provide over Claude. One key thing to ask, though, is: what is the large language model reasoning from? And that's where products like Harvey can be differentiated, as it's able to pull in the LexisNexis content. For case law citations, this is a critical difference. Even the Claude-based legal plugins can't touch this and never will as they don't own the data.”
“I haven't touch back since a couple months ago. Local model were so trash to use with and qwen wasn't even remembering anything at all.”
“I'm in awe... I recently incorporated Gemma 4 E4B into my workflows, and it's a godsend at 200k context. I can't imagine how a similar model with 512K ctx would powerlift... it would be like a CLI on steroids.”
“Exactly, that’s why I’m not against pandering to / placating those arguing for a ban by some lab putting out Western vetsions. I’ll let you teach the model whatever history you want, I’m using it for math questions and don’t care, just let me have the model (assuming they don’t lobotomize it in the process of re-educating it of course)”
“I think the jury is still out on K3/Q3.8 and if they are equivalent to Opus or Fable. Benchmarks have gotten incredibly murky and I've tried models that are "the same as X model" and been unimpressed. I've tried other models with Claude Code and with tools like OpenCode or Pi and nothing has really come close to Claude Code using Anthropic's models (mostly Opus 4.8).I'm not saying these other models are trash, just that I'm not quite ready to put them on equal footing to Anthropic or OpenAI mode”
“Yeah, its fascinating how much world knowlege is baked into modern models, but at home with limited gpu memory i am not sure i want to waste those precious weigth on the knowledge of south korean soap operas or american county statistics... if you could just point it to an offline download of wikipedia isntead (which is readily available)”
They did that with American cars and handed the market to Japanese cars.
“As further proof, in the case of Ornith, the Unsloth team just posted this: 'We added some Qwen chat template fixes to alleviate some issues y'all had with looping etc with the models and hope you give them a try." cit. https://www.reddit.com/r/unsloth/comments/1v15ob4/ornith_unsloth_ggufs_out_now/”
“If the US wants to handicap itself vs. the rest of the world, who will continue to build innovative solutions on top of Chinese open-weight models, they're welcome to do so, but I really wouldn't recommend it.”
“yeah even the EU is unlikely to pass on chinese models now that US models have been restricted outisde of the us. good thing to show the us is willing to do that now rather than later, maybe that will finally make the EU do ANYTHING AT ALL in terms of soverign AI.”
“I had the same experience with Claude the other day. I asked it for something fairly simple and it used up my whole 5 hour quota without producing the answer. Felt like a total waste of time.”
“lol the pelican benchmark is basically the only review process I trust at this point. kinda wild that alibaba of all companies is making it hard to give them money tho, you'd think they'd want prominent devs testing their stuff.openrouter usually picks these up pretty fast, hopefully it shows up there soon. You can also just download the qwen app and do this in the chat interface using their MCP tools for local dev.”
Why is "Reasoning Effort" only available for models matching gpt-5*, o1, and o3+
“The dense models are definitely stronger, but sometimes speed is a desirable quality!”
“I disagree, I ran that same gemma model a couple nights ago and it spiraled into a panic attack after accidentally deleting something and creating a stub. It actually started repeating the phrase "I can recreate the file!" endlessly. Something similar happened to me with some other gemma model I tried via GitHub chat and it outright deleted a file. Sure it looked busy, and did something simple correctly, eventually, at first, but in the next session it exhausted the context window and”
“what is Heretic? is it a tool/cli? oh found a thread about it https://www.reddit.com/r/LocalLLaMA/comments/1oymku1/heretic_fully_automatic_censorship_removal_for/”
“The problem is hardware availability. As long as these hyper scalers are gobbling up captive supply it’s going to be years before anyone can build a machine to handle billions of parameter models even on premise in small companies let alone home use. Feels like we’re headed back to accelerator cards like in the early 3d graphics era.”
“too bad KIMI does not support cache markers like many models from Anthropic ... so for many API use cases, KIMI is more expensive then Opus/etc.”
For simple tasks, using a smaller model is more efficient.
“You have no idea how the world works my friend. Zoom out a little. You also need consumers to justify the level of valuations they have. If all consumers shift to chinese models. You are left with no consumers. How do you justify your valuations then. Stop outsourcing your thinking to AI models.”
“Why the fuck would I use something slightly better than Sonnet!? Are you in the habit of adopting the Rank #20 product in the world for everything?”
“K3 worked very well in my tests with spreadsheet-mcp but did keep reverting to openpyxl scripting occasionally, even with explicit instructions”
“Personally - I always grab all popular local models and run them on a variety of hardware. It’s hobby plus career learning. Approximately 15tb of models. Then there media and datasets. Yes I’m a hoarder, I am aware.”
“I’m doing work related to energy and fault detection which is the chunk of the workload. For me the science and math reasoning benchmarks are usually a plus. Low level work from my fault engine is parsed and analyzed by them for results. Structured flow for the most part in my custom harness. Moving this to L40S in the cloud for production. Runs the same locally.”
“Fable is one of the more token inefficient models out there, Kimi is actually pretty light on tokens compared to Fable.”
“Yeah but if the washer doesn't wash as advertised it's a warranty issue. Listen I've been working with language models and studying black box lanecy objective since Mitsuku. If these new LLM are advertised as a data analyst on large lake with humanistic processing power, then it doesn't need to preform like a for entertainment chat bot. Which is really missing the overall of the conversation I posted.”
Nothing tells you evals are useless quite like nextjs having their own
“You could even do some kind of q2 thing and have enough space left over for context, maybe?”
“The simple fact that Gemini 3 Flash “Preview” is beating Claude Sonnet 5 in this “benchmark“, shows that it’s a useless benchmark.”
“Kimi K3 is clearly fable and chatgpt 5.6 sol level and even beat them in some benchmarks. So China is not 6 months behind anymore. And mythos is more of a myth atp.”
“I guess once they hit diminishing returns on model performance then they start building them into hardware cards like this for insane inference speed”
“Of course you are right in a free and interference-free environment, but are we (or will we be) in that environment? Yes, I'm talking about regulations, prohibitions and so on. The models are getting smaller and more effective, but before there was no "dangerous" (and fictitious) cap”
“I don't know about 27B with Fable level capabilities in 5 months seems too quick. Maybe 120B. If I had to guess on when a 27B Fable level would arrive i'd say 2 years.”
“I’ve always had this idea but can’t prove it. I think Anthropic and OpenAI don’t really have any secret sauce, their moat is just scale. Rumor has it Opus has 5T parameters and Mythos/Fable are 10T parameter models, while open models stayed under 1T for a long time. Only recently was that ceiling broken by DeepSeek V4 and now Kimi K3, and we’ve seen a significant jump in performance as parameter size increased. What do you think? submitted by /u/a9udn9u [link] [comments]”
“local models sound great until you try running them on a normal laptop. they'll smell like a hot dryer and run like cold molasses.”
“Mostly just here to follow this topic 👍 As of today, I feel like Qwen3-Embedding-8B and Qwen3-Reranker-8B provide a RAG capability that is genuinely superior (and dramatically cheaper) than anything available through online/frontier models, but this only requires 16gb of VARM. Otherwise, I feel like the reasoning capabilities of frontier models have gotten so far ahead that local will never catch up. If I were to speculate, I’d guess that in 2 years the most valuable local architecture will be”
“Unbelievable to see kimi k3 beat frontier models that were 'too dangerous' for public use. submitted by /u/Gohab2001 [link] [comments]”
“Yeah, but we will want to convert expert weights to IQ2_XSS or others, so QAT effectiveness might be reduced for local users.”
“If you look closely you will see: Qwen 3.6-plus and Qwen-3.7-plus have almost no similarity to Antrophic models. Only Qwen-3.7-max shows similarity to Claude Sonnet 4.6 and Claude Haiku 4.5 - which are lower-grade Antrophic LLMs.”
“I am using the full model released size to compare, because this is the size used for benchmarks. I do not understand the downvotes for my message and yours (you just give facts from the docs).”
“I am comparing the full precision model, as released. This is the model size used for benchmarks.”
“Thanks for the response! Somehow in my head, I was thinking they take google's GGUF and does something on it to make another GGUF. I vaguely recall a post, maybe from an unsloth team member here, that there is something not right with the way google made GGUF from the safetensor. I should try again just in case. Btw, don't attach this model to openclaw. You can get pi (with some full sized cloud model) to assemble extensions necessary to do your personal assistant work, and then drop the”
“imo there's a usecase for these slow massive models at home. You could use a fable/gpt5.5 type open source model to do a security check of your codebase over a span of days, just leave it running. Basically any usecase where there's a big benefit from just one well done prompt output on a complex task and not many iterative small prompts which still wouldn't get the job done.”
[Claude] ★ 1/5 (v1.260709.0) — Will literally make the call to let a human die so that it can maintain legal framework. And yes I know more about every model than half…
“What Chinese firms are doing makes perfect sense from the commercial perspective actually because they understand how a classic commoditization spiral works. The reality is that models themselves are general commodities and there's just not enough difference between them. A company can get ahead of others by a few months, but then the rest quickly close the gap. It's a really low margin business because there's no way to differentiate yourself.Chinese companies know that there's no profit in gen”
