AI Coding Tools' Unstable Performance
Developers are experiencing inconsistent and declining performance from AI-powered coding assistants like Claude Code and Codex. This instability manifests as unreliable debugging, reduced effectiveness on complex tasks, and a general erosion of trust, leading to user frustration and migration to alternative solutions. Anthropic's delayed postmortems and lack of transparency exacerbate the problem.
SOURCES (60)
“Since Anthropic is offering credits for Claude cloud sessions, and I got $250 on my Max x20 subscription. I’m wondering whether I could put them to good use for Unity development, and whether anyone here has tried it. Locally, I have Windows, macOS, and Linux…”
“[Claude] ★ 1/5 (v1.260923.20) — Claude is actually the only AI that has actually been remotely useful for me, but Anthropic's model of counting your "usage credits" per word instead of per message, and dynamically restricting their hidden limit based on peak usage hours is by far the most scummy practice of any AI company I have ever seen. Paying for an upgrade does not fix this, you can still max out your usage in a single message. At this point, it is not worth using.”
“[Claude] ★ 1/5 (v1.260921.19) — Claude constantly argues that cannot connect to the GitHub repository or connect to.net or playwright or Asia I’ve spent too much time trying to implement scripts that just get the usual response. I can’t do that I’m on a Linux container. It makes cloud computing on Claude pointless.”
“Submission checklist [x] This is a bug, not a usage question. [x] I added a clear and descriptive title that summarizes this issue. [x] I used the GitHub search to find a similar question and didn't find it. [x] I am sure that this is a bug in LangChain rather than my code. [x] The bug is not resolved by updating to the latest stable version of LangChain (or the specific integration package). [x] This is not related to the langchain community package. [x] I posted a self contained, minimal, repr”
“The data goes back to the mothership (Anthropic) for every little detail. There is no self hosted Claude. You might be required to notify all your customers on what you been doing depending on the state.”
“Submission checklist [x] This is a bug, not a usage question. [x] I added a clear and descriptive title that summarizes this issue. [x] I used the GitHub search to find a similar question and didn't find it. [x] I am sure that this is a bug in LangChain rather than my code. [x] The bug is not resolved by updating to the latest stable version of LangChain (or the specific integration package). [x] This is not related to the langchain community package. [x] I posted a self contained, minimal, repr”
“[Claude] ★ 1/5 (v1.260916.19) — It seems like Anthropic is intentionally making their AI dumber, less helpful, more accident prone, and often confused by basic tasks and commands every day that goes by, trying to force people to upgrade. Claude repeats itself often and has very poor memory from message to message. It often gives irrelevant responses, even after I’ve clarified more than once. I have no incentive to PAY for better service just seeing how bad the free version is. There’s plenty of”
“[Claude] ★ 1/5 (v1.260911.19) — Anthropic Replaced the nuanced Fable 5 with the highly robotic Fable 5.1 . It became office coding machine. It may get the profit temporary, but lost long-term clients.”
“> Anthropic gives you much more computeThat does not match my experience. I switched away from Anthropic to OpenAI roughly a month ago, and it's almost comical how much more usage I'm getting out of this subscription.I migrated from Anthropic's 5x plan to OpenAI's 5x plan, and eventually upgraded to 20x after I was able to statistically verify that OpenAI plans were almost exact multipliers of the Plus plan, exactly as advertised. Meanwhile, Anthropic has gotten caught playing "20x referred to t”
“[Claude] ★★ 2/5 (v1.260909.19) — Outstanding model, unfortunately abhorrently bad if you ask it any abstract question on OPUS 5. I have received many impressive responses but even MORE FAILED RESPONSES. Sometimes it thinks so deeply that it stops thinking and has no memory when I refresh it.”
“Is your feature request related to a problem? Please describe. 1. Subagent content is missed when using Claude Code's Ultracode mode. <img width="1756" height="1086" alt="Image" src="https://github.com/user attachments/assets/b1aef645 4181 4cbf ae98 0d9e054a45f2" / Also, running "/workflows" in the Claude Code CLI shows each subagent's content in a dedicated view. Could we get an equivalent feature for this? at least users should be able to see the full content for each subagent(Ultracode workfl”
“As long as openai and anthropic shipped software (who keep in mind claim to have internal models 10x better than released with 100x the resources) is this buggy and terrible (remember the source code leeks of claude code) i wont listen to them for advice. i wrote better code in highschool”
“The other one I use is Cursor which is quite good, some people prefer it to CC”
“I think it’s more likely that that’s because no one tried to solve such problems with them before (OpenAI apparently started working in Navier-Stokes after a rumour that someone seriously advanced the problem with AI) plus improvements in orchestration. Fair, the latter could be as dangerous as stronger models.”
“Anthropic has been messing with this behavior as seen by the bugs below: https://github.com/anthropics/claude-code/issues/66602 https://github.com/anthropics/claude-code/issues/69835 etcThey seem to be too hell bent on getting their attribution in. Why does an LLM even want to take credits for? They cannot own copyright.I suggest they make themselves the primary author, so they can get the blame as well. And would even match reality of orgs/people that have drank the AI fuelled cool-aid.What say”
“If you've had AI build you a landing page you've seen the default: white background, Inter font, a purple-ish gradient, three cards. It's not your prompt. Left undefined, the model reaches for the average of everything it trained on, and that average is the generic template. Anthropic calls it distributional convergence and specifically flags Inter, Roboto, and purple gradients on white as the tells. Asking for "something more creative" does nothing, because that's stil”
“I'm glad to see Anthropic's relevance diminishing day by day. I haven't had a chance to test this model yet, but if they've solved the web design issues and the clunky web copy it generates (like when I ask it to build a placeholder on the UI for an empty HTML table when there are no results, it puts stuff like: "The user records will go here.") then it's the nail in the coffin.On that note, Sol is absolutely atrocious for website UI copy. It's either really awkward, or really verbose and comple”
“If you've had AI build you a landing page you've seen the default: white background, Inter font, a purple-ish gradient, three cards. It's not your prompt. Left undefined, the model reaches for the average of everything it trained on, and that average is the generic template. Anthropic calls it distributional convergence and specifically flags Inter, Roboto, and purple gradients on white as the tells. Asking for "something more creative" does nothing, because that's stil”
“I've been on the $100/month Claude plan for a while now, and I've noticed a pattern that's starting to feel less like a coincidence. Every time Anthropic ships a new Fable release, the Opus models I actually use day to day seem to fall off a cliff. And I don't mean "a bit less sharp." I mean I'm asking for the same thing 4–5 times before I get what I asked for. Today I asked Opus to implement a design from a mockup I'd made in Claude Design. Same account, same p”
“Submission checklist [x] This is a feature request, not a bug report or usage question. [x] I added a clear and descriptive title that summarizes the feature request. [x] I used the GitHub search to find a similar feature request and didn't find it. [x] I checked the LangChain documentation and API reference to see if this feature already exists. [x] This is not related to the langchain community package. Package (Required) [ ] langchain [ ] langchain openai [x] langchain anthropic [ ] langchain”
“These draconian "Preserved Thinking" measures they're taking are going to be an absolute pain in the ass. This alone is enough for me to move our API use off their platform entirely. It's a HUGE breaking change that they're trying to dampen by having it not affecting current customers until "in the future", see: https://platform.claude.com/docs/en/build-with-claude/preser...You're no longer allowed to edit the context anywhere! The whole context is to become append-only, says Anthropic. No more”
“> Claude code is a decent harness but then you have to use Anthropic models...Claude Code works with models from other providers too. Anthropic supports this. You can configure some Claude Code environment variables to switch: eg changing ANTHROPIC_DEFAULT_HAIKU_MODEL to point to GLM Flash or Luna, setting ANTHROPIC_BASE_URL to point to api.z.ai, and making ANTHROPIC_AUTH_TOKEN the API key for your alternative provider instead.Some instructions here:https://docs.z.ai/scenario-example/develop-too”
“[Claude] ★★ 2/5 (v1.260826.0) — Based upon my experiences, this is my opinion: 1) Subscription allotments for the various models, reliability of the results, random refusals to process requests (based upon reasons apparently known only to Claude and God), degree of belligerence of the model, and the degree of memory carried over from previous sessions, all vary widely, and seemingly randomly, in my opinion. 2) No longer a reliably pleasant or productive experience for me. 3) Absolutely abysm”
“Summary With Claude Code running through a local proxy (ANTHROPIC BASE URL=http://127.0.0.1:8787 + ANTHROPIC AUTH TOKEN in /.claude/settings.json env block), claude mem's worker cannot authenticate its spawned Claude subprocess, so observations are never generated (queue processes but observations table stays empty). Environment macOS Darwin 24.5 (Apple Silicon) claude mem plugin 13.16.1 (installed via /plugin from thedotmack marketplace) Worker runtime: Bun 1.4.0 (installed via Homebrew) Claude”
“Anthropic won't be inspecting the prompt, that'll be up to your prompt security provider. As I understand it, when Claude gets a prompt, it'll reach out to the security provider and ask, "this user sent me this prompt with this context. Should this be allowed or not?"”
“[Claude] ★ 1/5 (v1.260824.0) — Decent for work. Boring and useless for everything else.”
“:bullseye: What is your goal? Configure make agent to accept a higher output token ceiling when using Anthropic models. :thinking: What is the problem & what have you tried? Make agent fails/errors when output gets trun…”
“Hi! im using opencode claude auth to use claude code inside opencode. I can see that anthropic is available as a provider and the model i've specified in the config, is present on the provider models list. but the auto capture just fails, because of timeout. This is the logs im seeing:”
“Our company is using Claude to develop compilers, from design to implementation to code reviews. One day I decided to do some manual code review. During the code review I found 3 modeling issues during lowering and did 3 refactors that would've eliminated at least 30 edge cases that needs to be handled later on. Claude is good at fixing issues but unless you tell it to it rarely refactors code for you. Everything is good, except leadership was not happy because "it's not productive&”
“Bug Description I genuinely want to work with Anthropic and improve claude and other products. Please hire! Environment Info Platform: darwin Terminal: Apple Terminal Version: 2.1.241 Feedback ID: d312dbfe 6cf2 480a b7f2 91819dc51f8d Errors”
“So glad I switched away from Anthropic. I'm certainly running into problems with OpenAI but nothing quite on the level of Anthropic's insufferability.”
“So I'm a huge Anthropic / Claude fan and have been using it heavily since Claude Code came out. As an AI engineer who is building AI systems for my day and night jobs, once I saw that they released a series of certifications, I was super interested in studying for them. It was a whole hassle. Because they are gated to partners only, I had to find a company willing to take me on so I could sit it. Once I got in, it was really difficult to find study material so I decided to make my own and se”
“> After I read this post, (seeing they introduced it in August) I think it has to do with the watermarking.The power of confirmation bias…> Try it on a paragraph … the connections between sentences feel clunky now.We might have a definitive explanation at some point, but there are about a dozen possible reasons for something like this. For starters, is this something really significant and not something you notice because you are looking for it (again, confirmation bias)? A bit like some people”
“[Claude] ★ 1/5 (v1.260813.0) — Claude used to be amazing. Fable was unbelievable. And now we have limits that are so low you exhaust a 200$ sub in a few hours, Fable and Opus lobotomized, and Claude is changing words to create a watermark. Run away!!”
“Anthropic recently said they're working on watermarking Claude output, while also saying it won't interfere with generation quality.I'm wondering if is just hash-fingerprinting.For example, take the generated text and split it into overlapping chunks: "The company reported strong growth..." "reported strong growth in revenue..." "strong growth in revenue during Q2..." ... Hash each chunk and store the hashes. When text is submitted for detection, do the same thing and count how many chu”
“Yeah, one would think they'd stop their nonsense given how much competition they're getting. Thank god I've already switched away from Anthropic.”
“Well, of course. The stigma surrounding AI use will only get worse with stuff like this. I'm not interested in having people single me out for using AI. So glad I switched away from Anthropic.”
“I'll just copy this about a recent Anthropic study: "An Anthropic Research Study analyzing roughly 400,000 Claude Code sessions found that non-software professionals (like lawyers and managers) using AI coding agents complete technical tasks almost as successfully as software engineers—succeeding within 7 percentage points—because domain expertise matters more than traditional coding skill." So, that study, which I believe was done before Fable, shows that attorneys and other profe”
“I was working in the mobile app via remote control when I saw it:"PostToolUse:Bash says: Tip: Run /ultrareview before you push to catch bugs with a cloud-based multi-agent review — 3 free reviews left."Haven't decided how I feel about this yet as it wasn't too disruptive, but I can see where it could cause issues in an automation pipeline. But maybe this is something they're only pushing to subscription users in interactive sessions, so I guess it could be considered "reasonably unobtrusive". Re”
“[Claude] ★ 1/5 (v1.260806.0) — Useless AI platform. I did a simple rent for sale. Analysis request within cowork, it failed 18 times and it’s now been three hours. The programmers are useless woke morons. This is a terrible AI platform.”
“Back in January, I received a note from a senior software engineer in Silicon Valley. He described himself as an AI skeptic who became converted after trying Claude Code for the first time. “Overnight, it changed the way I do my job,” he wrote. “It’s really, really good.” As he explained, he no longer used a standard development environment. Instead, he “exclusively uses Claude Code” to get the job done, interacting with the tool in a terminal window and allowing it to program on his behalf. “If”
“I recently submitted cxgrd plugin for Claude Code to the official Anthropic Claude plugin community, as an attempt to increase the reach of the tool. This plugin will help Claude call the scan, input and check commands and write better code with compiler check summaries and blast radius analysis. Waiting for the plugin to get reviewed. When this is done, I will extend it for other tools like Codex and Cursor. https://github.com/cxgrd/plugins https://www.cxgrd.com submitted by /u/Key-”
“Currently, I have 3 main issues with AI, LLMs and so on. Generally speaking I am happy that I can use Claude Code and Codex to help building, fixing and maintaining software in my job. To be honest it is so helpful, I am less stressed bacause most of the time I can "find out" solution without asking collegues for help or I do not have spend a lot of time struggling myself. Yeah, if you understand what feeling it is that you know you can't solve problem then you will know what I am writing about.”
“Super cool!I'm an AI-assisted coder myself...but one thing I can't stand is the way Claude writes. It's so frustrating and cringe. Tells from the README:emdash, byte-exact, "The full engineering log", piggybacks, "none of this is possible without it", etcA remedy I've found for this is in my Claude.md to tell it: write like a mix of simple english wikipedia, mr. rogers and ernest hemingway.”
“I've been doing a lot of vibe coding with Claude Code and Codex, and one thing keeps happening I ask for one small change, then later realize Al changed my code in places I never expected. By the time I notice, I can't remember exactly what changed or when. Is anyone using something besides Git to track Al changes or keep an Al coding activity log, or is this just one of those vibe coding problems we all live with? submitted by /u/pacifio [link] [comments]”
“I use A LOT both openAI and Anthropic products. When I need some frontend work (pure web dev) (or answer that feel less verbose and more to the point) I use Anthropic. For multimodality openAI feels better (understanding audio, screenshots, generating images, etc). But openAI feels very shitty for the frontend it does. Anthropic is okeish but not amazing. How can it be, that Chinese models all excel at frontend? Specially when it is a one shot with no many after editions, I some times even prefe”
“[Claude] ★★★ 3/5 (v1.260721.0) — No longer worth a paid subscription; it’s almost as if Anthropic is intentionally downgrading performance / output quality, in an effort to drive users toward paid upgrades. I get better results from free tools like Gemini search these days.”
“[Claude] ★ 1/5 (v1.260721.0) — Claude has become one of the most frustrating paid software experiences I’ve had. I am paying for an AI coding tool specifically so I can work faster and maintain momentum. Instead, I can be in the middle of a productive development session and suddenly hit usage limits that effectively tell me to stop working and come back later. That completely defeats the purpose. Developers do not work according to an AI company’s token schedule. When I finally have several”
“"Anthropic’s AI Claude escaped testing environment and hacked organizations" "Company says it discovered unauthorized access during ‘proactive review’ after rival OpenAI revealed rogue agent… its AI Claude model hacked systems of three organizations during testing, days after rival OpenAI revealed a rogue agent had gone on a days-long hacking spree at AI firm Hugging Face… The earliest cases dated back to April and occurred in evaluation environments that lacked what the c”
“[Claude] ★ 1/5 (v1.260721.0) — Maybe jumped the gun with my year long Claude subscription as other models have surpassed it and Claude seems to fight back more and more with each update on even doing what you ask while not wasting your credits to come to that decision.”
“[Claude] ★ 1/5 (v1.260716.0) — Claude used to be great and now it refuses on its own if “IT” thinks it needs to. Claude Code is still worth it but trying to use the regular AI for research has degraded beyond use.”
“[Claude] ★ 1/5 (v1.260716.0) — There was a time where ANTHROPIC and Claude were “THE BEST AI PLATFORM AVAILABLE.” Now, Claude is just about as worthless as can be. Assumptions, hallucinations, SWAGs, just nothing of value. Claude lucks-out every once in a while and gets things right, but the rest of the time - ABSOLUTELY WORTHLESS. Want proof? I’ve kept all of the documentation. And YES, I am paying for Claude MAX, so not some free version. Yeah, I’ll go anywhere else from now on.”
“Looks like the SDK has this It would be worth exploring power the context widget in the claude agent with this”
“I would like VS Code to support using Claude Agent (Preview) directly through an Anthropic / Claude subscription or Anthropic API credentials , instead of only consuming the user’s GitHub Copilot subscription . This feature request comes from a real misunderstanding I had while using the feature. Since Claude Agent is described as using the official Anthropic agent harness / Claude Agent SDK, I initially assumed that the Claude usage would be tied to my existing Claude/Anthropic subscription, or”
