AI Keeps Hallucinating in Front of Real Customers
AI users are frustrated that the focus on prompt engineering overshadows the broader context and understanding needed for effective AI interaction. The core issue isn't crafting the perfect prompt, but ensuring the AI has sufficient background knowledge and a clear understanding of the desired outcome, leading to outputs that are technically correct but ultimately unhelpful. This shift highlights a need to move beyond superficial prompt adjustments and address the underlying integration and contextualization of AI within workflows.
SOURCES (60)
“Spot on observation, but the hallucination isn't wasted params. it's what 1-bit does, factual recall lives in fine weight distinctions and gets flattened first while grammar/flow survive. hence confident fluent nonsense. which is exactly why your tool-loop idea is right: don't trust its memory, hand…”
“Running roughly this shape daily, so a few practical notes. It works, but there is a floor to how lightweight the brain can be. Below a certain size the model stops being reliable at the part you actually need it for: deciding when to search, forming decent queries, and not blending retrieved text with made up filler. That reliability is the "intelligence" in practice. Two things that mattered more than model choice for me: have the search tool return small clean extracts instead of wh”
“If a computer is doing the maths, then another computer (a bot) can also do that. What you're doing is not proving the difference between a human and a bot, you're just eliminating the crappy bots. Anyone could set up a bot that could bypass a Proof-of-Work mechanism, it's not difficult. In-fact, the ideas that people are explaining to you on this thread almost all utilise a headless browser. Further, those kinds of PoW (PoW not POW) mechanisms also push out real people using lower e”
“yeah the reason you only see ai slop in slide decks is because most pms are just using it to generate text they should have written themselves. where it actually helps is synthesizing massive piles of unstructured qualitative data. if you dump 50 customer interview transcripts or a mountain of chaotic slack feedback into claude and ask it to find the recurring friction points, it saves hours of manual tagging. using it to write the actual prd without human review is dangerous and usually leads t”
“Maybe I didn't write it clearly. I used Russel's teapot preemptively before someone tries to shift the burden of proof (that is why I wrote "before you say: How do you know"). In the context of your actual comment: yes, Russel's teapot is irrelevant. I'm not the original commentor you intially responded to, by the way. My problem with your thought experiment is, that it constructs a case without explaining how anything of that is relevant to the real AI discussion. Beca”
“Agreed, and the accountability piece is the part my framing skipped. I deliberately kept it narrow because I was solving for "how do you stop the tool war," but you're right that a clean acceptance bar is worthless if the person feeding it nonsense acceptance criteria never feels it. What worked for us: the owner who accepts an output owns everything downstream of it too, not just the decision at the sign-off moment. So bad AC isn't "the AI got it wrong," it's &qu”
“Does it explain how this demonstrates bias? If you’re forcing the model to make an arbitrary choice then it’s a coin flip on whether it chooses option A or option B. What methodology does it use to prove there is a bias?”
“Why are textbooks relevant here? Even if you repeated the experiment with a textbook instead of the AI and got the same result, what conclusion would you draw from this? The general conclusion of the study seems to be "giving people access to authoritative-seeming but wrong tools for answering questions outside their area of expertise reduces their ability to say they don't know the answer, even when the answer is wrong". So yeah, don't buy bad textbooks for your employees if you don't want them”
“Personally, I find the technology very interesting. I want to understand how LLMs work, and when it doesn’t work, which is not infrequently, want to understand why it didn’t work. AI has made my work product better and added to my enjoyment of practicing law (which I’ve done for 34 years) and by making some of the drudgery more interesting and less time-consuming. Sure, LLMs are overhyped, and I don’t believe it will lead to AGI. Skeptical about labor market disruption predictions. And there is”
“The other day I asked Claude Haiku (the dumbest model) if dark energy was just the self-energy of the Higgs field. I know barely anything about quantum physics. I just wanted some example of AI writing style - something like "You're absolutely right! That's a key insight into a load-bearing ..."Instead it spat back a bunch of physics related stuff, alleged problems with the idea ("there are 30 orders of magnitude difference between that and dark energy, and why the Higgs field and not any of the”
“The key insight a lot of people miss is that Claude or any LLM shouldn't be doing the reconciliation directly — it should be flagging anomalies in a reconciliation you already know is correct. Think of it as a three-layer architecture: 1) Deterministic extraction layer: rules-based (VBA, Python, Power Query) that always produces the same output from the same input. This is your existing Excel automation, and it's the right foundation. 2) Cross-document validation layer: This is where LLM”
“honestly the review data here matches my experience, review is the constraint, not writing the code. that's the whole reason i spend my time on it.where i'd push on the conclusion is the 400 lines/hour ceiling, which assumes you're reading an undifferentiated diff top to bottom. most AI PRs aren't uniform, a big chunk is mechanical (renames, boilerplate, repeated patterns) and a small chunk is the actual logic that needs a careful human look. We see more of these PRs now as companies use AI to c”
“I tried this one you suggested the Arobis AI visibility checker, but I also tried a few more. My two cents are that while it really gives you a comprehensive look on your current state in AI visibility I would sttil want to see more data showed, like besides the things you need to fix. What about things like your status in Microsoft Co-pilot? or within Grok? Also what about like hav ing 2-3 competitors where the checker can tell you where they stand vs where your brand stand? Again, I don't”
“The issues I find in setups like these is that there’s not enough clarity between instruction and context, and the models are good enough on their own that I’m not sure it’s obvious whether the output is actually a result of the harness you’ve put together or not. Since you’re prompting through Notion AI, there’s also the added friction of having to prompt it with specific instruction each time to ensure it follows the flow you’ve built for it through Notion pages. I use Cowork, but most agent h”
“When AI image gen first reached the public, they used it for haircut ideas. It just frustrated them and their hairstylists, though, because the AI's recommendations were impossible to recreate in real life. Are you taking any measures to ensure that the haircuts your system suggests are actually realistic?”
“> And then I discovered SO meta. Holy cow. Those people were so far up their own butts, they couldn't see daylight. I was morbidly transfixed.It's validating to see that I wasn't wrong in my assessment. The comments and takes under the post which showed the drop in traffic was eye opening. That website was just a walking corpse if they weren't seeing the plain truth and spinning the drop in traffic as a good thing, because apparently it meant that they were not getting stupid questions anymore a”
“I'm an Ironclad user. The AI playbook allows me to flag certain clauses, and on what conditions they should (or should not) be in the contract. Both missing clauses and those that don't match our pre-approved positions are flagged.”
“About a year ago we added AI features to the product. And overall I think the rollout has been solid, but it has introduced many tickets around incorrect information or flatout escalations. One of our main issues early on was that we had no way of knowing what the problems with the responses were without screenshots from the customers. We'd have the support ticket, but not the actual conversation between the customer and the AI. We needed to escalate all the tickets to engineering, who would”
“AI is good, but inside it, there is a problem called 'Lost in the Middle,' and the percentages for the first tokens are always high. As the conversation goes on, the accuracy degrades so much that it becomes so mute [incapable] that it fails to understand either logic or semantics, and then it gives such a weird thing that it makes no sense. Furthermore, AI normally understands linguistic logic, not facts, so you have to somehow graft a real logic onto the linguistic one. What I personal”
“All AI models have biases and defaults. I wanted to explore these, and what started as some some tinkering, turned into a rabbit hole and a full blown project. I asked 100 AI models the same 100 simple questions three times, producing 30,000 answers in total. I organized the results by prompt and model and added an unofficial benchmark called ConsensusBench, which measures how often each model gives the most common answer. The prompts were ran through OpenRouter, and the original dataset is avai”
“Sure they do, but a completely generic statement with a very-likely made up example is not at all useful for this sub. Since this is marked as "Need Help", the author should provide more detail in order to achieve their stated goal.”
“This has happened to me too. LLMs hallucinate. It happens. Are you a bot? You're being weirdly defensive.”
“The Google DeepMind-sponsored Kaggle challenge "Measuring Progress Toward AGI - Cognitive Abilities" asked participants to design new cognitive-science-based AI benchmarks and they just announced the results this week. In my two posts I present evidence that deepmind & kaggle rewarded a nonsensical number generation machine and a litany of unfounded claims with 25k and a grand prize stamp. What the authors of the work I analyze intended to do was to present an LLM with alternative”
[ChatGPT] ★★★ 3/5 (v1.2026.188) — I’m not writing this because I want an AI that sounds smarter. I want an AI I can trust. I’ve spent hundreds of hours working with ChatGPT…
“Hello my friend and I promise you what you're going through is more common than you know and it's not your fault, it is the result of deliberate training that makes these models genuinely infuriating to do anything other than code with, but I know ChatGPT is a household name but it is far from the only end if you look for genuine collaboration I promise you will find it just not here anymore unfortunately. I'm not gonna suggest platforms because I don't wanna give the system an e”
“The AI only mentions him "peak"ing inside the box. It is weird that it uses that word incorrectly. He doesn't just peek inside. He removes it, and the AI makes no mention of this, because the AI makes mistakes.”
“I don’t know why it produces concise and better-reasoned responses. I use Cowork for larger drafting tasks or for agentic-oriented tasks. My experience is output for Claude model inside Perplexity is better (my subjective evaluation mostly but I think objectively in many cases). I’m curious of others experience but doesn’t look like I’m going to find out. Try it.”
“What is the implication to Notion AI? submitted by /u/louisleung1991 [link] [comments]”
“This is the one that gets me too. The wrong clause you catch on the read. The assumption everyone in the room shared and nobody wrote down — you can’t catch that on a read, because you’re reading with the same assumption. The one thing that helps, a little: check the doc against an actual list of what should be in it. A playbook, a precedent set, whatever. At least then the known-required stuff that’s just missing gets flagged instead of relying on someone happening to notice the hole. Doesn’t t”
“Agreed, if the code is written to be easily testable, it becomes easier for AI to write meaningful tests.”
“I liked your post and I enjoyed reading it, who doesn't enjoy reading about Feynman?But let's get scientific about this, you say:> The test is absolute. There's no way to beat it; no way to fakeWhere's the proof of that?”
“I'd love to see the 'AI as personal tutor' approach. Even incorporating things like spaced repetition or the testing effect, or evaluating free-written responses.A lot of potential that's currently unrealized. It takes a student to swim upstream to get there. The convenience of cognitive offloading is difficult to say no to. For evidence, I see it everywhere at work, including (at least in some cases) in my own work, for matters I don't care to invest effort in learning because it's a one-off.Th”
“The honest conclusion is not that AI is secretly draining the country dry, but it is also not that its water use is trivial. Data-center water demand is growing quickly, is poorly disclosed, and can be significant in the communities where facilities are concentrated. New AI datacenters that are being built by Microsoft, NVIDIA, Oracle, and others use closed loop cooling which uses almost no water. Hot water is send to chillers and then recycled back through the system. That doesn't address t”
“Yes, absolutely. Everybody’s being encouraged to use AI and it is making red lines more verbose. incense everything is a rush. You might be seeing more of that. I’m also seeing much more sophisticated red lining from unsophisticated parties. But that’s the name of the game, and you absolutely don’t have standing to ask if there is human review. That would just ruined the negotiation.”
“I’ve been using AI tools daily for client work for about a year, and for most of that time my prompts were basically stream-of-consciousness. I’d type what I wanted, get something mediocre back, adjust, retry, adjust again. On a good day it’d take 3-4 rounds. On a bad day I’d give up and rewrite the output manually. A few months ago I got frustrated enough to sit down and actually work out what my best prompts had in common — the ones that gave me what I wanted first try. Turns out they all had”
“Time and practice and best practices/forms and luck. I’ll say though: this is always the hardest part of the job to come to terms with. There could always be something missing. Always something no one anticipated. Always an unstated assumption you and the drafters made that is not obvious to later readers.”
I think a lot of people are using AI wrong for this, and I say that as someone who ran programs for years. The mental model that fixed it for me is…
“How do you run Perplexity on your dataset? Is this “Perplexity Computer for Counsel”? I heard that perplexity was coming out with a legal harness (e.g., something like Claude for Legal running inside of Claude Code/Cowork) to be rolled out to their enterprise/max users, but didn’t know it was available yet. So when you say, “ChatGPT 5.6 inside Perplexity Max,” is that GPT-5.6 inside of Computer for Counsel?”
“> Amusingly, when I know my peer is just going to point his AI at my feedback, I write for their AI, not for them.At this point, why not just talk directly to the AI?Asking as someone who is likely leaving dev, and maybe tech completely, very soon, possibly to go wait tables, as he hates everything about the way things have gone in the last half decade or more with remote work, stupid levels of unnecessary complication everywhere (people architecting to be the next Amazon as soon as, or even bef”
“The project involves creating unambiguous STEM prompts that fail both AI models. I've been on it for days now and both models have gotten the answer right each time, how can I get this done, please anybody know something that could help? submitted by /u/WesternBaker9913 [link] [comments]”
“Hi everybody, Hopefully, you can help me with this issue—or perhaps it isn't an issue at all. When AI suggests a solution, I find myself unable to come up with a better one for the problem at hand. On one hand, I appreciate this because the problem gets solved. On the other hand, it bothers me because I don't feel like I’ve really done the work myself or that I fully stand behind the code being written. submitted by /u/ifstatementequalsAI [link] [comments]”
“I've been trying to create STEM prompts with one verifiable answer that stumps the reasoning of the AI of the models but they always seem to get it right even after layering so many obscuring observations. Can anyone help? submitted by /u/WesternBaker9913 [link] [comments]”
“Who detests the dreadful AI descriptions? How do I change it back to the original text? I am selling cards on Etsy and tried the AI script but am now having to manually change it back to plain English. Because I can't bring myself to sell cards with these descriptions! These are the mills and boon type of descriptions! Who writes English in this way?!! "Crafted with remarkable attention to detail, these cards are perfect for every occasion, whether it's a birthday, thank you note, o”
[Claude] ★★★ 3/5 (v1.260709.0) — Still needs much work to be the best AI
[ChatGPT] ★ 1/5 (v1.2026.188) — For something that should be able to research and be knowledgeable even something as basic as getting an email or website setup is a challenge for it.…
“This is also how I have seen this particular thought terminating cliche used. The problem with the framing is that on the one side you have someone complaining about unrestrained slop and then this thought terminating cliche is offered. Why yes, it matters how you use it and the complaint is that users are not refining the output enough before presenting it to others.The thought terminating issue with the phrase is that it isn't just a tool. Once you automate its use (automated PR reviews, ticke”
“Let's turn this around.I don't use any paid AI or any agents. I just use some of their free interactive question answering interfaces.There have been some times where I've been doing a project in a language or environment that I do not use a lot, and maybe need to use language or environment features I've never had occasion to learn.I don't ask AI to write it for me, or even to write any particular functions. I might ask it some syntax questions, or what data structure in the language's standard”
“Yes. Absolutely.I feel like it's also making people even more lazy. It feels like people just ask questions without putting ANY effort in themselves beforehand to find the answer, like they assume everything is an AI and we're all just going to drop everything and give them a gushing answer. Like the whole concept of a manual or documentation for something seems like the biggest waste of time now, because it feels like no one reads it (or maybe haven't got the attention span to read anything any”
“The core thesis of this essay is reminiscent of the Lisp Curse [1] / Bipolar Lisp Programmer [2].It's been a few years since I read these, but if I recall the argument there, it was that Lisp makes it so easy to build stuff and scratch exactly your own itch, that there's no real strong push for lisp programmers to come together and collaborate to build non-trivial and general purpose artifacts. And that is why the landscape of public lisp software is poorer as a result, compared to languages whi”
“Yeah I have found that AI in general has a procrastination nullifying effect.Before dealing with anything that might put me off. I can just ask the agent to do it for me. And then, do something else, take that break, but regardless in a few minutes I will have something to jump on instead of the same blank terminal with the same blinking cursor judging me. It really makes taking the first step, much easier and then the ball just gets rolling.I see what his point is to be honest though, it's easy”
“Honestly, I wish that LLMs were better at challenging their own assumptions, or even just stating them for me to validate before rushing ahead. By far the biggest aggregate waste of time for me with them is how they all seem to be tuned to try to guess what I'm going to want next and give it to me in advance, when in reality what I want is very commonly dependent on what I get back from the current thing. Sometimes I swear they must have been explicitly trained to treat as many questions as poss”
[Claude] ★★ 2/5 (v1.260709.0) — The AI models are overall very good. However, it randomly stops mid task and says “response incomplete” even when limit is still not at 100%. And the…
[Claude] ★ 1/5 (v1.260702.3) — I would recommend any ai platform over this one. Very lazy answers. I constantly get answers like “i have no way of knowing that” if i push…
[Perplexity] ★ 1/5 (v26.26.0) — I was using Perplexity to help me write a fishing guide. This AI MADE UP imaginary lures and supporting statements for me to use. Going back, I…
[Perplexity] ★★ 2/5 (v26.25.0) — Fails basic memory, fails simple tasks. ChatGPT, Manus, and Claude are 10x better than this one. Perplexity is fast but honestly it’s just another chump AI trying…
