AI Image Users Fighting Unpredictable Output
Users are frustrated with the unpredictable quality and lack of control in AI image generation tools like ChatGPT and Gemini. They struggle to achieve desired results, encountering issues like illogical outputs, artificial-looking images, and text errors, leading to wasted time and unmet expectations. The evolving landscape of models and tools adds to the confusion.
SOURCES (60)
“You’re on to something. I’ve messed around with this idea in my head and I’m amazed it’s actually working so well. Keep it up! I’m really curious where this goes.”
Thanks for the advice, I will definitely check out DaVinci Resolve.
“It's always the high-quality 3D (AI?) simulations that require me to install Chrome... thank you anyway! I will try it”
“I have a question I sometimes will make pictures and does the thing make like individual pictures sometimes I’ll ask it to make six individual pictures or something like that. Sometimes it doesn’t. Is there a way to get to listen to me or is there an actual limit on that because sometimes I’ll ask our sixth individual pictures I’ll make a outfit but I want like six varieties of the outfit in the same style or something and it will give me like six outfits in like one picture I want like six sepa”
“I've been working on a small project called "Veramanu". The basic problem I started from was this: we're getting very good at generating images, but increasingly bad at knowing where an image actually came from. Trying to detect AI from the finished image alone didn't seem like a very solid long-term answer. So I take a different approach. Part of the creative process happens inside the platform, and that process can stay attached to the finished work. It's still early.”
GPT Image 2 produces plastic-looking faces in product edit
“Well try pokemon, not really that good in chatgpt, I prefer gemini for pokemon, but for other anime chatgpt is better, this is my example https://preview.redd.it/ud01ld461bnh1.png?width=843&format=png&auto=webp&s=955f91eb5a2fd7cd66b48913b81cfd6b0b72150e”
hmmm got it, so instead of "create an image of a fishing reel", I would share my pictures and with my prompt try to increase image quality, style or something like this.…
“I can’t decide which image better captures the atmosphere of our ant colony RTS/sim. I feel like the bottom one represents the game better, but it also has less contrast and doesn’t stand out as much. What do you think — A or B? Thanks for the feedback! Tomas Steam page: https://store.steampowered.com/app/3016940/Garden_of_Ants/ submitted by /u/Able-Sherbert-4447 [link] [comments]”
“the more copyrighted or popular something is the easier it is to create lol it is just more training data. It is not an ability you add, but an ability you would restrict, and ig openai didn't care to restrict that side of things as much.”
“I have not noticed any image generation model that can replicate at all a copyrighted character with this perfect detail or almost perfect to the background as well. What kind of black magic they did? Because bo model i saw is not even capable of making this perfectly unless very specific loras trained with 1 character is put in the middle of a model. And almost none can make them perfectly like almost all make completly random bodyparts not reflecting the actual character Is there any image gen”
“Be specific, which opacity, background, texture or tint? Can you post a screen recording.”
“It may help that this happen to me too when I don't have opacity set to 1.0(opaque).”
“After generating anywhere from 10 to 100 AI images a day since the early days of Stable Diffusion (well before tools like ChatGPT image generation or Nano Banana), here are my top three takeaways for e-commerce store owners: Always start with authentic product photos. Even with top-tier AI image generators, you still need real photos of your actual product. Inaccuracy leads to mismatched expectations and high return rates—and every e-commerce seller knows returns kill margins. Master product-to-”
“Perhaps a stupid question, or me just overthinking it 😅 Since uploads follow each other fast (which is off course great!) and I upload manually with tagged versions: should I update version per version, or can I skip them and immediately go to the most recent one at the time of update? I don't manage always to update immediately and I'm anxious to break anything... 🤐 Thanks for your hard work and keep it up! It's really unbelievable how good the tool has become so fast!”
“Instead of generating each image individually. Just have the model generate a collage of 64 then you pick and choose which ones you like. Upscale and boom you are ready to go. 😊 submitted by /u/Full_Supermarket_109 [link] [comments]”
Hey /u/MrJuart , If your post is a screenshot of a ChatGPT conversation, please reply to this message with the conversation link or prompt. If your post is a DALL-E 3 image…
“This is getting asked from time to time, but since models changed a lot, I wanted to reask it. I'm looking for a trained model that can give short descriptions about an image, simply for an alt text of pictures taken with a smartphone. Should I just throw it at Qwen3.8/Qwen3-VL or are there better models trained for it ? Similar to https://www.reddit.com/r/LocalLLaMA/comments/1oar481/what_is_currently_the_best_model_for_accurately/ submitted by /u/AnyNameFreeGiveIt [link] [”
“After posting this Mythbusting ChatGPT "Secret Slash Commands" + Free Image Preset Keyword List , I am thinking about what are the poweful keywords that makes UGC-style Ads feel realistic? SO, I run deep research for any repeatable prompting patterns behind realistic AI-generated UGC. The useful part wasn't one "magic keyword." It was combinations like: handheld smartphone slightly off-center framing natural room lighting everyday background clutter natural skin texture r”
“Forget pelicans. What do your models produce in a single turn, no harness, for this prompt (include your exact model Hugging Face ID or equivalent, with quantization and runtime ): Draw an ASCII art house in the woods with a chimney, two windows and a door between them and two horses in front of it. And what do you get with your harness of choice, same model? I found the reasoning to be quite insightful. Let's see which models / responses get the most upvotes. PS: This post only low effort i”
“Hi, Thanks for the detailed message, and for mentioning Pixort. That 1/2/3/4 keyboard sort is the right way to think about it. Your setup will be fine. The vision model runs on the 3060 and everything stays on your machine, nothing gets uploaded. It's built to be left running on a big folder, so if you'd rather have accuracy than speed, letting it go overnight is exactly how I'd use it. The Windows build is live. It does the dedupe properly. It groups near-identical bursts and can ke”
“How does this picture looks in stepping art https://preview.redd.it/708251t8w3nh1.jpeg?width=716&format=pjpg&auto=webp&s=2beeab18bcfff366478c0ad486f6a36452f3cfc4”
“[ChatGPT] ★ 1/5 (v1.2026.237) — When given very simple assignments like adding a logo, it can’t do it. It will waste and plow through your photo and image quota by giving you sloppy iterations. And ultimately, it will simply apologize and say sorry without ever getting the task done; pretty crazy that powerful AI can’t do the most basic things.”
“been playing with GPT Image 2 and really liked how this one turned out. I wanted a quiet Japanese travel-zine kind of look, simple train interior, lots of empty space, imperfect ink lines, soft watercolor, slightly aged paper. The orange seat ended up being a nice little focal point too. Prompt below if anyone wants to try it: Create a minimalist vintage travel-journal illustration of a modern city train interior: a row of empty blue-and-white seats beside large windows, one distinctive orange s”
“Can you share some of the prompts you used? Specifically regarding the sketches… or can you point me in the direction for resources to learn? This looks really helpful regarding character design (obviously).”
“Working with Chat, it took a lot of back and forth to get the image generator to get ahold of this one. It seemed to get the hang of it until trying to generate a sprinting scene/pose. Otherwise this was a fun stupid project today. Link to convo below. submitted by /u/Optimus_Spider07 [link] [comments]”
[ChatGPT] ★★★ 3/5 (v1.2026.237) — Make the photo look more realistic
“Give it a vector file (svg works well) and explicitly tell it not to draw the shape, but use the exact vector. It should be allowed to scale or color the vector as necessary based on the image context, but it should never be allowed to edit the shapes/points that make up the vector. It will still only get it right half the time.”
“I have a project set up with all of my orgs branding guidelines and I use it to generate thumbnails for organization data. It works great, but one particular thing I struggle with is getting it to draw non-standard shapes. For example, we have a city that I would love for it to be able to reliably draw the boundary of, but it just can't do it. I've given it png files, vector files, GIS files, and everything in between in an attempt to get it to understand the shape, but every time I ask”
“It’s been nearly two days since I generated an image. Last night it said resets at 3am. I check at 8am, it says it resets at 3pm. I check at 3pm. It says it resets at 9pm. Why does it keep pushing the generation time? submitted by /u/PregKittyGal [link] [comments]”
“[ChatGPT] ★ 1/5 (v1.2026.237) — Thumbs down I type what I would like my image to be and it gives me the opposite of what I want”
“That's exactly what I did. My point was that in the beginning the image generated by my gpt was by default, white, thin, ect...I constantly criticize my GPT for the lack of diversity in the image generation and visibly after years of doing it. It registered”
“First off I'm an artist not an engie. I'm building an app and needed an icon set. I kind of wanted abstract, lo-fi looking graphics, something out of the movie The Andromeda Strain (1971). Went to Recraft thinking I had a good prompt. I didn't. First thing I got wrong was describing how to build the shape instead of naming the thing. I wrote "a thick rounded bell shape with a squared arc handle on top" and got slop. Changed it to "depict a kettlebell icon, front view&q”
“Using Chatgpt to change discoloration in hands , but Chat only gives me a version with far less quality Is there an AI photo editor that keeps 4k quality in export pic? Or is there an editor i can feed GPT pics too to upscale? I'm a newbie to AI submitted by /u/Ecstatic_Dot_4256 [link] [comments]”
“Using Chatgpt to change discoloration in hands , but Chat only gives me a version with far less quality Is there an AI photo editor that keeps 4k quality in export pic? Or is there an editor i can feed GPT pics too to upscale? I'm a newbie to AI submitted by /u/Ecstatic_Dot_4256 [link] [comments]”
“12B is an odd choice for vision. Even with resolution maxed out, I find that E4B —or 26A4B if your machine affords— can see more details in the pictures. E4B may not recognize a Vermeer painting and attribute it to Lorenzo Linguini the famous Renaissance painter from Pizza but it will see the girl.”
“That's kinda what I thought too but still wanted to ask, as I saw Photoprism does it (downloads tensorflow...) and other apps do it for models (not code though, afaik)”
“been testing a prompt that turns each uploaded photo into its own editorial-style poster. The idea is simple to reinterpret the same image as a very minimal hand-drawn illustration. Lots of negative space, paper texture, restrained colors, and very little typography I wanted it to feel more like an independent art book / contemporary publication cover than a commercial poster. also, if you upload multiple photos, the prompt tells the model to make one separate poster per image , rather than comb”
Hey /u/Zachary_Lee_Antle , If your post is a screenshot of a ChatGPT conversation, please reply to this message with the conversation link or prompt. If your post is a DALL-E 3 image…
“[ChatGPT] ★★ 2/5 (v1.2026.230) — It’s constantly distorting photos and changing faces. Very frustrating to use because when it works, it does a good job and fixing old pix but sometimes it just goes off the rails and they look ridiculously bad.”
“I asked Qwen to make a .txt file and make ascii art of a goose. It made large ascii letters saying OBBBI and then below that a diamond shape made out of / and \ Why is my Qwen retarted”
“First prompt: "generate an image that is the hardest possible for you to make" Next: "Now make the easiest" chat link: https://chatgpt.com/share/6a956bfd-9e2c-83ea-b800-f1caa97a9871 submitted by /u/BruhMamad [link] [comments]”
“Is there an AI app that is free, and does not require any subscription, that will do image editing for me where I can tell it what to do verbally? I want to do several editing stages involving a few images and some basic modifications, adding and customizing some artificial elements, and producing a resulting image that I can save. Using the voice mode in chatgpt, it was able to confirm and understand what I wanted for the first step, but could not actually perform the editing and produce an ima”
“Well its latest imagegen model seems to be flux 1, and the official text bridge basically only supports kobold (whick can not into batching) and has not been updated for half a year or so”
“[Claude] ★ 1/5 (v1.260828.0) — This should be used with caution with focused outcomes only. Do not share every screengrab. Problem is- it doesn’t know how to actually do anything.”
“[ChatGPT] ★★★ 3/5 (v1.2026.230) — Image generation needs improvement to reflect the real image or at least not to degrade.”
“Prerequisites [x] I am running the latest code. Mention the version if possible as well. [x] I carefully followed the README.md. [x] I searched using keywords relevant to my issue to make sure that I am creating a new issue that is not already open (or closed). [x] I reviewed the Discussions, and have a new and useful enhancement to share. Feature Description Lllama.cpp version b10121 Summary It would be extremely useful if the llama.cpp WebUI could render inline images that are referenced insid”
