I use gpt image 2 very heavily for my current project (ai UI design tool). The biggest improvement I'm seeing with this is in speed. I've generated around 50k images with gpt-image-2 via api, the the average latency has held at around 104s.
It's wild how much of a difference this is - images are coming in at around 35-40s. Very noticable, and makes a difference when you're iterating quickly: https://jjcm.org/gpt-image-2.5-speed.mp4
Overall it used the reference images I gave it a bit better than gpt-image-2. I noticed 2 had issues getting the blue button just right. 2.5 nailed it.
Did very well modifying the pose while keeping the appearance of Glenn Powell. gpt-image-2 had a lot of the "fried" look for some of his skin in prior designs I did for johnpolitics.com
Dark mode sites surfaced the fried look quite a bit in prior models, but this definitely looks better on that front. One thing that looks perhaps worse though is the microglyphs - note the teal lines to the bottom right of the ramen, they're kinda blurry / not straight.
Overall fixed some of the main issues / gripes I had with gpt-image-2
It's both. I have several models loaded into diffui. Which one each prompt uses is determined based on user preference over time. Whichever image is currently at the top of any image node stack marks a win for the model that generated it, I assign each model an ELO score based on that, and I bias the chance each model is selected based on their ELO score. Right now gpt-image-2 is better than my own model, and it services around ~96% of the requests in diffui.
I'll also be adding in microsoft's mai-image-2.6 soon, but I need to update my SOC2 to add MS as a provider before I turn that on for other users. The full list of models in rotation is here: https://image.non.io/d53a9760-8b74-4386-b032-d59da2cd5319.we...
Oh, you're the dev of diffui? It's completely off-topic, but I've notice that RMB -> Download image will download an .png but it's actually a .webp and it can cause issues (e.g. file explorer doesn't show thumbnail correctly).
I've really enjoyed using AI to generate images. For example, Long time ago, I read an amazing five-book saga called Riverworld, and I used AI to recreate many scenes, places, and ideas from the story. Seeing the books come to life through hundreds of images was a great experience. Definitely one of the coolest things to enjoy in 2026. Try it with your favorite books.
I just finished the 2nd book of the Stormlight Chronicles by Brandon Sanderson. Sadly I have had aphantasia all my life and therefore I can not visually imagine things - I do have a very loud inner monologue though and can imagine the voices of the characters.
I just tried out the new version of ChatGPT to generate images for the main characters in a concept art style and for me it makes the story come to life a bit more.
I also run a Pathfinder Campaign every other week and have found great pleasure visualizing scenes in this way for myself. Sometimes I share them with the players and so far the feedback on that has been very positive.
I wish I could picture things in my mind but at least now I have something that can help me with my handicap!
If it’s a good film then those are works of art and interpretation in themselves, but I suspect you have something which is more the product of a mechanical bureaucracy in mind than works of art: maybe like a production of Harry Potter or The Hunger Games.
If those are the kinds of books you have in mind I can see why you get confused/frustrated by how others feel about slop, since you’ve been shoulder deep in it for your whole life!
FWIW, there’s a famous(-ish) fantasy author with aphantasia (Mark Lawrence), and I also have aphantasia, yet I absolutely love reading and if I would have to choose only one entertainment for the rest of my live, it’d be reading.
But I can’t stand elaborate descriptions (it’s almost always landscapes, rarely do authors feel the need to describe anything in as excruciating details as landscapes), which might be related to those descriptions not building up to some kind of image in my head.
This is quite common if you have full on aphantasia.
I have a friend who only reads fiction for which he's seen the accompanying show first because he otherwise has no visual reference for the characters and environments.
> This is quite common if you have full on aphantasia.
Is it? I’ve never really researched common aphantasia behavior, but I’ve seen many mentions of aphantasia in fantasy/scifi subreddits where those of us with it still love reading. I can imagine (heh) it, but I wonder if there’s some kind of data regarding it?
This is a nice showcase of what it can do but so little of it seems to be actually useful or rather odd choices. Some examples were already mentioned, but also the Gdocs presentation: apart from the terrible layout, why would one use fake photos of the sun in a "science" talk when there is an abundance of real high-res photos available? For educational purposes these are even free to use.
Omg, I love how the first examples just show how easy you can fake things. Fake being at a party with your friends. Didn't make your bed, no problem, just fake it.
The sad part is my mother would love "remixing" my old child photos of me.
Seriously, are these really the best examples they could come up with?
What is the point of having a fake picture of your dog in a costume? What’s the point of having a fake picture about being at a party?
The only use case I can think of for this is for someone who likes to make up stories and lie about what they’ve done. Is that really the target market?
In a world were there is a significant amount of people that do things just to take photos of themselves doing things rather than experiencing the thing, this may be the lesser evil.
Trying to be as least cynical as possible, imagine if you were a party with 3 or more people, and you really wanted a nice pic like that to commemorate the night, but never got to take one all together because someone left early, and you were in the bathroom when they announced their departure.
Yes, it's a major first-world problem to not get a photo, but to the right person, it could mean a lot to them to "fix" the photo they took with A and B to add C to their 'rightful place.'
Is spending time with your friends not enough in 2026? Do we must provide photographic evidence of all the fun you have been having or it did not happen?
You’re mistaking your preferences for the majority’s preferences. Huge chunk of humanity now chooses to experience in-person events recording it through their phones.
So many people I just stopped spending time with altogether as instead of being present they’d be posing for their 90th instagram shot by 10am - and then of course they want you to pose, too.
So no. Spending time with your friends is not enough in 2026. I’m not even convinced people have friends now - just colleagues and mutual parasites.
I think it depends on who your friends are, and/or culture.
I was in France the other month and walking down to the river in Saumur, I noticed an area serving drinks (out of an old converted bus) with lights and suchlike (opposite the castle on the hill), with a small acoustic band. Nobody (and I mean nobody) had their phone out. The couple of hours I spent there, nobody was observable taking selfies. The age range was from young to old.
I was shocked. It was very different from the UK. Refreshingly so.
> The only use case I can think of for this is for someone who likes to make up stories and lie about what they’ve done. Is that really the target market?
As far as I can tell, the biggest uses of these sort of models is to pump out lots of content on multimedia-based social media, so yeah, it's basically for the sort of person who doesn't shy away from exaggerating, omitting or outright lying in order to get more views.
I’m less concerned about faking being at a party with friends and more with the 2006 timestamp on said photo!
Can you imagine finding a stack of photos in the basement with timestamps of the Before AI times and wondering whether they are real or just got swapped with generated and printed fakes? Scary!
It’s cool, folks… nothing bad is going to happen. Right? Right?!
> Omg, I love how the first examples just show how easy you can fake things. Fake being at a party with your friends. Didn't make your bed, no problem, just fake it.
Well, that's like the entire point of social media anyway, isn't it? People there will _love_ this. Ugh ...
My secret hope is that this kills the influencer industry and gets things back to just ads. Just the ads happen to be fake people, instead of real people pretending to be ads.
For thousands of years people have told embellished stories, and humanity has thrived upon it. Tall tales, fish stories. Heck, that's still the average person's experience unless they've really worked their critical thinking muscles.
Today's working adults are used to the short thirty-year "safe space" of smartphones and internet. We grew up in a temporary meta stability where "truth" was "recorded" and could be "relied upon", and now that the fundamentals are shifting, we're complaining that the physical world is amenable to storytelling and imagination once again.
Cry me a river. This is awesome and is a direct consequence of everything I ever wanted the future to be: magical creative superpowers. I wanted to graduate into the world that is emerging now rather than spend the first third of my career in incrementalism and slow progress.
2008 - 2020 sucked. What a total lull. AI is healing these things and putting us back on track for the jet pack future we grew up dreaming about. It's taking us back to a creative world without shitty platforms controlling what we say and do, and without a ceiling on what we can accomplish.
It feels like every day we're unwrapping a new present or several. Not small things, but reality-shattering things that fill me with inspiration to build and explore. It's so much fun.
To tell a story, to even exaggerate one is something completely different from providing "prove" it did really happen exactly the way it was told.
There are many many people that will take an image as an undeniable fact of the real world. While before you could fake(photoshop) things it took more time then 30 seconds.
I wouldnt regulate these aspects, because I believe pandoras box is already wide open. Regulating anything wouldn't change much.
Pandora's box is already wide open? The lid was blown off the hinges long before AI. There were people eating Tide Pods in the 2010s. People were drinking bleach and taking Ivermectin to kill a virus about five years ago.
AI, in some ways, might be our only saving grace at this point. It sure as Hell cannot cause people to do much worse.
Because she will pounder in old memories that never existed in the first place. Instead of cherrying the current times. Past moments of long gone time will be elongated beyond their actual existence.
Imagine a time where a small moment is actually a smaller amount experienced then the remixed one. At one point you will have more memories of fake events, and more emotional beats for said fake events than real ones.
AI videos showing you a second life if you just had chosen a different path.
I made an AI video of a picture of me when I was about 2. Had a huge smile on my face in the photo so as a lark I made a 5 second video of me laughing. The look in my mom's eyes when I showed her and the way she grabbed the phone from me to look closer was such a good moment.
It's not an alternate reality. I was laughing at the time almost certainly. New ways of seeing old memories really needn't be depressing at all. It's so strange how so many people seem to only see the bad in a technology that can be used in countless different ways.
There’s a wonderful Ted Chiang short story called “Anxiety is The Dizziness Of Freedom” that explores how people become enmeshed and addicted to alternate versions of how their life could have gone. It’s reminiscent of what you’re describing.
Because the mother's joy should hinge on the actual photo of the actual event of the actual person, not a fictional imaginary event or fake childhood of their child. The experience of being human should be tethered to reality.
Because elderly people who never had a connection to tech or sci-fi will have no feel for, and therefore no resistance to the kind of trickery that seems fun but is ultimately corrosive. Cheerfully meddling with the artifacts of memory at that life stage is corrosive.
The "composite party photo", while impressive, shows that still the miniscule details are being lost, like the teeth structure of the guy in the middle or the fact that the guy on the left is holding the cup with three fingers. Wondering why they chose this edit for the showcase.
I've also found the OpenAI image models to lose fine detail on image edits compared to Nano Banana or Flux models which faithfully retain input source image geometry and details. I was hoping this might be different but it sounds similar to previous OpenAI image models where something is lost in translation during image editing.
I know the API (assuming you're using it) lets you set the output size of the edited images - I haven't done a huge amount of testing with > 1mp, so I'd be curious if this might mitigate some of that.
Even with LM Arena being flawed, this is significant. I was planning to do a writeup on the original gpt-image-2 as it crushed every complex image comprehension benchmark I had...I'm glad I procrastinated since ChatGPT Images 2.5 seems like an even better starting point to test out what these models can actually do nowadays.
Many people still think AI images output the wrong number of fingers on a regular basis. (EDIT: this was an ironic comment to make in hindsight and I own it)
So this is a valid point (and I admit I eat crow on my earlier statement), but not for that reason. In that photo, there are three fingers in front of the cup, but you would not expect 5 fingers because the way humans hold cups, the thumb will be occluded by the cup itself. That said, I don't think there is a way to hold a cup with both thumb and index occluded, so the correct number of fingers would be 4 in that case.
If you zoom in, the index finger is just obscured by the middle finger. You can also see the knuckle, so it's not incorrect, just an awkward way to hold a cup.
There's maybe something cool at the core of this - help you ideate by seeing things in images! - but it's lost in all the other awful ways this gets used. More fake menus with food that looks nothing like the real thing. More fake book covers displacing real artists.
I can admire and enjoy the work of coding agents, which mostly just help solve problems and save me time. I can appreciate that there are creative uses for agents with text, that they could save people time, and that the environmental and safety concerns potentially have solutions.
I just can't see my way to appreciating these art agents, or any use of AI as a replacement for a human creative process. It's not because their output is bad, anymore (though sometimes it still isn't great). It's because it is bereft.
Even if I could, nobody in my life would be okay with my using them alongside my actual creative work. In fact, they'd be pretty upset if I did. And if I found out they'd been sharing stories with me written by AI, I'd be mad, too.
AI art is making me realise that through my whole artistic career, the reason it's so hard to convince people that they need good design is because no one really appreciates good art and design other than people who have studied art and design.
I am building an app and I need illustrations. Hiring someone is out of the question.
Previously, I would either not have illustrations or try to fit a stock image to my use case.
Now I can just generate illustrations that fit my use case well, and it takes very little effort.
Mine is just one out of thousands of conceivable use cases for image generation and editing. I appreciate that this capability is now available to everyone.
As a software user, I feel just as frustrated and disappointed when I look at a projects website, code, and commit history and it’s slop, why in the world would I use this garbage when I can just prompt my model to make my own?
Realize the people using this to make slop menus and book covers feel the same way, wow this is useful and saves me so much time/money!
I'm a solo coder. I love coding. Never used (or trusted) AI to generate code, only to review.
But good art and especially animation, beyond basic UI and logos, is a full-time commitment that requires giving up coding while you're working on the visuals (yes many great solo devs did their own art, but I prefer using the extra time on other shit, like making music and getting on with life)
Getting a good meat-based artist to hop on my Awesome Game Idea #458235 is easier if I can show them a running gameplay example, so they can see what the actual game will be like.
That works well but carries a risk of "painting yourself into a corner"; the longer you test a game in a particular art style (or music) the more you subconsciously tend to develop the rest of the game around that style! At least for me
AI-generated PLACEHOLDER art lets me iterate on a game closer to how I envision the final product to look like, right from Day 1.
I'd still hire/kidnap an actual artist before publishing.
Did anyone else notice that all the use cases they show here are "use ChatGPT Images to imagine all the things you'd like to have, but don't"? For some, they then show the thing being done, since just imagining it is a bit sad. But ChatGPT Images can't help you with that part. At best, it's providing creative inspiration - but that's often the most rewarding part of the process for a human... None of the examples show things where you just need an image directly, e.g. product labels, signage, sprites, textures, images for websites etc. I guess aspirational stuff sells?
In order, there's:
- Imagine you had better decorations for your fish.
If someone shows me a picture of a haircut, or a dress and asked me if it would look good on them, I can't imagine it. My mind does not work like that. I just can't visually see it.
There's no way to win in that situation anyway - if you said "yes I think it'd look great!" and then they went and got their hair cut to match your suggestion but didn't like it, is it your fault for saying "yes"?
The Facebook Marketplace experience for used items has depreciated considerably with the widespread adoption of LLMs. Between placing attractive models in photos to help sell items to "improving" the visual condition of items that completely misrepresent the real condition of the item, it's becoming a trickier landscape to navigate to find what you're looking for.
I just saw a Twitter thread of someone calling out realtors with absolutely egregious editing examples. Think pools and landscaping being added to backyards, a background being changed to a mountainous region, and furniture being “staged” in an attic space that definitely couldn’t fit that furniture.
A realtor in the comments said they legally need to also post the same images but unedited, but I’m sure when you’re doing this the first 50 are the edited ones and the last 50 are the unedited. How many people make it to the unedited?
I mean, it would probably be better than many tattoo artist’s art. At the very least it’s not a bad way to see how something would look on you before you permanently do it.
Look at the kid in a costume. The shoulders of ChatGPT’s output are just wrong. Sure, the kid looks prettier, stronger with a wider upper body, closer to society’s ideal of a child. But it’s not him. If you put that child in a suit, he’d still have a smaller upper body and the arms wouldn’t fill out the suit as nicely as the child in the AI image does.
It’s just not the same child.
I know we’re gonna get there eventually, but still, seeing as this is their first demo image on the page I just can’t help but scream inside “HOW CAN YOU NOT SEE THIS?”
It really makes me question whether the people building, or at least the people marketing this, actually understand their product. I love image generation for all kinds of use cases, but I find that particular example creepy.
You can get semi-decent sprite work out of GenAI models, but you still have to put in some manual work (scale normalization, palette reduction, grid alignments, etc). It's definitely not "out-of-the-box" yet.
Yeah, I think you’re right. I think the hardest part would honestly be maintaining some level of consistency as your project grows its number of assets - otherwise there's no sense of a common stylistic theme.
The agentic tooling in some of the latest models is getting pretty crazy too - somebody on Reddit recently shared an example of using GPT-6 Astra to generate sprites with Aseprite, an open-source, popular pixel art editor and it worked shockingly well.
The key is making them one frame at a time, rather than asking for the whole sheet. Adherence frame-to-frame with a reference image is really good, so just prompt with the previous + direction. I finally got Nano Banana to make fairly decent fluid animations that way.
I've been struggling with this for a few years for Yoto LED icons. Design rules didn't cut it, but a pymagick workflow did, more or less.
I have abandoned the work I was doing, but if you provide the image constraints along with some examples and descriptions of those examples using the terminology of your constraints, you should get something iterable.
Interesting. I'm surprised they haven't prioritised/deliberately trained for this more because of how useful it is to ask codex to generate some images/assets/sprites when building sill games. It would dramatically enhance how polished it's games were.
That said for single images the old model was already okayish for prototypes
Should probably mention this as a root comment, but various local image gen models like fooocus [1] run on basically any remotely modern hardware, have phenomenal output quality, and cost basically nothing.
I think companies trying to sell image gen at this point are just banking on ignorance.
Even GPT Images 2.0 was already the best image generator that I have tried, however it has one massive disadvantage: draconian censorship. Even rendering a gun in a scene will get flagged as "violent" and I need to create a battle scene... there's no way to do that with the censorship the way it is
What kind of person cares enough to buy flowers for an arrangement, but is uninterested in arranging them? What's next, will we buy music for our AI's to listen to? Food for them to eat?
There's a lot of drudgery that can be automated away, but this seems to be focused on automating the fun parts away.
I believe the next step will be from Douglas Adam's "Dirk Gently" book where an electric monk exists to believe things for you, to save you the effort and "bother" of having to have religious belief.
Surprised to see no acknowledgement of how "AI Menu Slop" has become deeply associated with ChatGPT.
Everywhere I go now I see places that blatantly used ChatGPT for their menus or posters and they all look the same.
It feels almost existential to the service, I know its a bit of survivor bias but so many images you can tell immediately are ChatGPT vs other image providers, and I feel many people are sick of them.
It's not like a lot of these small-time outfits used their own photography to begin with.
They'd pick a kebab from a menu of professionally made kebab pictures the printer has in stock. Or worse, they'd take pics of their own plated food with a dead-centre point flash in a dark cupboard or something and you end up with awful-looking food, no matter how good.
I'm not sure if I'm working with ChatGPT Images 2.5 or 2.0 here...
The original photo I took was at https://www.reddit.com/r/Tovala/comments/1pfuwrb/meatloaf_pa... - it's a cheeseburger meatloaf with potato wedges taken with a phone camera. And I was going for a consistent documentation approach for the photograph, not trying for menu proper.
I suspect that someone doing a menu could take a properly plated meal from the kitchen (rather than me photographing on top of my oven) and have it get redone for a good image for a menu without fundamentally changing what is being served.
I think quite a lot of people can distinguish between AI and real (maybe 30%?) but almost no one is worried about distinguishing between ChatGPT and other image generators
What if I told you.. the average person LIKES ai slop images and prefers them to "regular" ones.
The same way the average person prefers McDonalds food to healthy food
It's the great desemantification machine: want to look, read, sound, feel just like the median, without personal traits or individual expression? This is for you! (And people used to talk about communism and everything being same. Well, tech-broism actually achieved this.)
Can someone recommend a correct way for interior design ideas?
Yesterday tried GPT-6-Astra Light to have a kitchen remodel... well it did understand 2D space, an improvement when tried on GPT-5.6-Sol
But the image representation was way off and I was not satisfied... misplaced the window, drew a wall which was not on 2D image, cabinet size issues. Eh.
None of the improvements are particularly useful for my principal use-case. Creating infographics to visualize how things interact and so on. But it’s always nice to see improvement here.
One of my locally owned fried rice places doodled a simple pencil cartoon of them cooking fried rice to satisfy their evil landlord and it has cascaded into a series of mildly unhinged pencil cartoon doodles every few days on Facebook. It is an incredibly refreshing break from the local food/eats FB groups being utterly flooded with identical looking slop posters and has earned my business several times. Also it's just... incredibly in-touch marketing in general.
I can't disagree more. A lot of shops in our city have done this and I actually miss the shitty photoshopped images that did not even look like the real life dish anyway. AI signs are way worse.
I thought that was parents point, that currently they're using so shitty AI generated logos, that even if we despise that as a concept, at least better models output slightly less sloppy shit. But re-reading it, I'm not sure that was the right initial reading.
I am wondering if in 2-3 years AI will be heads and shoulders above us in both images and text. Then will we continue to see the slop accusations? Maybe its slop is better than most of our work.
Yes, even though fast food places would previously use every food photography trick in the book to make the food look good, there was at least some relation between the menu and what you actually get, which people just seem completely fine with dispensing because the affordances of image generators work against including product photography
(Have also heard about people - not just scammers/spam operations - using generated pictures and descriptions for dating profiles and real estate ads, and I cannot understand what outcome they expect when someone follows up and immediately realizes what's up)
> "Here is a picture of the dish, make it look nicer and use that as a reference"
That's not a function of improvement in models, though, which was the question. You could literally do that today, no need to wait for models in "2-3 years".
Not OP but my guess is that they're already using bad AI images for their kebab business, and OP wants them to use better AI for their images (so they may actually lose fewer customers because the AI imagery isn't so awful)
Why do you care what your kebab shop uses for their menus? These people are trying to run a business. If this can allow them to clearly communicate prices in a well designed pleasant way, it's great. Why must they slave away in design and technology or pay someone to do it for them? I think it's great
If I eat at a restaurant and that restaurant has images of their food, I want those to be “real images” of what they will bring me when I pay. Otherwise it’s misleading.
Not sure exactly what images we're talking about here, places above "medium-scale" typically don't have images on their menus at all, and mostly have real iamges on their Instagram and website, of course professionally taken and edited though.
I see a ton of AI generated images and descriptions on DoorDash. I get the sense the restaurant is not doing it themselves, I’m not sure if they have any control over disabling it or perhaps they are opting in.
It must be, I'm not on any plan (free user) and got a modal on ChatGPT.com that let's me use it (for 3 images a day). Should be much higher for Plus users.
Surprised there is no mention of the Sketch feature so far. That's very powerful! I do art on the side and I know no better of getting the result I want than passing in a first draft sketch.
I super encourage others to learn just a little bit of art technique and your images are going to go to the next level. You need a little construction, perspective, gesture, anatomy, and you're pretty much off to the races.
Knowing names of illustration styles is great too. Look up "style transfer" if you are not familiar.
ChatGPT Images is for me the most impresive of all models, but also the one I hate the most, because of what it's used for: either funny wasteful use, or nefarious use, basically nothing else.
Just like WordPerfect enabled your mom to create a professional-enough cover letter, ChatGPT allows your mom to send you a more or less visually pleasing virtual birthday card of you wearing a party hat.
Earlier today I couldn't get ChatGPT to draw a horse on the moon. Idea taken from Stable Diffusion's wikipedia page. I tried three times with that prompt, "a photograph of an astronaut riding a horse" and again with a horse on the moon. I got it to do another image. However, the original prompt just worked, though it didn't happen to choose the moon, it didn't say the moon.
That’s surprising even gpt-image-1 was able to handle the equestrian astronaut prompt pretty trivially, albeit yellow-tinged as all hell.
For the heck of it, I ratcheted up the difficulty of the prompt (two-headed horse, astronaut with the visor up, etc.), and gpt-image-2 got it right every time.
Wow, that's amazing. I was about to tell it to make the helmet similar to the human one, with the gold visor, but thought it wouldn't convey it as well, but not only is it obvious, it still has the feeling of being a horse.
Coming to HN to read comments on releases like this reminds me that there is a small, vocal minority that is pushing so hard for more and more features like this.
The majority of the world does not want or need any of this, yet the nerds in SF that can hardly hold a conversation with another human being are pumping it out as quickly as they can.
Please don't post unsubstantive comments to HN, and particularly not unsubstantive + negative ones, which kill the kind of discussion we're hoping to have here (i.e. curious and thoughtful).
What? He quotes the opening line of the article, and notes that he reacts negatively to it. Interpreting his sentiment as "curmudgeonly" is a clear exaggeration.
And sure--it isn't a high-substance comment, but it isn't irrelevant either. It was, for a while, the top-voted comment on this whole thread. That shows people resonate with the thought, and it obviously sparked quite a lot of discussion.
This seriously makes me question your, and our other moderators, motivations. Sad times.
This was not a borderline call! The GP comment is not in keeping with the kind of discussion we're hoping for, especially not at the top of a thread—which unfortunately is where cheap negative comments often end up.
The problem isn't primarily with the comment, and in that sense conradludgate did nothing wrong. The problem is the upvotes. If it weren't for those, such comments wouldn't rise to the top and choke out more interesting conversation. But they routinely do, and that's one of the biggest problems we have here.
This i did like 15 images just today, mostly just editing and fixing one of my wifes photos that she wanted the lighting and stuff leaned up with so i decided to test out 2.5... it did a great job, but that was easily 15 images 1 of which was the end result but all 15 count toward that figure, i'd honestly guess its more 1:100-1:500 that actually makes it outside the chatgpt servers
I assume the comment I am responding to was made in sarcasm, so I'm basing my response on that register, so apologies in advance if I have got this wrong.
I don't understand what is dystopian about this: With the exception of mildly elevated resource consumption to achieve the end goal (a nicer photo), is this not the sort of things that humans have been wanting and doing for as long as photographs have been a thing?
It was previously mostly unachievable for regular people to have their snapshots processed in such a way, but it feels like it's just making what used to be special mildly trivial. Anyone who has had their wedding photographed for the past 20 years has had the photos touched up on the way through. Any school portrait for just as long.
100 years ago people would dodge and burn photographic prints to improve them as much as the technology allowed. This is just the forward evolution of that old tale.
Maybe slightly odd is how anyone could possibly think I was being sarcastic. :)
I'm not attacking you here or anything :)
I'm just a bit shocked at how cynical people can be. The thought would not cross my mind ... and I'm a skeptic of many things!
There's a video somewhere of some people in Wyoming, riding horses.
All of the comments are young people shock of these 'boomers' committing 'animal abuse'.
I mean ... just ranchers riding horses.
People think that 'Ward Cleaver' form 'Leave It To Beaver' is some 'creepy twilight zone figure' ... when he's just literally a nice guy. Or rather the character is.
We're getting way to jaded!
So - I really mean it: AI is going to be mostly very pedantic stuff.
I witnessed the birth of computers and the internet - and while AI is definitely a big deal ... I think it's just going to feel like 'more computing stuff' in the end.
Computers started out for Science - then it became about video games, socialization, design, commerce, and of course 'gambling and porn' and all the 'same vices'.
Every step of the way there were people thinking it was being 'debauched' etc..
AI is going to be about the most mundane things, just as most technology is today.
It will be boring and in the background, and probably weird like a poorly designed, slow web-page full of cruft, but it will 'do the thing'.
That's also what my real life social circle does. My friends and family test out home reno/decoration ideas, and as I posted in another comment, my wife likes to visualize different layouts and changes to her flower gardens.
The 'environmental' consideration is much less than the argument over 'paper or plastic'.
That's going to be an issue of regulation and cost - not some arbitrary decision over 'whether I should make an image or just use text'.
Second - there are really not much in the way 'social or moral implications' for the most part.
This is a dystopian delusion that I'm hinting at here.
Folks are blowing this way out of proportion, and cynically so.
People now have 'automated Photoshop' and a slightly more power to express themselves. It's not going to change that much, and it's not a bad thing, it's a good thing. Mildly.
It has been possible to create convincingly cheap image misinformation for over a year (Nano Banana last August) but civil society has not collapsed. There are also open-weight models of similar quality and functionality without any safeguards.
Checks and balances in trust still exist, believe it or not.
The "society has not collapsed" goal post seems odd to me.
More and more people are getting caught up in weird echo chambers, extremist groups, and spewing misinformation. The amount of pure slop on Facebook, either for engagement or manipulation is so high, there is more fake content on FB than real content. The conspiracy side of me believes the CIA is deliberately forcing us against each other, but even if no underlying entity is behind this, the same outcome is happening. This is a byproduct of both Social Media & Generative AI.
In my country, Canada, the number one supplier of fake news and misinformation is the government news company, the CBC. It is an ocean compared to the fake news I've ever seen generated by chatgpt.
It is 100% not. You can create a whole lot of misinformation with an image. You cannot do that by visiting a website. To dismiss the concern with “people are having fun” is, at best, a naive idea and, at worst, a deliberately deceitful and malicious comment.
"which anyone can use to do, relatively easily" this is not true at all. Not anyone can use Photoshop, you need to first buy or get the software, then you need to learn how to use it, then work on the images. Whereas now you can just use Grok and immediately create pedophilia. Another deliberately deceitful and malicious comment
I hate it because I hate having to look at these AI generated images. In fact I want an internet where I have to explicitly opt-in to see AI generated content in general.
I was watching a friend stream a dead game lately.
That game was developed before generating images became widespread, and sunset as that "art form" started really taking off (game-as-a-service... probably partly reverse engineered and retooled with AI too). So during its lifetime, it was pretty much AI free.
I remember having been seduced by that game's art style at the time. It wasn't "the usual guys with futuristic guns and armor", it had some stylization as well, nice color...
Watching the stream, it occurred to me that it was in fact artwork from "the days before", in a way I hadn't been hit with until that point, even as a returning artist looking at pre-AI works.
"People came up with the concepts and developed and rendered them themselves", I thought to myself. Not in a "oh look, floppy disks. How quaint!" way, either.
Today I got stuck with a hairy frontend problem, which even AI couldn't help me with. Luckily, I managed to find a Medium article from like 2019. It was joyfully absent of 'comparisons that actually hold up' and 'X that is only a part of the picture' and nothing was load-bearing in there.
I forgot that once there existed a magical time when human beings used to write Medium articles by hand.
I would imagine a lot of it is mundane shit like getting ChatGPT to render the same room with different coloured walls or provide inspiration on plant pots. I have a decent amount of image usage and it's all things like this, not generating content I intend to distribute.
Yeah seriously, the cultural mindset is so damn kneejerk negative to everything right now. It's pathetic. This is an exciting time, and lots of good things are happening.
Maybe? Transformers and neural networks in general can do lots of interesting things for sure. But an image with the default ChatGPT aesthetic mindlessly produced by a generic prompt isn't one of them.
Seriously, it’s so fucking lame and boring to read these cynical takes. I might just make a HN comment that filters these people out, they’re ruining the site.
To be clear, I don’t mind a well-reasoned negative take on something, but much of the negatives takes I see here these days isn’t much better than Reddit, just hyper-exaggerated bitching and moaning about everything.
The weirdly narrow and cynical purview of your example proves my point.
It echos sentiments of painters, intellectuals and odd religious moralists and other rabble rousers who spent decades raving against 'tHe eVIL pHoToGRaPH' in late 19th and early 20t century.
People can now make more expressive images than they could before, which is generally a good thing.
Of course, people could always use new technology for negative purposes - and - people have been using Photoshop for decades already to make 'fake nudes' - this is not a new phenomenon.
None of this is even a very big deal. We can make images now much more easily, that's it.
> imagine being someone being blackmailed by photos created by ChatGPT Images
If anything, the fact that I know for a fact that I can, in 45 seconds, generate an image of any arbitrary person holding Solo cups at a party arm in arm with Vladimir Putin (to crib an idea from one of the demos) should make everyone less and less susceptible to blackmail in general - in fact, even if the source material is real! There was a time 40 years ago when a photograph's existence was considered proof of a fact - the existence of Photoshop-like tools made that into a weak proof, and the existence of genAI has made "a photograph" into a format that is known to be as malleable as a .txt file.
Whether we like that or not, no one will stop bad actors from attempting to abuse this technology, so democratizing its use does help people to understand how they work and make potential victims much less credulous.
this is an oddly dismissive comment that says more about the commenter then what is commented on.
generating an ai image is vastly more costly than loading a webpage…but still not that costly compared to most other things.
this is depressing in the same way content slop or fyp feeds are fun.
if you want a silly pic with ur kids, just pay someone on fiver. it will be only slightly more expensive and likely far more exciting for kids as they have to wait for the dopamine hit.
"if you want a silly pic with ur kids, just pay someone on fiver. it will be only slightly more expensive and likely far more exciting for kids as "
???
You mean to imply there's something 'more authentic' about paying $5 for a random Philippino guy to use Photoshop to mix-up some family photos with sticker overlays?
What are you even talking about?
On what level is that either more authentic, better for the environment, more novel ... of frankly better in any way?
That is 'not even an argument' - and so you're making my point for me:
People can now make images.
It's mostly for fun, some of it will be used professional, some Tweets will be a bit more visual - but that's it.
It's not going to change that much, and it's generally better on the whole for people to be able to create what they want.
People are putting this out of proportion.
Maybe I'm too old, but I've been around - this is not a big deal in the grand scheme.
>if you want a silly pic with ur kids, just pay someone on fiver. it will be only slightly more expensive and likely far more exciting for kids as they have to wait for the dopamine hit.
Or, y'know, learn how to draw, or accept that it's not exactly going to be Chuck Jones caliber stuff.
Yea, “fun”, sure. I’m sure none of those billions of images were used to lie, mislead, or cheat anyone. I’m sure none of those images were used to create fake news narratives, used in advertisements to bait people into buying fake products, or create misleading mockups of rental properties. Definitely not.
Nevermind the fact that the examples in OpenAI’s blog post are examples of creating images that have no plausible purpose other than to fake stuff that didn’t happen.
Yesterday my wife used ChatGPT to generate an image of our cat as a toilet, which was very funny for the both of us. She's also used it to turn him into a basketball getting dunked, a tsetse fly, a clown and a napkin. It's hard to say whether she uses it more to generate inside joke images about our cat, or to generate mockup plans and layouts for her flower gardens, but it is safe to say she does it all for fun.
The sheer waste that accompanies image and video genAI is actually terrifying.
I challenge OpenAI to put a "your carbon footprint" field next to each generation. If you have nothing to hide, more information is surely better, right?
Every SDXL image uses something like 0.5 Watt-hour. That means that 2000 sdxl images uses about 30 cents of electricity. If we multiply that by 30x for video (a fair assumption for models like minimax H3, which take that much longer to make video), that means you get almost 700 5-second videos for a dollar of electricity, which is about a kilo of carbon dioxide. More or less depending on where your electricity comes from.
> If we multiply that by 30x for video (a fair assumption for models like minimax H3, which take that much longer to make video)
It takes about 7 minutes for my RTX 5090 to generate a 15 second video with MiniMax H3. At 600 watts, that's 252,000 watt-seconds, or 0.07 kilowatt-hours.
My electricity is 18 cents/kWh, so $0.0126. Just a little over a penny.
Someone needs to name this fallacy. The existence of existing harmful norms does not, ex ante, justify introducing additional harmful norms. I propose naming it the "GenAI Defender Fallacy".
Separately--yes, I would love carbon receipts for everything. Such a system would be fascinating and do real good in this world.
> The existence of existing harmful norms does not, ex ante, justify introducing additional harmful norms
The point isn't to justify anything.
It's to put things into perspective, and give people a more informed and holistic point of view, and then to ask them to reconsider their opinions in light of more information. There's nothing bad about that. In fact it's a good thing.
Especially in a world where people are so easily misled by others who cherrypick unexceptional statistics and then intentionally present them as exceptional in order to generate outrage, retweets, clicks, whatever.
It's a combination of argument from analogy and reductio ad absurdum. It's not a fallacy. You're free to dispute the validity of the analogy, or you can dispute the absurdity, which you did, by confirming that you would indeed love carbon receipts for everything.
What's the goal of complaining about it though? What's more likely, you persuade people to stop all of these habits that use electricity that burns carbon, or you convert the electricity supply to renewables?
Whataboutism isn’t a fallacy and is widely misused online.
If someone says “I think x should stop because of y”. It is a valid argumentative response to say “y also occurs with z, so should z also stop?”. It is argumentatively revealing either hypocrisy, or that the original imperative relies on additional, unstated arguments/assumptions/biases.
Maybe I’m confused, but I thought whataboutism was more like “you are saying I am doing X bad thing? What about Y bad thing you’re doing, let’s discuss that instead”.
It’s a deflection technique, not a check on logical consistency.
> “you are saying I am doing X bad thing? What about Y bad thing you’re doing, let’s discuss that instead”.
no one is generally expecting or even asking for "let’s discuss that instead". what they are doing is trying to point out that your lack of of care about Y implies that you never actually cared about X in the first place but rather you only care about me doing X.
often times X is actually bad but human nature is such that no one actually cares about it while wanting to appear to when locked down on it. for this reason arguing about X directly is bad optics.
whataboustism is basically an effort to draw attention to the selective enforcement of norms, laws, morality, whatever. hypocrisy is implied.
It can also be used to derail a conversation. I.e. "I find subject X unpleasant, so I am going to imply that you don't care enough about subject Y, which has a passing similarity to it, so that the subject of the conversation turns to Y rather than X. I do not actually care about subject Y, but I won't say that out loud". See also "concern trolling".
Considering that especially leftists have spent literal decades hating on vegans for telling them about their individual footprint, there is no real fallacy here.
AI haters have somehow discovered personal responsibility. Awesome. They've never ever applied that to anything else they PERSONALLY do, though. It's only things others do. For any personal choices, suddenly it's all the governments fault or the corporations fault. No personal responsibility to be seen.
My 5-ish years of veganism have easily made up for a life time of prompting. Yet, I still would never generate video or images in part because of their impact.
How many of you non-vegan AI haters are going to be vegan now?
This is such a strange reply I don't even know what to do with it. Good on you for being vegan, and not prompting?
I have gone lacto-ovo vegetarian in the past year, so I'll take your 5 years of veganism as inspiration. Otherwise, I welcome people waking up to the harm they cause, regardless of their past hypocrisy.
Being flawed isn't the issue. Everyone is flawed. The issue is blaming people for problems like environmental damage, while at the same time needlessly causing more damage yourself. The problem is being hypocritical. Some people are certainly more hypocritical than others.
I mean, I actively advocate for better fuel efficiency standards but also haven't entirely eliminated red meat from my diet. If I'm discussing fuel efficiency standards and you call me hypocritical, even if I granted that, it doesn't seem like a useful turn of the conversation for anyone unless it's a rhetorical move to change the topic of conversation?
Advocating vs blaming. Blaming people for using inefficient vehicles when you eat red meat (assuming you eat enough red meat to have more of an impact than the fuel) is hypocritical. Advocating for better alternatives to inefficient vehicles while eating red meat is not hypocritical.
I've gone chicken only. Mostly because pork and beef are too expensive. It is very hard to get 150+ grams of protein a day on a vegan diet. And butter is in almost every delicious baked good. Avoiding it is like having the worst allergy.
Bad examples. Use meat consumption. It will blow any other kind of "non-essential" resource use out of the water, on ALL fronts.
It's really hard to believe that people care about the personal impact of their decisions, when it's only ever other peoples decisions that get talked about.
Meat is particularly ironic because it uses an obscene amount of water. Even worse, a lot of the feed for cattle comes from the imperial valley in CA, which has first dibs on the Colorado River, which is on the verge of collapsing. Obviously they aren't cutting usage.
But datacenters using water is what the public needs to fear...
No one was talking about the ethical part (which you may or may not agree with). Carbon footprint is an actually measurable quantity that we can compare - what's your problem?
that would be actually pretty cool and would cause more awareness for a lot of people. but something like this would go in both directions i am afraid, as with everything there surely would be people trying to do carbonmaxxing just for the sake of it, kind of like breaking a highscore. i am sure people like this already exist anyway, though.
It is because you know this and you continually choose to keep using it.
It’s be like if your 2001 BMW car was super fuel efficient, and then for whatever reason the 2028 model guzzled gas and you chose to buy it anyway
Just because the thing you did once upon a time was efficient, does not mean you have a lifetime pass to engage with it no matter the updated circumstances
reddit used to shared how much gold purchases went towards their server costs (pretty sure it wasn't 100% accurate but at least good to an order of magnitude). It was approximately $1/hour and they served ~100M monthly active users. So I don't think those metrics would demonstrate how wasteful traditional data centers are like you think it would.
this would be amazing, absolutely. Especially if they were able to see specifically where my power comes from, the embedded carbon in the chipsets used to serve my requests, all the costs associated with the datacenter(s), etc. A weekly breakdown would be excellent.
You must have your units really messed up because there's no way that's right. A GPU drawing 700W for 10s would be 1.9444 Wh or 0.0019444 kWh and that's if the GPU was only serving your image gen for the whole 10s which is unlikely, most of the time is probably queuing since even a much lower powered GPU in a desktop PC can generate the same image in well under 10s.
Do you have a source for 1-2.5kWh for 10s of content? It takes about a minute or two to generate, so you'd be talking about a GPU consuming 30_000W-150_000W, it's just not possible. Even if it was distributed (which I don't think it is), that'd be 30-100 datacenter GPUs running at 100% to generate one clip? There's no way that would make financial sense.
I propose they use my body for compost to power their electricity turbines, in exchange for a advance in tokens while I am alive...I am overweight and that is good, more burning mass. We could settle at 2 million USD.
A round trip cross country plane ride for one passenger is equivalent to about 500k - 1 million image generations. So all one has to do to offset their lifetime usage of image generators is forego one vacation. And this of course is only for current day carbon usage, that could go down (or up I guess but compute-wise usually things get cheaper over time).
The waste overall for data centers is beyond terrifying. Enough water for 1.2 billion people. It’s like all the GenX and Boomers looked at the challenges younger people will face and decided make it as bad as possible start kicking them in the ribs for good measure.
I know of entire media departments that went from - let’s hire an artist and see what they come up with to - let’s produce hundreds if not thousands of images to explore all possible ideas (and end up with something bland anyway).
Do you want to know how many identical questions ChatGPT gets asked every week, recomputing its answer each time because people stopped sharing results? I don't, but I'm sure it far exceeds 3 billion. "How to center a div", "how to install this and that package", "what is this pop culture reference", "who is this famous person",...
I ask a question to google and i spend 10-100 times longer visiting multiple websites making multiple servers generate webpages wasting electricity and time reading them finding my answer, trying different solutions
Sounds like a complex question. On average you'll visit like 1.5 sites. Those SQL queries are peanuts compared to an LLM burning through tokens. Besides, for hard questions, the LLM itself might download 30-60 websites to find the info on your behalf, e.g. checking the state of RAM prices.
Ok I'll bite, how do you reuse the answer to "what color is the sky?" while also layering in memory, custom system instructions and custom response styles?
It's probably only memory that'd need an answer for some form of caching to be worthwhile (though I'd be curious what that answer is) since memory is on by default and rarely identical between users.
The rest are nice to haves, but not everyone customizes settings just because they are there so there will be some large pool of users with the defaults who could hit cache without them.
I’m doing renovation . I slap a prompt in a tile store over my room and design several types of tiles and paint from one prompt while causally browsing in almost real time .
Isn’t that the future ? I generated 200 images alone in this use case .
What’s the alternative to that ? Send it tomorrow someone who hate their job moving tiles in the toilet , may money , wait 1 week
Groan. Your comment is the most depressing thing I've read today. What a sad way to look at the world. People are doing stuff and having fun. Meanwhile a whole sub-cult is sucking from the doom straw and mumbling carbon, water, greed, blah. The sport of Golf consumes far more resources than any AI data center.
We need more green spaces, we need people to be active. Golf is a fairly quite sport, animals do foster around the greenery with golf courses. I would much rather a golf course than hot concrete everywhere.
I'm not hating on golf. My grandfather would be appalled! In crowded urban areas they don't work so well anymore. And I'm also not against using our resources for human pleasure!
its because "green spaces" and "spaces with greens" are not the same. just imagine a 150 acre generic park or forest instead of an over fertilized over watered field for a single game played by only .002 percent of people.
golf is generally a huge waste of space for the amount of value we get from it per person and should have no place anywhere near a population center where we actually need more green space.
Golf courses are not green spaces, come on now. They are most often limited in their access and devoid of greenery native to the area or that supports pollinators. Grass does not equal green space, and certainly not the flavor of green that is present in golf courses. The alternative to a golf course is not hot concrete either, that's a false dichotomy.
the golf stat is actually way worse than this... because you aren't accounting for active users. how many people
actually benefit from all that water.
a golf course... at BEST has two or three foursomes on each hole on average at a time. so 36-48 players concurrently. accounting for less than perfect utilization it rounds to a rough throughput of about 1000 people per day per course. winter Exists and so you probably get like 20-50k people per year.
or maybe an easier term would be the estimated 160m people who played golf at least once last year.
which means that not only are we using that much water for so few people (.002% of the population) we are also reserving all of that space for so few people. for courses out in the middle of nowhere that doesn't matter but for example are thousands of acres of courses inside and near cities. one in LA reported 117k visits one year and that required dawn to dusk 365 play only possible in a climate like LA.
A typical course is ~150 acres so we could fit 50 soccer fields into that same space and service 3.2 million players. Ill spare the envelope math but its pretty hard to come up
with any other use for that space/water that could serve less people.
ok one more example that doesn't need water either. just up the street from the high volume golf course is runyon canyon. similar size 150 acres, serves 1.5m visits a year.
Basically golf is a major waste of resources for very little benefit in basically every way possible, and should be zoned out of existence.
Yes. And right after we kill golf lets pave over Central Park. And we can tear down sports arenas and concert venues after that. It's a huge waste when you think about how few people actually go there and what benefit they actually have. Sports overall are a waste of time. Everyone should just walk in the woods instead. Except soccer. You're right that we need thousands more soccer fields. Its way more efficient than any other kind of exercise. And there is so much demand for soccer in the US, you can't even find an empty field. In contrast to golf courses which are mostly empty like you noticed.
you aren't countering my core point that golf courses is the possibly the worst use of space and basically anything else would serve the same or better utility for way more people.
heck keep it golf but switch to driving ranges and its still like 10x more capacity for users than golf courses.
counter arguments that point out other things that are still 10x better uses of space than golf are not very compelling. a sports arena that serves 2.5m people/year is a terrible argument against a golf course that maxes out at 117k. (using sofi in LA numbers as its a similiar footprint as well 150acres)
im also not saying abolish golf, just only play it in places where the space/water use dont matter. like rural scotland. don't put golf courses in/near cities where we should use our space better.
I’m not commenting on whether generative AI is especially resource-wasteful or not, but a lot of people online seem to have newly discovered the existence of data centers.
That's fair and I left it as is knowing that. But it was still factually correct.
The main point is people worrying about resources and how to allocate them. We are creatures who want joy and I will not apologize for that. Certainly in face of abstract things such as "the environment". Which environment? I am a humanist. That definition varies. Golf is not the problem sorry to pick on it!
I was curious and tried to come up with a crude estimate - all of the golf courses in the US use approximately the same amount of energy (mowing, fertilizer, etc) as one 1GW DC does.
K, but they're paying for that energy though. And the land and the building costs, etc. I really don't want to live in a country where the government just arbitrarily says "Oh, sorry we have 'enough' datacenters / Subway sandwich shops / golf courses / storage facilities so you aren't allowed to build one" -- even if it becomes a populist belief that they should.
The fact is, at this point there is so much demand for DC capacity that the prices are super high, and thus it's worth building them. Many people believe that it's a bubble or whatever. If they're right, well then a lot of people building DCs right now will lose their shirts. If they're wrong and the demand is sustained, DCs are exactly what we need to be building.
>where the government just arbitrarily says "Oh, sorry we have 'enough' datacenters / Subway sandwich shops / golf courses / storage facilities so you aren't allowed to build one" -- even if it becomes a populist belief that they should.
It's not really arbitrary though, is it? It's decided by duly-elected representatives who answer to the citizens of the municipality. A government "of the people", etc.
I feel like what you're arguing against is precisely why we have local governments in the first place.
Well, you could definitely do policies like that in a plain democracy or republic without any guaranteed rights of the citizens. In our particular constitutional systems, and many others, the founding documents say that the government can't do certain things, even if they pass a law saying they can. (Of course, the Constitution can be amended to remove rights, but there is an intentionally high bar to meet in order to do that.)
Congress could pass a law tomorrow, or California could pass a ballot measure with a 100% Yes vote, saying (for instance) that some random company now belongs to the government, or to Gavin Newsom personally, or to me, but the courts would still be obligated to strike it down since that grossly violates the Seventh Amendment. Ideas like this are commonly referred to as being checks against "mob rule" - the system recognizes that 'what the people want' isn't necessarily supreme when it runs up against other people's recognized rights.
Is there such a thing as a 1GW data center? Like a beefy gaming PC easily uses 1 KW According to Google, an individual GPU tower consumes anywhere between 10-100kW. An 1GW data center is can fit in a size of a mid size apartment.
Lol. Let's not pick on golf? All sports are pointless depending on how you look at them. But here we are thousands and thousands of years later and humans want what humans want.
to be fair you need to create a dozen images to get a decent one. Many times i need to get to almost what i want using paint.net and hope chatgpt removes the aliasing/cleans up the pixel boundaries without messing anything up - i'm not complaining though, still faster and cleaner than photoshop -- when it works
This is really depressing to me. I think this is abusive.
Your son has no concept of the implications of his body, face and identifying features being used here. He has no ability to opt-out, refuse consent, and avoid his data (biometric features) being swooped up in a data-centers for further training, further data collection & further monetization.
Your son didn't agree to the privacy policy of OpenAI. This is an extremely high level of disrespect to a person you created.
If you're not familiar of the pitfalls of this tech, and the fact that there are more ethical tools with which to do those things, I'm not sure what to tell you.
Some people feel the need to attach an image to almost everything the post - even in chat rooms (including slack at work). These images serve no purpose, but the poster feels it is useful - though it was often reaction gifs in the past I'm seeing a lot more AI content than I ever saw reaction gifs, maybe because of novelty or maybe because it's ultra personalised.
Sometimes images help to visualise something or to get a point across, but I see so many people who think it's necessary to reply to a discord message with a cat with human limbs doing a dance, or a photo of "themselves" climbing a mountain with the Rust logo to show them mastering Rust... Ok?
Image models are useful and I'm thankful for much better visual reasoning but it really really frustrates me the constant need to burn money for all of this slop.
>Generating 1,000 images with a powerful AI model, such as Stable Diffusion XL, is responsible for roughly as much carbon dioxide as driving the equivalent of 4.1 miles in an average gasoline-powered car. In contrast, the least carbon-intensive text generation model they examined was responsible for as much CO2 as driving 0.0006 miles in a similar vehicle
But 3 billion pictures per week is equivalent to 12.4 million miles or 639.6 million miles extra per year.*
It's not about you. There are no 3 billion people each creating one picture per week. Its probably more like 10.000 assholes creating 2.5 billion pictures per week to satisfy some stupid online feed and the other users creating a picture per month on average.
It equals around 0.004% of actual driven miles if my estimates are right.
That is with an estimated ~270 billion miles vehicle miles every week (I put commercial in there.. about ~200 billion miles if you only include passenger vehicles.)
That study is from 2023 when the cost to produce 1000 images was 2.91 Wh per image. More recent numbers for Stable Diffusion's 2025 models that puts it at 1.3 Wh per image and that number continues to decrease.
For context that means a days worth of image generation emits about the same amount of carbon dioxide as a single transatlantic flight (New York to London). There are approximately 1500 transatlantic flights per day.
So every day the people of Vermont emit more CO2 just commuting than all image generation through OpenAI per week. Nice. That does put it in perspective. It’s really efficient.
yep, what our climate really needs is another carbon emitter similar to an entire state's worth of vehicles just so people can not pay artists or make dumb images of themselves ripping off some artistic style like Studio Ghibli. that is so much more of a social good than people being able to make it to work and earn a living
Oh I think paying people to make 3 billion images would result in far more carbon emissions. Just think about how much CO2 a human emits while painting. The paint, the food. Just for the sake of the climate I would never pay an artist for something like this.
for sure, because every one of those 3 billion images rendered by human hand and not automated with a tool would exist. we wouldn't have a magnitude smaller number of focused designs, of course, that's definitely not how art and design processes happen
Ours cars are atrocious. We can have the most powerful artificial minds imagine 1000 images from simple prompts, and that takes as much energy as moving a human being 4 miles.
And that will get more energy efficient. Gas cars have barely budged.
This is an embarrassingly idiotic article. They're comparing against the least intensive text model, ie some million parameter model nobody uses. Stable Diffusion is a 3.5 billion parameter model, while the GPT models a billion people are using for text generation are over 10 trillion parameters. To say nothing of the differences in average context size usage. The actual ratio of SDXL to text generation pollution is probably literally reversed from what the article claims by lying with statistics.
(Note, however, that OpenAI's image model is much much larger than SDXL; however, we don't have precise numbers for it. Nonetheless, misinformation is misinformation.)
I wouldn't/don't really feel comfortable uploading private pictures to any other cloud provider either :). I find OpenAI especially untrustworthy because of their known business practices, unestablished business model and unknown future capabilities and role in society. You cannot take that data back once they have it. The condescending tone was unasked for, apologies to OP for that.
Not sure you can compare huge ML clusters of GPUs running models in order to do these images, and the typical "meme generator" which is basically a call to imagemagick passing an image and some text.
I use gpt image 2 very heavily for my current project (ai UI design tool). The biggest improvement I'm seeing with this is in speed. I've generated around 50k images with gpt-image-2 via api, the the average latency has held at around 104s.
It's wild how much of a difference this is - images are coming in at around 35-40s. Very noticable, and makes a difference when you're iterating quickly: https://jjcm.org/gpt-image-2.5-speed.mp4
Some UI tests with it:
Warcraft 3 style agentic dev interface: https://image.non.io/cd9ea5cd-8ed7-44e0-ad3f-480ff0e51875.we...
Overall it used the reference images I gave it a bit better than gpt-image-2. I noticed 2 had issues getting the blue button just right. 2.5 nailed it.
A "John Politics" meme site: https://image.non.io/d2922164-fa96-4d07-a141-2febadb02939.we...
Did very well modifying the pose while keeping the appearance of Glenn Powell. gpt-image-2 had a lot of the "fried" look for some of his skin in prior designs I did for johnpolitics.com
A cyberpunk inspired ramen website: https://image.non.io/8d5d8f10-0f0f-4d91-b5ea-33af7538b150.we...
Dark mode sites surfaced the fried look quite a bit in prior models, but this definitely looks better on that front. One thing that looks perhaps worse though is the microglyphs - note the teal lines to the bottom right of the ramen, they're kinda blurry / not straight.
Overall fixed some of the main issues / gripes I had with gpt-image-2
The WC3 one is fun. Definitely feels close to what I remember.
>Warcraft 3 style agentic dev interface
I can forgive the randumb placement of chains but not that sovlless icon of a person from the wow era. wc3 era UI would've used a character portrait.
Love the John Politics page, utterly slick completely bland and phoney. Assume it’s a meme i missed, love it.
I'm confused. Isn't diffui using its own model?
It's both. I have several models loaded into diffui. Which one each prompt uses is determined based on user preference over time. Whichever image is currently at the top of any image node stack marks a win for the model that generated it, I assign each model an ELO score based on that, and I bias the chance each model is selected based on their ELO score. Right now gpt-image-2 is better than my own model, and it services around ~96% of the requests in diffui.
I'll also be adding in microsoft's mai-image-2.6 soon, but I need to update my SOC2 to add MS as a provider before I turn that on for other users. The full list of models in rotation is here: https://image.non.io/d53a9760-8b74-4386-b032-d59da2cd5319.we...
Oh, you're the dev of diffui? It's completely off-topic, but I've notice that RMB -> Download image will download an .png but it's actually a .webp and it can cause issues (e.g. file explorer doesn't show thumbnail correctly).
Indeed I am! Also nice catch - should have a fix up in the next hour. Hit me up with any other reqs - j@diffui.ai
Edit: confirmed that the download should be fixed now - should be properly a png now.
I've really enjoyed using AI to generate images. For example, Long time ago, I read an amazing five-book saga called Riverworld, and I used AI to recreate many scenes, places, and ideas from the story. Seeing the books come to life through hundreds of images was a great experience. Definitely one of the coolest things to enjoy in 2026. Try it with your favorite books.
[flagged]
I just finished the 2nd book of the Stormlight Chronicles by Brandon Sanderson. Sadly I have had aphantasia all my life and therefore I can not visually imagine things - I do have a very loud inner monologue though and can imagine the voices of the characters.
I just tried out the new version of ChatGPT to generate images for the main characters in a concept art style and for me it makes the story come to life a bit more.
I also run a Pathfinder Campaign every other week and have found great pleasure visualizing scenes in this way for myself. Sometimes I share them with the players and so far the feedback on that has been very positive.
I wish I could picture things in my mind but at least now I have something that can help me with my handicap!
[flagged]
If it’s a good film then those are works of art and interpretation in themselves, but I suspect you have something which is more the product of a mechanical bureaucracy in mind than works of art: maybe like a production of Harry Potter or The Hunger Games.
If those are the kinds of books you have in mind I can see why you get confused/frustrated by how others feel about slop, since you’ve been shoulder deep in it for your whole life!
[dead]
[flagged]
Then reading a book is outsourcing imagination or enjoying it's illustrations by someone else, as you rely on the others for imagining stories
About 300 million people literally can't imagine.
We’ve outsourced everything else.
Why is that such a bad thing? Imagination is not truly useful for anything without skills to complement it, at least not in my experience.
>Imagination is not truly useful for anything without skills to complement it
techbro final boss
Outsourcing re-imagination.
[flagged]
I wish I could do it better. There seems to be a lot of variability. This episode of radiolab on the topic is fascinating: https://radiolab.org/podcast/aphantasia/transcript?utm_sourc...
The apple scale:
https://www.reddit.com/r/HelloInternet/comments/f0wuej/are_y...
Wow I haven’t thought about HI in years!
I feel like the reason I dislike reading or find reading boring is because I can't. But it's just a theory. Meanwhile I absolutely love movies.
FWIW, there’s a famous(-ish) fantasy author with aphantasia (Mark Lawrence), and I also have aphantasia, yet I absolutely love reading and if I would have to choose only one entertainment for the rest of my live, it’d be reading.
But I can’t stand elaborate descriptions (it’s almost always landscapes, rarely do authors feel the need to describe anything in as excruciating details as landscapes), which might be related to those descriptions not building up to some kind of image in my head.
This is quite common if you have full on aphantasia.
I have a friend who only reads fiction for which he's seen the accompanying show first because he otherwise has no visual reference for the characters and environments.
> This is quite common if you have full on aphantasia.
Is it? I’ve never really researched common aphantasia behavior, but I’ve seen many mentions of aphantasia in fantasy/scifi subreddits where those of us with it still love reading. I can imagine (heh) it, but I wonder if there’s some kind of data regarding it?
Not all people can!
I did it as well. But it was fun to imagine it through another type of eyes.
[dead]
One man's treasure is another man's trash
One man's cliche is another man's hacker news comment
This is a nice showcase of what it can do but so little of it seems to be actually useful or rather odd choices. Some examples were already mentioned, but also the Gdocs presentation: apart from the terrible layout, why would one use fake photos of the sun in a "science" talk when there is an abundance of real high-res photos available? For educational purposes these are even free to use.
Omg, I love how the first examples just show how easy you can fake things. Fake being at a party with your friends. Didn't make your bed, no problem, just fake it.
The sad part is my mother would love "remixing" my old child photos of me.
Seriously, are these really the best examples they could come up with?
What is the point of having a fake picture of your dog in a costume? What’s the point of having a fake picture about being at a party?
The only use case I can think of for this is for someone who likes to make up stories and lie about what they’ve done. Is that really the target market?
In a world were there is a significant amount of people that do things just to take photos of themselves doing things rather than experiencing the thing, this may be the lesser evil.
Trying to be as least cynical as possible, imagine if you were a party with 3 or more people, and you really wanted a nice pic like that to commemorate the night, but never got to take one all together because someone left early, and you were in the bathroom when they announced their departure.
Yes, it's a major first-world problem to not get a photo, but to the right person, it could mean a lot to them to "fix" the photo they took with A and B to add C to their 'rightful place.'
What's the point of having a photo like that?
Is spending time with your friends not enough in 2026? Do we must provide photographic evidence of all the fun you have been having or it did not happen?
You’re mistaking your preferences for the majority’s preferences. Huge chunk of humanity now chooses to experience in-person events recording it through their phones.
Did you miss the last 15 years of social media?
So many people I just stopped spending time with altogether as instead of being present they’d be posing for their 90th instagram shot by 10am - and then of course they want you to pose, too.
So no. Spending time with your friends is not enough in 2026. I’m not even convinced people have friends now - just colleagues and mutual parasites.
I think it depends on who your friends are, and/or culture.
I was in France the other month and walking down to the river in Saumur, I noticed an area serving drinks (out of an old converted bus) with lights and suchlike (opposite the castle on the hill), with a small acoustic band. Nobody (and I mean nobody) had their phone out. The couple of hours I spent there, nobody was observable taking selfies. The age range was from young to old.
I was shocked. It was very different from the UK. Refreshingly so.
and then as you look back at that wonderful night you "remember" with all your friends, you will reinforce a false picture and reality strays further
This is such a contrived imaginary situation that is pretty divorced from reality. No one wants a fake AI photo to commemorate a real event.
But still what's the point? To look back on false memories that didn't happen?
That was exactly my first reaction!
Mark where you at the party today? Yes of course look at these pictures I took (╥﹏╥)
> The only use case I can think of for this is for someone who likes to make up stories and lie about what they’ve done. Is that really the target market?
As far as I can tell, the biggest uses of these sort of models is to pump out lots of content on multimedia-based social media, so yeah, it's basically for the sort of person who doesn't shy away from exaggerating, omitting or outright lying in order to get more views.
Didn't make your bed, no problem, just fake it.
Except you need to have taken a photo of the unmade bed, uploaded it to ChatGPT, prompted for the 'fake made bed' version, and downloaded the image.
Surely it's less effort to just make the bed?
It would use far less energy to make the bed too, and help preserve Earth's future by avoiding needlessly expending energy.
You can take the photo, inside the chatgpt app and then talk to it without much effort. Perfect for a lazy teenager .
> Surely it's less effort to just make the bed?
It's not about the bed.
People areaddicted to generating now.
Slopoholics the whole bunch of them.
Not in the past I guess?
Yeah, the "3 lonely selfies -> fake 2006 party photo" is the most 2020s thing I've ever seen in my life.
I’m less concerned about faking being at a party with friends and more with the 2006 timestamp on said photo!
Can you imagine finding a stack of photos in the basement with timestamps of the Before AI times and wondering whether they are real or just got swapped with generated and printed fakes? Scary!
It’s cool, folks… nothing bad is going to happen. Right? Right?!
True, I didn't thought about that. Just imagine the scale of "historic" images that will be created over the next 100 years.
There won’t be a next 100 years.
> Omg, I love how the first examples just show how easy you can fake things. Fake being at a party with your friends. Didn't make your bed, no problem, just fake it.
Well, that's like the entire point of social media anyway, isn't it? People there will _love_ this. Ugh ...
My secret hope is that this kills the influencer industry and gets things back to just ads. Just the ads happen to be fake people, instead of real people pretending to be ads.
Is that 5 pillows on the original (messy) bed down to 4 pillows in the clean version?
Apple Intelligence received the same criticism when they launched Siri AI. Seems like the faking it, is a common product UC.
And the fan favourite 'Imagine the apartment we're trying to sell doesn't look like a superfund site'
This is an artificial modern world problem.
For thousands of years people have told embellished stories, and humanity has thrived upon it. Tall tales, fish stories. Heck, that's still the average person's experience unless they've really worked their critical thinking muscles.
Today's working adults are used to the short thirty-year "safe space" of smartphones and internet. We grew up in a temporary meta stability where "truth" was "recorded" and could be "relied upon", and now that the fundamentals are shifting, we're complaining that the physical world is amenable to storytelling and imagination once again.
Cry me a river. This is awesome and is a direct consequence of everything I ever wanted the future to be: magical creative superpowers. I wanted to graduate into the world that is emerging now rather than spend the first third of my career in incrementalism and slow progress.
2008 - 2020 sucked. What a total lull. AI is healing these things and putting us back on track for the jet pack future we grew up dreaming about. It's taking us back to a creative world without shitty platforms controlling what we say and do, and without a ceiling on what we can accomplish.
It feels like every day we're unwrapping a new present or several. Not small things, but reality-shattering things that fill me with inspiration to build and explore. It's so much fun.
To tell a story, to even exaggerate one is something completely different from providing "prove" it did really happen exactly the way it was told.
There are many many people that will take an image as an undeniable fact of the real world. While before you could fake(photoshop) things it took more time then 30 seconds.
I wouldnt regulate these aspects, because I believe pandoras box is already wide open. Regulating anything wouldn't change much.
Pandora's box is already wide open? The lid was blown off the hinges long before AI. There were people eating Tide Pods in the 2010s. People were drinking bleach and taking Ivermectin to kill a virus about five years ago.
AI, in some ways, might be our only saving grace at this point. It sure as Hell cannot cause people to do much worse.
AI makes it easier to fact check. It’s not perfect but it’s better than nothing.
This is the right perspective.
While {AI, the Internet, Smartphones} can cause harm, {AI, the Internet, Smartphones} are vastly more beneficial to society than not.
[Cry me a River]
No thank you
Good, because we need that river to make AI images!
>The sad part is my mother would love "remixing" my old child photos of me.
How is that sad? Why is your mothers joy sad?
Because she will pounder in old memories that never existed in the first place. Instead of cherrying the current times. Past moments of long gone time will be elongated beyond their actual existence.
Imagine a time where a small moment is actually a smaller amount experienced then the remixed one. At one point you will have more memories of fake events, and more emotional beats for said fake events than real ones.
AI videos showing you a second life if you just had chosen a different path.
People will be depressed from it.
I made an AI video of a picture of me when I was about 2. Had a huge smile on my face in the photo so as a lark I made a 5 second video of me laughing. The look in my mom's eyes when I showed her and the way she grabbed the phone from me to look closer was such a good moment.
It's not an alternate reality. I was laughing at the time almost certainly. New ways of seeing old memories really needn't be depressing at all. It's so strange how so many people seem to only see the bad in a technology that can be used in countless different ways.
There’s a wonderful Ted Chiang short story called “Anxiety is The Dizziness Of Freedom” that explores how people become enmeshed and addicted to alternate versions of how their life could have gone. It’s reminiscent of what you’re describing.
Because the mother's joy should hinge on the actual photo of the actual event of the actual person, not a fictional imaginary event or fake childhood of their child. The experience of being human should be tethered to reality.
Because elderly people who never had a connection to tech or sci-fi will have no feel for, and therefore no resistance to the kind of trickery that seems fun but is ultimately corrosive. Cheerfully meddling with the artifacts of memory at that life stage is corrosive.
Why do you assume her joy is the sad part?
Memory already declines with age. AI slop threatens to warp those memories further.
The "composite party photo", while impressive, shows that still the miniscule details are being lost, like the teeth structure of the guy in the middle or the fact that the guy on the left is holding the cup with three fingers. Wondering why they chose this edit for the showcase.
I'm guessing if you work on these images the whole day you start to lose the ability to judge image quality. like your brain becomes slopified
> Wondering why they chose this edit for the showcase.
They probably didn't care to check those images in detail.
I've also found the OpenAI image models to lose fine detail on image edits compared to Nano Banana or Flux models which faithfully retain input source image geometry and details. I was hoping this might be different but it sounds similar to previous OpenAI image models where something is lost in translation during image editing.
I know the API (assuming you're using it) lets you set the output size of the edited images - I haven't done a huge amount of testing with > 1mp, so I'd be curious if this might mitigate some of that.
https://developers.openai.com/api/reference/python/resources...
The OpenAI PR/doc team seem to have a history of this kind of attention lapse - their 4-panel comic strip using gpt-image-2 was an absolute mess too.
https://mordenstar.com/blog/gen-failures
Another example was the doc for GPT-5, which had all its graphs messed up, but they said it was human error rather than AI.
And don't forget the 'Thing' from Addams Family in the shoulder of the guy on the left.
The guy on the left in that same photo only has three fingers (and a thumb, I suppose). I thought image generation has already outlived that.
Edit: I feel stupid I didn't see the original OP already mentioning three finger issue. I'll just leave it here.
The dog is even more on the nose. The shadow shape is almost the same despite the dog getting quite a lot of volume around it's original body parts.
Haha yes, I immediately saw the fingers and was surprised 'cos I thought that was something they had "solved" by now
Maybe it's like how spammers intentionally make their emails more obvious, because they have a very specific target audience.
Exactly this. "All this progress and they're back to miscounted fingers again?!"
The guy’s arm is also extremely long.
Dogs shadow didn't change
A big wtf at the LM Arena scores: https://arena.ai/leaderboard/text-to-image
gpt-image-2.5-sunburst: 1421
gpt-image-2.5-flare: 1399
gpt-image-2 (medium): 1381
mai-image-2.6: 1331
Even with LM Arena being flawed, this is significant. I was planning to do a writeup on the original gpt-image-2 as it crushed every complex image comprehension benchmark I had...I'm glad I procrastinated since ChatGPT Images 2.5 seems like an even better starting point to test out what these models can actually do nowadays.
Many people still think AI images output the wrong number of fingers on a regular basis. (EDIT: this was an ironic comment to make in hindsight and I own it)
> Many people still think AI images output the wrong number of fingers on a regular basis.
There's literally an image of a dude with 3 fingers in the Composite Party Photo.
So this is a valid point (and I admit I eat crow on my earlier statement), but not for that reason. In that photo, there are three fingers in front of the cup, but you would not expect 5 fingers because the way humans hold cups, the thumb will be occluded by the cup itself. That said, I don't think there is a way to hold a cup with both thumb and index occluded, so the correct number of fingers would be 4 in that case.
If you zoom in, the index finger is just obscured by the middle finger. You can also see the knuckle, so it's not incorrect, just an awkward way to hold a cup.
> Many people still think AI images output the wrong number of fingers on a regular basis.
Last time I used a frontier image model it made me a seal with three hands so…
I mean there's literally one of their example images with hands with the wrong number of fingers.
There's maybe something cool at the core of this - help you ideate by seeing things in images! - but it's lost in all the other awful ways this gets used. More fake menus with food that looks nothing like the real thing. More fake book covers displacing real artists.
I can admire and enjoy the work of coding agents, which mostly just help solve problems and save me time. I can appreciate that there are creative uses for agents with text, that they could save people time, and that the environmental and safety concerns potentially have solutions.
I just can't see my way to appreciating these art agents, or any use of AI as a replacement for a human creative process. It's not because their output is bad, anymore (though sometimes it still isn't great). It's because it is bereft.
Even if I could, nobody in my life would be okay with my using them alongside my actual creative work. In fact, they'd be pretty upset if I did. And if I found out they'd been sharing stories with me written by AI, I'd be mad, too.
I'd rather just see the prompts.
AI art is making me realise that through my whole artistic career, the reason it's so hard to convince people that they need good design is because no one really appreciates good art and design other than people who have studied art and design.
It’s okay to make nice things for a small group of people, the challenge, obviously, is finding those people and exposing them to your work.
And getting paid…
I am building an app and I need illustrations. Hiring someone is out of the question.
Previously, I would either not have illustrations or try to fit a stock image to my use case.
Now I can just generate illustrations that fit my use case well, and it takes very little effort.
Mine is just one out of thousands of conceivable use cases for image generation and editing. I appreciate that this capability is now available to everyone.
As a software user, I feel just as frustrated and disappointed when I look at a projects website, code, and commit history and it’s slop, why in the world would I use this garbage when I can just prompt my model to make my own?
Realize the people using this to make slop menus and book covers feel the same way, wow this is useful and saves me so much time/money!
Notice that not a single business on earth so far is passing those time and money savings on however.
These AI tools primarily benefit single person businesses and they are creating a race to the bottom so that's not entirely true
I'm a solo coder. I love coding. Never used (or trusted) AI to generate code, only to review.
But good art and especially animation, beyond basic UI and logos, is a full-time commitment that requires giving up coding while you're working on the visuals (yes many great solo devs did their own art, but I prefer using the extra time on other shit, like making music and getting on with life)
Getting a good meat-based artist to hop on my Awesome Game Idea #458235 is easier if I can show them a running gameplay example, so they can see what the actual game will be like.
Up until recently I was just using free third-party asset packs like Kenney's: https://kenney.nl/assets/1-bit-pack
That works well but carries a risk of "painting yourself into a corner"; the longer you test a game in a particular art style (or music) the more you subconsciously tend to develop the rest of the game around that style! At least for me
AI-generated PLACEHOLDER art lets me iterate on a game closer to how I envision the final product to look like, right from Day 1.
I'd still hire/kidnap an actual artist before publishing.
Did anyone else notice that all the use cases they show here are "use ChatGPT Images to imagine all the things you'd like to have, but don't"? For some, they then show the thing being done, since just imagining it is a bit sad. But ChatGPT Images can't help you with that part. At best, it's providing creative inspiration - but that's often the most rewarding part of the process for a human... None of the examples show things where you just need an image directly, e.g. product labels, signage, sprites, textures, images for websites etc. I guess aspirational stuff sells?
In order, there's:
- Imagine you had better decorations for your fish.
- Imagine you had a tattoo of your pet.
- Imagine you had a unique candle holder.
- Imagine you had a worse haircut.
- Imagine your flowers were arranged.
- Imagine your child had a suit.
- Imagine your dog had a costume.
- Imagine you owned a scanner.
- Imagine you could hang out with friends.
- Imagine your bed was made.
I think you've picked the most cynical interpretation for each one. More charitably, you could say:
- Ideate on things you can build for your fish tank
- Iterate quickly on ideas for a tattoo
- Get inspiration for a new metalworking (or 3D printing?) project
- Get a sense for what a new haircut might look like, before you commit
- Iterate on arrangements for a bouquet of flowers
- Do silly things with childhood photos
- Do silly things with pet photos
- Clean up low-res, damaged, or weirdly formatted photos
- Do some creative storytelling with your long-distance friends
- Okay yeah fine they should just make the damn bed, but I can imagine small tweaks being useful for staging a home (easy to abuse though)
This is fair criticism, I was being pretty harsh.
> Did anyone else notice that all the use cases they show here are "use ChatGPT Images to imagine all the things you'd like to have, but don't"?
What else could they highlight? The other major use case for these models for normal people (generating NSFW images) is explicitly blocked by them.
If someone shows me a picture of a haircut, or a dress and asked me if it would look good on them, I can't imagine it. My mind does not work like that. I just can't visually see it.
At least it saves them asking you now, I guess?
There's no way to win in that situation anyway - if you said "yes I think it'd look great!" and then they went and got their hair cut to match your suggestion but didn't like it, is it your fault for saying "yes"?
> None of the examples show things where you just need an image directly, e.g. product labels, signage, sprites, textures, images for websites etc.
I think this is because recently there has been extreme backlash to this sort of thing.
The Facebook Marketplace experience for used items has depreciated considerably with the widespread adoption of LLMs. Between placing attractive models in photos to help sell items to "improving" the visual condition of items that completely misrepresent the real condition of the item, it's becoming a trickier landscape to navigate to find what you're looking for.
I just saw a Twitter thread of someone calling out realtors with absolutely egregious editing examples. Think pools and landscaping being added to backyards, a background being changed to a mountainous region, and furniture being “staged” in an attic space that definitely couldn’t fit that furniture.
A realtor in the comments said they legally need to also post the same images but unedited, but I’m sure when you’re doing this the first 50 are the edited ones and the last 50 are the unedited. How many people make it to the unedited?
This is going to get regulated pretty quick.
The introductory video is a bit... daring. Generating tattoos that clearly look AI-generated is not one of the use-cases I'd try to sell.
It hit me just now that people might actually getting tattooed with AI slop. That’s… depressing
I mean, it would probably be better than many tattoo artist’s art. At the very least it’s not a bad way to see how something would look on you before you permanently do it.
Think on the bright side, it will be easier to judge them based on how they look, possibly saving you time and effort.
Great idea, get an AI art tattoo so you can always be reminded how AI art looked in 2026.
Lots of people have tattoos of what cutting edge video game graphics looked like in 1992.
Look at the kid in a costume. The shoulders of ChatGPT’s output are just wrong. Sure, the kid looks prettier, stronger with a wider upper body, closer to society’s ideal of a child. But it’s not him. If you put that child in a suit, he’d still have a smaller upper body and the arms wouldn’t fill out the suit as nicely as the child in the AI image does.
It’s just not the same child.
I know we’re gonna get there eventually, but still, seeing as this is their first demo image on the page I just can’t help but scream inside “HOW CAN YOU NOT SEE THIS?”
It really makes me question whether the people building, or at least the people marketing this, actually understand their product. I love image generation for all kinds of use cases, but I find that particular example creepy.
Still can't make sprite sheets :(
You can get semi-decent sprite work out of GenAI models, but you still have to put in some manual work (scale normalization, palette reduction, grid alignments, etc). It's definitely not "out-of-the-box" yet.
https://mordenstar.com/other/hobbes-animation
Really nice share. I bet you could automate the manual work with today's models
Yeah, I think you’re right. I think the hardest part would honestly be maintaining some level of consistency as your project grows its number of assets - otherwise there's no sense of a common stylistic theme.
The agentic tooling in some of the latest models is getting pretty crazy too - somebody on Reddit recently shared an example of using GPT-6 Astra to generate sprites with Aseprite, an open-source, popular pixel art editor and it worked shockingly well.
Shameless plug, but I built an image generation app with tools for sprite sheet and cutting out the sprites automatically.
Added support for GPT 2.5
http://batchbanana.com/
It's better to use a specialized model for that https://retrodiffusion.ai
>"Still can't make sprite sheets :("
The key is making them one frame at a time, rather than asking for the whole sheet. Adherence frame-to-frame with a reference image is really good, so just prompt with the previous + direction. I finally got Nano Banana to make fairly decent fluid animations that way.
No but it can write a tool that can.
Go on... I'm listening.
I've been struggling with this for a few years for Yoto LED icons. Design rules didn't cut it, but a pymagick workflow did, more or less.
I have abandoned the work I was doing, but if you provide the image constraints along with some examples and descriptions of those examples using the terminology of your constraints, you should get something iterable.
Interesting. I'm surprised they haven't prioritised/deliberately trained for this more because of how useful it is to ask codex to generate some images/assets/sprites when building sill games. It would dramatically enhance how polished it's games were.
That said for single images the old model was already okayish for prototypes
[dead]
Composite party photo...this is going to revolutionize tinder profiles.
Will be updating my dating app photo generator asap: https://apps.apple.com/us/app/pull-ai-dating-app-photos/id67...
The societal hurt coming out of public releases of image/video generation models must surely outweight the gain. That said, this is super impressive.
I've done a side by side comparison of the two new models across all quality levels. Versus all the previous models
https://generative-ai.review/2026/09/rush-openai-image-gen-2...
right, off to bed now
Still far too expensive.
At this point, I'm much more interested in seeing releases for models chasing down prices on "good enough" to unlock new use-cases than I am SOTA.
I've been using Grok Imagine V1 in prod for months now as its quality for my use-case is already more than enough.
Should probably mention this as a root comment, but various local image gen models like fooocus [1] run on basically any remotely modern hardware, have phenomenal output quality, and cost basically nothing.
I think companies trying to sell image gen at this point are just banking on ignorance.
[1] - https://github.com/lllyasviel/Fooocus
I can't stop thinking that these use cases aren't real, and OpenAI is just gaslighting people. Humans don't need AI for these trivial things.
This ad feels like a science fiction, where humans can't do anything without the help of their AI assistants.
- Design a tattoo of my cat
- What hairstyle should I have
- Arrange these flowers
What next? Who should I date? Open the door?
You're right that no one needs this, but OpenAI are right that many people will use it
the clothing examples are super odd because its not actually useable. the clothes will not look like that on you, it just fits them to your body.
Even GPT Images 2.0 was already the best image generator that I have tried, however it has one massive disadvantage: draconian censorship. Even rendering a gun in a scene will get flagged as "violent" and I need to create a battle scene... there's no way to do that with the censorship the way it is
Does it still have that distinct off-white shading of previous models though?
I'm 99% sure that that is the main tell AI sniffers rely on.
> that distinct off-white shading
Let's see Paul Allen's card.
allllll the way at the bottom
> Pricing and availability
> Images 2.5 is rolling out today to ChatGPT, ChatGPT Work, and Codex users across all tiers on desktop, mobile, and web.
> GPT‑Image‑2.5 Sunburst and GPT‑Image‑2.5 Flare are available in the API. See pricing details here.
and the link to the pricing details is a page where i am either too dumb to find the pricing details, or they don't exist
https://developers.openai.com/api/docs/pricing#image-tokens
Looks like they updated the link to that page now too. Thanks.
Same price as image-2, and I'm guessing it's not more tokens (given the speedup).
> Arrange these flowers
What kind of person cares enough to buy flowers for an arrangement, but is uninterested in arranging them? What's next, will we buy music for our AI's to listen to? Food for them to eat?
There's a lot of drudgery that can be automated away, but this seems to be focused on automating the fun parts away.
I believe the next step will be from Douglas Adam's "Dirk Gently" book where an electric monk exists to believe things for you, to save you the effort and "bother" of having to have religious belief.
It's very sad.
Surprised to see no acknowledgement of how "AI Menu Slop" has become deeply associated with ChatGPT.
Everywhere I go now I see places that blatantly used ChatGPT for their menus or posters and they all look the same.
It feels almost existential to the service, I know its a bit of survivor bias but so many images you can tell immediately are ChatGPT vs other image providers, and I feel many people are sick of them.
It's not like a lot of these small-time outfits used their own photography to begin with.
They'd pick a kebab from a menu of professionally made kebab pictures the printer has in stock. Or worse, they'd take pics of their own plated food with a dead-centre point flash in a dark cupboard or something and you end up with awful-looking food, no matter how good.
I'm not sure if I'm working with ChatGPT Images 2.5 or 2.0 here...
The original photo I took was at https://www.reddit.com/r/Tovala/comments/1pfuwrb/meatloaf_pa... - it's a cheeseburger meatloaf with potato wedges taken with a phone camera. And I was going for a consistent documentation approach for the photograph, not trying for menu proper.
https://chatgpt.com/share/6aa05d33-3e6c-83ea-9e32-f1371c1ad6... ( https://imgur.com/a/fPB5VKP for just the image)
I suspect that someone doing a menu could take a properly plated meal from the kitchen (rather than me photographing on top of my oven) and have it get redone for a good image for a menu without fundamentally changing what is being served.
It's still more honest than generating virtual kebab.
Is it though? I don’t find using stock art to be any better myself.
Stock art food doesn't usually look unappetizing and bizarre.
Sounds like a great self-selection problem.
I think quite a lot of people can distinguish between AI and real (maybe 30%?) but almost no one is worried about distinguishing between ChatGPT and other image generators
What if I told you.. the average person LIKES ai slop images and prefers them to "regular" ones. The same way the average person prefers McDonalds food to healthy food
It's the great desemantification machine: want to look, read, sound, feel just like the median, without personal traits or individual expression? This is for you! (And people used to talk about communism and everything being same. Well, tech-broism actually achieved this.)
Can someone recommend a correct way for interior design ideas?
Yesterday tried GPT-6-Astra Light to have a kitchen remodel... well it did understand 2D space, an improvement when tried on GPT-5.6-Sol
But the image representation was way off and I was not satisfied... misplaced the window, drew a wall which was not on 2D image, cabinet size issues. Eh.
None of the improvements are particularly useful for my principal use-case. Creating infographics to visualize how things interact and so on. But it’s always nice to see improvement here.
I hope my local kebab shops switch to this for their signs.
This is hilarious and it's killing me people are taking it literally
Sarcasm is the lowest form of wit, but...
One of my locally owned fried rice places doodled a simple pencil cartoon of them cooking fried rice to satisfy their evil landlord and it has cascaded into a series of mildly unhinged pencil cartoon doodles every few days on Facebook. It is an incredibly refreshing break from the local food/eats FB groups being utterly flooded with identical looking slop posters and has earned my business several times. Also it's just... incredibly in-touch marketing in general.
I can't disagree more. A lot of shops in our city have done this and I actually miss the shitty photoshopped images that did not even look like the real life dish anyway. AI signs are way worse.
I thought that was parents point, that currently they're using so shitty AI generated logos, that even if we despise that as a concept, at least better models output slightly less sloppy shit. But re-reading it, I'm not sure that was the right initial reading.
I am wondering if in 2-3 years AI will be heads and shoulders above us in both images and text. Then will we continue to see the slop accusations? Maybe its slop is better than most of our work.
> Then will we continue to see the slop accusations?
If it doesn't look like the food you're about to purchase, yes? Why is this even a question in this context.
Yes, even though fast food places would previously use every food photography trick in the book to make the food look good, there was at least some relation between the menu and what you actually get, which people just seem completely fine with dispensing because the affordances of image generators work against including product photography
(Have also heard about people - not just scammers/spam operations - using generated pictures and descriptions for dating profiles and real estate ads, and I cannot understand what outcome they expect when someone follows up and immediately realizes what's up)
"Here is a picture of the dish, make it look nicer and use that as a reference"
> "Here is a picture of the dish, make it look nicer and use that as a reference"
That's not a function of improvement in models, though, which was the question. You could literally do that today, no need to wait for models in "2-3 years".
"Users will voluntarily prompt better to avoid misleading people" will be the "people can simply write memory safe C" of the late 2020s
“Stop calling it slop you guys, I’ll have you know I am even sloppier”
It might be a sub-comment on a tweet that was fairly viral last week. Or it might be serious. We’ll never know.
If they wanna loose customers...
Not OP but my guess is that they're already using bad AI images for their kebab business, and OP wants them to use better AI for their images (so they may actually lose fewer customers because the AI imagery isn't so awful)
The normier the user the crazier the pics.
Why do you care what your kebab shop uses for their menus? These people are trying to run a business. If this can allow them to clearly communicate prices in a well designed pleasant way, it's great. Why must they slave away in design and technology or pay someone to do it for them? I think it's great
If I eat at a restaurant and that restaurant has images of their food, I want those to be “real images” of what they will bring me when I pay. Otherwise it’s misleading.
I don't think I've ever eaten at a restaurant with real images of their food, pre or post AI.
Not sure exactly what images we're talking about here, places above "medium-scale" typically don't have images on their menus at all, and mostly have real iamges on their Instagram and website, of course professionally taken and edited though.
I care, the same way I wont read your blog, if you use AI photos. Same mechanism.
It would be an improvement, probably, because they are still using Dall.e 1 apparently.
Loose them on what?
loose? only if they don't cook it well
I see a lot of the slop posters in real life now and I hate them
I think every restaurant on Uber Eats in my area uses AI slop or, minimally, AI "enhanced" food pictures.
I see a ton of AI generated images and descriptions on DoorDash. I get the sense the restaurant is not doing it themselves, I’m not sure if they have any control over disabling it or perhaps they are opting in.
Probably some other type of ads '[..] in your area' will use this as well.
I still don't know if I hate them because they are visually slop, or that the businesses farmed out sign creation to the slop machine
I love them because while not perfect it’s a good way to filter out businesses that don’t have good taste.
All of the above for me.
If it was only the small shops doing it... companies that could easily hire ad designers use AI slop images full of errors.
[dead]
I feel it's more like a Nano Banana 2 update. It's much faster than gpt-image-2 but the quality isn't much different.
The AI tattoo. Hope you don’t have any ragrets.
Stu Hamm (a bass player) has a massive tattoo down his arm saying "no regerts"...
"I don't know why all the other shops turned me down..."
Don't tap the aquarium glass man
I used it to visualise different arrangements of my furniture and it worked nicely. Personal interior designer
Is this included in the 20/month plan?
I had trouble weeding through all the marketing speak.
It must be, I'm not on any plan (free user) and got a modal on ChatGPT.com that let's me use it (for 3 images a day). Should be much higher for Plus users.
Did they fix the noise gradient?
doesn't look like it, looking at the motocross photo (0:19 in the "Structure your prompts for better results" video)
drink more water is a crazy thing to ai generate
I plan to try something for designing the UI of my website.
Surprised there is no mention of the Sketch feature so far. That's very powerful! I do art on the side and I know no better of getting the result I want than passing in a first draft sketch.
I super encourage others to learn just a little bit of art technique and your images are going to go to the next level. You need a little construction, perspective, gesture, anatomy, and you're pretty much off to the races.
Knowing names of illustration styles is great too. Look up "style transfer" if you are not familiar.
ChatGPT Images is for me the most impresive of all models, but also the one I hate the most, because of what it's used for: either funny wasteful use, or nefarious use, basically nothing else.
Just like WordPerfect enabled your mom to create a professional-enough cover letter, ChatGPT allows your mom to send you a more or less visually pleasing virtual birthday card of you wearing a party hat.
Do all Silicon Valley corporations use the same jingle for their product promotional videos? The creator must be really rich by now.
If they all use it, then it is free
I hate so much how every food menu is now AI slop. I wish that was outlawed under false advertising.
Ah. Another 3 months delay for gemini-3.5-pro.
wtf are these comments what happened to hn.
Where is Dall-e?
feels this something backhanded
Earlier today I couldn't get ChatGPT to draw a horse on the moon. Idea taken from Stable Diffusion's wikipedia page. I tried three times with that prompt, "a photograph of an astronaut riding a horse" and again with a horse on the moon. I got it to do another image. However, the original prompt just worked, though it didn't happen to choose the moon, it didn't say the moon.
Edit: finally, I have my horse on the moon. https://chatgpt.com/share/6aa0627e-b168-83e9-98e5-d5ad5a9cad... Though the horse doesn't have a spacesuit, neither does it have one in the Stable Diffusion wikipedia page.
Edit 2: finally got the kind of result I wanted https://chatgpt.com/share/6aa06813-4818-83e9-9c19-8c8ac9a348... https://chatgpt.com/s/p_359c0c38939c81918addaaa9de196f4e
That’s surprising even gpt-image-1 was able to handle the equestrian astronaut prompt pretty trivially, albeit yellow-tinged as all hell.
For the heck of it, I ratcheted up the difficulty of the prompt (two-headed horse, astronaut with the visor up, etc.), and gpt-image-2 got it right every time.
https://imgpb.com/HwDVcJo
Nothing will ever beat this one https://cf.preview.redd.it/a-photo-of-an-astronaut-riding-a-...
Wow, that's amazing. I was about to tell it to make the helmet similar to the human one, with the gold visor, but thought it wouldn't convey it as well, but not only is it obvious, it still has the feeling of being a horse.
Does multi turn image consistency mean they have solved the yellow tint issue?
Fine tuned slop. It still isn’t very visually appealing, what’s the end game here?
Good, keep going.
text-to-text is a solved problem.
Trump is going to have so much fun with this!
Is this even public?
Coming to HN to read comments on releases like this reminds me that there is a small, vocal minority that is pushing so hard for more and more features like this.
The majority of the world does not want or need any of this, yet the nerds in SF that can hardly hold a conversation with another human being are pumping it out as quickly as they can.
I truly think we're fucked.
Do you travel? I've visited several countries in the past couple of years, and everywhere I go I see AI-generated images used widely by businesses.
People love being able to generate flyers and stuff with AI. This isn't an SF only phenomenon.
[flagged]
[flagged]
[flagged]
Please don't post unsubstantive comments to HN, and particularly not unsubstantive + negative ones, which kill the kind of discussion we're hoping to have here (i.e. curious and thoughtful).
If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful. Note this one:
"Don't be curmudgeonly. Thoughtful criticism is fine, but please don't be rigidly or generically negative."
What? He quotes the opening line of the article, and notes that he reacts negatively to it. Interpreting his sentiment as "curmudgeonly" is a clear exaggeration.
And sure--it isn't a high-substance comment, but it isn't irrelevant either. It was, for a while, the top-voted comment on this whole thread. That shows people resonate with the thought, and it obviously sparked quite a lot of discussion.
This seriously makes me question your, and our other moderators, motivations. Sad times.
This was not a borderline call! The GP comment is not in keeping with the kind of discussion we're hoping for, especially not at the top of a thread—which unfortunately is where cheap negative comments often end up.
The problem isn't primarily with the comment, and in that sense conradludgate did nothing wrong. The problem is the upvotes. If it weren't for those, such comments wouldn't rise to the top and choke out more interesting conversation. But they routinely do, and that's one of the biggest problems we have here.
This an oddly cynical comment, that I think says more about the observer than the observed.
People are having fun making images.
Like they do all sorts of things on the web.
Making an image is as simple as visiting a web-page.
That's it.
This i did like 15 images just today, mostly just editing and fixing one of my wifes photos that she wanted the lighting and stuff leaned up with so i decided to test out 2.5... it did a great job, but that was easily 15 images 1 of which was the end result but all 15 count toward that figure, i'd honestly guess its more 1:100-1:500 that actually makes it outside the chatgpt servers
This is the most normal thing, and this is the best example
AI is: 'my wife wanted better lighting in our reno pic thingy'
This is healthy and way in the range of normal, not dystopian.
I assume the comment I am responding to was made in sarcasm, so I'm basing my response on that register, so apologies in advance if I have got this wrong.
I don't understand what is dystopian about this: With the exception of mildly elevated resource consumption to achieve the end goal (a nicer photo), is this not the sort of things that humans have been wanting and doing for as long as photographs have been a thing?
It was previously mostly unachievable for regular people to have their snapshots processed in such a way, but it feels like it's just making what used to be special mildly trivial. Anyone who has had their wedding photographed for the past 20 years has had the photos touched up on the way through. Any school portrait for just as long.
100 years ago people would dodge and burn photographic prints to improve them as much as the technology allowed. This is just the forward evolution of that old tale.
No sarcasm.
Maybe slightly odd is how anyone could possibly think I was being sarcastic. :)
I'm not attacking you here or anything :)
I'm just a bit shocked at how cynical people can be. The thought would not cross my mind ... and I'm a skeptic of many things!
There's a video somewhere of some people in Wyoming, riding horses.
All of the comments are young people shock of these 'boomers' committing 'animal abuse'.
I mean ... just ranchers riding horses.
People think that 'Ward Cleaver' form 'Leave It To Beaver' is some 'creepy twilight zone figure' ... when he's just literally a nice guy. Or rather the character is.
We're getting way to jaded!
So - I really mean it: AI is going to be mostly very pedantic stuff.
I witnessed the birth of computers and the internet - and while AI is definitely a big deal ... I think it's just going to feel like 'more computing stuff' in the end.
Computers started out for Science - then it became about video games, socialization, design, commerce, and of course 'gambling and porn' and all the 'same vices'.
Every step of the way there were people thinking it was being 'debauched' etc..
AI is going to be about the most mundane things, just as most technology is today.
It will be boring and in the background, and probably weird like a poorly designed, slow web-page full of cruft, but it will 'do the thing'.
[dead]
If twitter and reddit have demonstrated anything is that the fun people are having with image generation is not healthy.
How so? I don't have twitter or reddit, what are people doing?
If my social circle (in real life) is anything to go by, mostly testing out home renovation/decoration ideas.
That's also what my real life social circle does. My friends and family test out home reno/decoration ideas, and as I posted in another comment, my wife likes to visualize different layouts and changes to her flower gardens.
I make my brother look older and plumper in most pics he sends me. It makes me laugh.
Yep, renovation and landscaping idea mostly.
Well, that and infographics on what to do when the AI's start to take over.
Those monsters
I don't have twitter or reddit either, I've just seen what people post online
Well what are they posting that concerns you? That's what I'm wondering.
not that i have a problem with any of these, but i believe he’s talking about anime titties, futas, lolis, etc. and their more realistic variants
Ah okay, I was thinking stuff like political propaganda or something, but I think yours makes more sense.
I know, right? And the worst part is, they didn't check with you first.
> That's it.
Yeah ignoring the massive environmental, social, and moral implications I suppose it’s not a big deal at all.
The 'environmental' consideration is much less than the argument over 'paper or plastic'.
That's going to be an issue of regulation and cost - not some arbitrary decision over 'whether I should make an image or just use text'.
Second - there are really not much in the way 'social or moral implications' for the most part.
This is a dystopian delusion that I'm hinting at here.
Folks are blowing this way out of proportion, and cynically so.
People now have 'automated Photoshop' and a slightly more power to express themselves. It's not going to change that much, and it's not a bad thing, it's a good thing. Mildly.
[dead]
[flagged]
It has been possible to create convincingly cheap image misinformation for over a year (Nano Banana last August) but civil society has not collapsed. There are also open-weight models of similar quality and functionality without any safeguards.
Checks and balances in trust still exist, believe it or not.
The "society has not collapsed" goal post seems odd to me.
More and more people are getting caught up in weird echo chambers, extremist groups, and spewing misinformation. The amount of pure slop on Facebook, either for engagement or manipulation is so high, there is more fake content on FB than real content. The conspiracy side of me believes the CIA is deliberately forcing us against each other, but even if no underlying entity is behind this, the same outcome is happening. This is a byproduct of both Social Media & Generative AI.
In my country, Canada, the number one supplier of fake news and misinformation is the government news company, the CBC. It is an ocean compared to the fake news I've ever seen generated by chatgpt.
Big claim, care to back it up?
The 'environmental' consideration is much less than the argument over 'paper or plastic'.
Citation needed.
More to the point: both can be bad. We don't let people run stop signs because running red lights is worse.
oh no... it's fun... you got us...
It is 100% not. You can create a whole lot of misinformation with an image. You cannot do that by visiting a website. To dismiss the concern with “people are having fun” is, at best, a naive idea and, at worst, a deliberately deceitful and malicious comment.
Here is a guy tweaking photos of 'girls on a beach' in 1990 with Photoshop, which anyone can use to do, relatively easily ...
... for almost the last 40 years
"a deliberately deceitful and malicious comment."
This is a 'lost contact with reality' type of thing to say.
This thread is not about AI, it's entirely about people's ability to ground reality.
[1] https://petapixel.com/2015/02/20/a-blast-from-the-past-demos...
"which anyone can use to do, relatively easily" this is not true at all. Not anyone can use Photoshop, you need to first buy or get the software, then you need to learn how to use it, then work on the images. Whereas now you can just use Grok and immediately create pedophilia. Another deliberately deceitful and malicious comment
I hate it because I hate having to look at these AI generated images. In fact I want an internet where I have to explicitly opt-in to see AI generated content in general.
I was watching a friend stream a dead game lately.
That game was developed before generating images became widespread, and sunset as that "art form" started really taking off (game-as-a-service... probably partly reverse engineered and retooled with AI too). So during its lifetime, it was pretty much AI free.
I remember having been seduced by that game's art style at the time. It wasn't "the usual guys with futuristic guns and armor", it had some stylization as well, nice color...
Watching the stream, it occurred to me that it was in fact artwork from "the days before", in a way I hadn't been hit with until that point, even as a returning artist looking at pre-AI works.
"People came up with the concepts and developed and rendered them themselves", I thought to myself. Not in a "oh look, floppy disks. How quaint!" way, either.
Weird feeling.
Today I got stuck with a hairy frontend problem, which even AI couldn't help me with. Luckily, I managed to find a Medium article from like 2019. It was joyfully absent of 'comparisons that actually hold up' and 'X that is only a part of the picture' and nothing was load-bearing in there.
I forgot that once there existed a magical time when human beings used to write Medium articles by hand.
I would imagine a lot of it is mundane shit like getting ChatGPT to render the same room with different coloured walls or provide inspiration on plant pots. I have a decent amount of image usage and it's all things like this, not generating content I intend to distribute.
Yeah seriously, the cultural mindset is so damn kneejerk negative to everything right now. It's pathetic. This is an exciting time, and lots of good things are happening.
Maybe? Transformers and neural networks in general can do lots of interesting things for sure. But an image with the default ChatGPT aesthetic mindlessly produced by a generic prompt isn't one of them.
Seriously, it’s so fucking lame and boring to read these cynical takes. I might just make a HN comment that filters these people out, they’re ruining the site.
To be clear, I don’t mind a well-reasoned negative take on something, but much of the negatives takes I see here these days isn’t much better than Reddit, just hyper-exaggerated bitching and moaning about everything.
[dead]
[flagged]
The weirdly narrow and cynical purview of your example proves my point.
It echos sentiments of painters, intellectuals and odd religious moralists and other rabble rousers who spent decades raving against 'tHe eVIL pHoToGRaPH' in late 19th and early 20t century.
People can now make more expressive images than they could before, which is generally a good thing.
Of course, people could always use new technology for negative purposes - and - people have been using Photoshop for decades already to make 'fake nudes' - this is not a new phenomenon.
None of this is even a very big deal. We can make images now much more easily, that's it.
> imagine being someone being blackmailed by photos created by ChatGPT Images
If anything, the fact that I know for a fact that I can, in 45 seconds, generate an image of any arbitrary person holding Solo cups at a party arm in arm with Vladimir Putin (to crib an idea from one of the demos) should make everyone less and less susceptible to blackmail in general - in fact, even if the source material is real! There was a time 40 years ago when a photograph's existence was considered proof of a fact - the existence of Photoshop-like tools made that into a weak proof, and the existence of genAI has made "a photograph" into a format that is known to be as malleable as a .txt file.
Whether we like that or not, no one will stop bad actors from attempting to abuse this technology, so democratizing its use does help people to understand how they work and make potential victims much less credulous.
this is an oddly dismissive comment that says more about the commenter then what is commented on.
generating an ai image is vastly more costly than loading a webpage…but still not that costly compared to most other things.
this is depressing in the same way content slop or fyp feeds are fun.
if you want a silly pic with ur kids, just pay someone on fiver. it will be only slightly more expensive and likely far more exciting for kids as they have to wait for the dopamine hit.
> if you want a silly pic with ur kids, just pay someone on fiver.
Incredible alternative. Talk about depressing.
"if you want a silly pic with ur kids, just pay someone on fiver. it will be only slightly more expensive and likely far more exciting for kids as "
???
You mean to imply there's something 'more authentic' about paying $5 for a random Philippino guy to use Photoshop to mix-up some family photos with sticker overlays?
What are you even talking about?
On what level is that either more authentic, better for the environment, more novel ... of frankly better in any way?
That is 'not even an argument' - and so you're making my point for me:
People can now make images.
It's mostly for fun, some of it will be used professional, some Tweets will be a bit more visual - but that's it.
It's not going to change that much, and it's generally better on the whole for people to be able to create what they want.
People are putting this out of proportion.
Maybe I'm too old, but I've been around - this is not a big deal in the grand scheme.
And it's 'net slightly positive'.
>if you want a silly pic with ur kids, just pay someone on fiver. it will be only slightly more expensive and likely far more exciting for kids as they have to wait for the dopamine hit.
Or, y'know, learn how to draw, or accept that it's not exactly going to be Chuck Jones caliber stuff.
its not “oddly” and u know full well
>People are having fun making images.
3 billion images per week isn't people, it's automated farms, dogshit marketing companies and other attacks on your brain.
All of it?
What percentage is benign?
Do you have a source?
[flagged]
You made a specific claim (~100% of ChatGPT image use is non-organic), and they asked you to support your claim. That is not sealioning.
They're having fun asking a robot to make a picture when they could have learned to make it themselves.
It's incredibly depressing.
Yea, “fun”, sure. I’m sure none of those billions of images were used to lie, mislead, or cheat anyone. I’m sure none of those images were used to create fake news narratives, used in advertisements to bait people into buying fake products, or create misleading mockups of rental properties. Definitely not.
Nevermind the fact that the examples in OpenAI’s blog post are examples of creating images that have no plausible purpose other than to fake stuff that didn’t happen.
But yea sure, it’s all just “for fun”. Sure.
Yesterday my wife used ChatGPT to generate an image of our cat as a toilet, which was very funny for the both of us. She's also used it to turn him into a basketball getting dunked, a tsetse fly, a clown and a napkin. It's hard to say whether she uses it more to generate inside joke images about our cat, or to generate mockup plans and layouts for her flower gardens, but it is safe to say she does it all for fun.
This was routed internally to CatGPT which runs only Vera ribbon.
Do you purposefully not use any products that could be used to do anything harmful?
A hammer is great for pushing nails into wood. It also can crack someone’s skull with efficiency. I’m assuming you do not use hammers then, right?
Knives are freely sold at shops on every street.
...making a kitten tangled up on a networking cabinet was fun.
Some of us find opportunities for the small joys that help offset the background radiation of despair.
The sheer waste that accompanies image and video genAI is actually terrifying.
I challenge OpenAI to put a "your carbon footprint" field next to each generation. If you have nothing to hide, more information is surely better, right?
Every SDXL image uses something like 0.5 Watt-hour. That means that 2000 sdxl images uses about 30 cents of electricity. If we multiply that by 30x for video (a fair assumption for models like minimax H3, which take that much longer to make video), that means you get almost 700 5-second videos for a dollar of electricity, which is about a kilo of carbon dioxide. More or less depending on where your electricity comes from.
> If we multiply that by 30x for video (a fair assumption for models like minimax H3, which take that much longer to make video)
It takes about 7 minutes for my RTX 5090 to generate a 15 second video with MiniMax H3. At 600 watts, that's 252,000 watt-seconds, or 0.07 kilowatt-hours.
My electricity is 18 cents/kWh, so $0.0126. Just a little over a penny.
No way gpt-image-2 is as cheap as SDXL. (I dont concern about gpt-image-2's consumption either, just pointing out it's not a good comparison)
People always napkin math something low, and then every academic who formally studies it ends up with a much higher estimate.
I await OpenAI's carbon footprint field.
would you like a carbon receipt for your video gaems, netflix, and reddit sessions?
Someone needs to name this fallacy. The existence of existing harmful norms does not, ex ante, justify introducing additional harmful norms. I propose naming it the "GenAI Defender Fallacy".
Separately--yes, I would love carbon receipts for everything. Such a system would be fascinating and do real good in this world.
> The existence of existing harmful norms does not, ex ante, justify introducing additional harmful norms
The point isn't to justify anything.
It's to put things into perspective, and give people a more informed and holistic point of view, and then to ask them to reconsider their opinions in light of more information. There's nothing bad about that. In fact it's a good thing.
Especially in a world where people are so easily misled by others who cherrypick unexceptional statistics and then intentionally present them as exceptional in order to generate outrage, retweets, clicks, whatever.
It's a combination of argument from analogy and reductio ad absurdum. It's not a fallacy. You're free to dispute the validity of the analogy, or you can dispute the absurdity, which you did, by confirming that you would indeed love carbon receipts for everything.
selective outrage, on the other hand, already has a name.
What's the goal of complaining about it though? What's more likely, you persuade people to stop all of these habits that use electricity that burns carbon, or you convert the electricity supply to renewables?
But is fairness between carbon uses justified? Or we cut the line at some point.
> Someone needs to name this fallacy.
Whataboutism?
Whataboutism isn’t a fallacy and is widely misused online.
If someone says “I think x should stop because of y”. It is a valid argumentative response to say “y also occurs with z, so should z also stop?”. It is argumentatively revealing either hypocrisy, or that the original imperative relies on additional, unstated arguments/assumptions/biases.
I don’t think that’s whataboutism?
Maybe I’m confused, but I thought whataboutism was more like “you are saying I am doing X bad thing? What about Y bad thing you’re doing, let’s discuss that instead”.
It’s a deflection technique, not a check on logical consistency.
often times X is actually bad but human nature is such that no one actually cares about it while wanting to appear to when locked down on it. for this reason arguing about X directly is bad optics.
whataboustism is basically an effort to draw attention to the selective enforcement of norms, laws, morality, whatever. hypocrisy is implied.
It can also be used to derail a conversation. I.e. "I find subject X unpleasant, so I am going to imply that you don't care enough about subject Y, which has a passing similarity to it, so that the subject of the conversation turns to Y rather than X. I do not actually care about subject Y, but I won't say that out loud". See also "concern trolling".
Considering that especially leftists have spent literal decades hating on vegans for telling them about their individual footprint, there is no real fallacy here.
AI haters have somehow discovered personal responsibility. Awesome. They've never ever applied that to anything else they PERSONALLY do, though. It's only things others do. For any personal choices, suddenly it's all the governments fault or the corporations fault. No personal responsibility to be seen.
My 5-ish years of veganism have easily made up for a life time of prompting. Yet, I still would never generate video or images in part because of their impact.
How many of you non-vegan AI haters are going to be vegan now?
This is such a strange reply I don't even know what to do with it. Good on you for being vegan, and not prompting?
I have gone lacto-ovo vegetarian in the past year, so I'll take your 5 years of veganism as inspiration. Otherwise, I welcome people waking up to the harm they cause, regardless of their past hypocrisy.
I can acknowledge that I should be vegan and still be unable/unwilling to make the sacrifices to do it.
People are flawed.
Being flawed isn't the issue. Everyone is flawed. The issue is blaming people for problems like environmental damage, while at the same time needlessly causing more damage yourself. The problem is being hypocritical. Some people are certainly more hypocritical than others.
I mean, I actively advocate for better fuel efficiency standards but also haven't entirely eliminated red meat from my diet. If I'm discussing fuel efficiency standards and you call me hypocritical, even if I granted that, it doesn't seem like a useful turn of the conversation for anyone unless it's a rhetorical move to change the topic of conversation?
Advocating vs blaming. Blaming people for using inefficient vehicles when you eat red meat (assuming you eat enough red meat to have more of an impact than the fuel) is hypocritical. Advocating for better alternatives to inefficient vehicles while eating red meat is not hypocritical.
I've gone chicken only. Mostly because pork and beef are too expensive. It is very hard to get 150+ grams of protein a day on a vegan diet. And butter is in almost every delicious baked good. Avoiding it is like having the worst allergy.
Bad examples. Use meat consumption. It will blow any other kind of "non-essential" resource use out of the water, on ALL fronts.
It's really hard to believe that people care about the personal impact of their decisions, when it's only ever other peoples decisions that get talked about.
Meat is particularly ironic because it uses an obscene amount of water. Even worse, a lot of the feed for cattle comes from the imperial valley in CA, which has first dibs on the Colorado River, which is on the verge of collapsing. Obviously they aren't cutting usage.
But datacenters using water is what the public needs to fear...
Which is, BTW, a funny way of thinking of living creatures.
Jeez where did the vegan touch you?
Care to refute the actual numbers?
No one was talking about the ethical part (which you may or may not agree with). Carbon footprint is an actually measurable quantity that we can compare - what's your problem?
I personally would love them
that would be actually pretty cool and would cause more awareness for a lot of people. but something like this would go in both directions i am afraid, as with everything there surely would be people trying to do carbonmaxxing just for the sake of it, kind of like breaking a highscore. i am sure people like this already exist anyway, though.
And when my "Reddit energy consumption" rose because they migrated from old.reddit.com, to their new bloated frontend, that's _my fault_, right?
It is because you know this and you continually choose to keep using it.
It’s be like if your 2001 BMW car was super fuel efficient, and then for whatever reason the 2028 model guzzled gas and you chose to buy it anyway
Just because the thing you did once upon a time was efficient, does not mean you have a lifetime pass to engage with it no matter the updated circumstances
So your suggestion is all responsibility is shifted to the consumer? Absolutely braindead take.
reddit used to shared how much gold purchases went towards their server costs (pretty sure it wasn't 100% accurate but at least good to an order of magnitude). It was approximately $1/hour and they served ~100M monthly active users. So I don't think those metrics would demonstrate how wasteful traditional data centers are like you think it would.
I'd love that. What if anything entertainment was required by law to have some value like that, with a standardized calculation?
this would be amazing, absolutely. Especially if they were able to see specifically where my power comes from, the embedded carbon in the chipsets used to serve my requests, all the costs associated with the datacenter(s), etc. A weekly breakdown would be excellent.
Absolutely, I'd love to see how tiny they are in comparison.
If said game, video, or social media site used 1-2.5 kWh of electricity for every 10s of content then I'd say definitely!
Thankfully web browsing is no where near that.
You must have your units really messed up because there's no way that's right. A GPU drawing 700W for 10s would be 1.9444 Wh or 0.0019444 kWh and that's if the GPU was only serving your image gen for the whole 10s which is unlikely, most of the time is probably queuing since even a much lower powered GPU in a desktop PC can generate the same image in well under 10s.
Do you have a source for 1-2.5kWh for 10s of content? It takes about a minute or two to generate, so you'd be talking about a GPU consuming 30_000W-150_000W, it's just not possible. Even if it was distributed (which I don't think it is), that'd be 30-100 datacenter GPUs running at 100% to generate one clip? There's no way that would make financial sense.
Sure, why not.
_This comment increased sea levels by 0.000000000003m_
Yes.
even the whataboutism is slop
You can generate images on 5060ti, it'd take less than minute per image. Trivial environmental footprint.
I propose they use my body for compost to power their electricity turbines, in exchange for a advance in tokens while I am alive...I am overweight and that is good, more burning mass. We could settle at 2 million USD.
If they could harness the energy from all your atoms it would be worth significantly more than 2 million.
About the same amount of energy per image as a second of hot shower.
Is it fundamentally different to burning wood just to enjoy watching the flames?
A round trip cross country plane ride for one passenger is equivalent to about 500k - 1 million image generations. So all one has to do to offset their lifetime usage of image generators is forego one vacation. And this of course is only for current day carbon usage, that could go down (or up I guess but compute-wise usually things get cheaper over time).
[dead]
The waste overall for data centers is beyond terrifying. Enough water for 1.2 billion people. It’s like all the GenX and Boomers looked at the challenges younger people will face and decided make it as bad as possible start kicking them in the ribs for good measure.
Theres going to be a reckoning.
You understand that any water used is not destroyed, right?
Where does the concern for all that water go?
Most are closed systems. Water gets recycled. Water gets used like bathtubs and water bottles, right?
You cannot believe water is destroyed.
I know of entire media departments that went from - let’s hire an artist and see what they come up with to - let’s produce hundreds if not thousands of images to explore all possible ideas (and end up with something bland anyway).
Do you want to know how many identical questions ChatGPT gets asked every week, recomputing its answer each time because people stopped sharing results? I don't, but I'm sure it far exceeds 3 billion. "How to center a div", "how to install this and that package", "what is this pop culture reference", "who is this famous person",...
I ask a question to ChatGPT and i get an answer.
I ask a question to google and i spend 10-100 times longer visiting multiple websites making multiple servers generate webpages wasting electricity and time reading them finding my answer, trying different solutions
Sounds like a complex question. On average you'll visit like 1.5 sites. Those SQL queries are peanuts compared to an LLM burning through tokens. Besides, for hard questions, the LLM itself might download 30-60 websites to find the info on your behalf, e.g. checking the state of RAM prices.
OpenAI does cache the most frequent questions like this.
There is no way to cache the answer to questions that differ by even one token reliably without introducing potential problems.
there absolutely are ways to cache common-routes to fields of commonly requested paths.
(hint : it's not all regex and checksums.)
Ok I'll bite, how do you reuse the answer to "what color is the sky?" while also layering in memory, custom system instructions and custom response styles?
It's probably only memory that'd need an answer for some form of caching to be worthwhile (though I'd be curious what that answer is) since memory is on by default and rarely identical between users.
The rest are nice to haves, but not everyone customizes settings just because they are there so there will be some large pool of users with the defaults who could hit cache without them.
Doubtful. Considering features like memory, response style, custom system instructions, etc. can heavy alter a users output.
Imagine if they used the stackoverflow model of "closed as dupe"
I’m doing renovation . I slap a prompt in a tile store over my room and design several types of tiles and paint from one prompt while causally browsing in almost real time .
Isn’t that the future ? I generated 200 images alone in this use case .
What’s the alternative to that ? Send it tomorrow someone who hate their job moving tiles in the toilet , may money , wait 1 week
Reading the comments here is worse.
Why? Because you disagree with them? Would you rather it be an echo chamber?
It's like I am on Reddit or Facebook.
In my home city, I started to see banners all made with AI.
From my kids schools to banners for food places.
Groan. Your comment is the most depressing thing I've read today. What a sad way to look at the world. People are doing stuff and having fun. Meanwhile a whole sub-cult is sucking from the doom straw and mumbling carbon, water, greed, blah. The sport of Golf consumes far more resources than any AI data center.
Wow, just a quick check .... it appears you are correct.
Indirect Water footprint via Electricity Usage:I don't really get the golf course hate.
We need more green spaces, we need people to be active. Golf is a fairly quite sport, animals do foster around the greenery with golf courses. I would much rather a golf course than hot concrete everywhere.
I'm not hating on golf. My grandfather would be appalled! In crowded urban areas they don't work so well anymore. And I'm also not against using our resources for human pleasure!
its because "green spaces" and "spaces with greens" are not the same. just imagine a 150 acre generic park or forest instead of an over fertilized over watered field for a single game played by only .002 percent of people.
golf is generally a huge waste of space for the amount of value we get from it per person and should have no place anywhere near a population center where we actually need more green space.
Golf courses are not green spaces, come on now. They are most often limited in their access and devoid of greenery native to the area or that supports pollinators. Grass does not equal green space, and certainly not the flavor of green that is present in golf courses. The alternative to a golf course is not hot concrete either, that's a false dichotomy.
the golf stat is actually way worse than this... because you aren't accounting for active users. how many people actually benefit from all that water.
a golf course... at BEST has two or three foursomes on each hole on average at a time. so 36-48 players concurrently. accounting for less than perfect utilization it rounds to a rough throughput of about 1000 people per day per course. winter Exists and so you probably get like 20-50k people per year.
or maybe an easier term would be the estimated 160m people who played golf at least once last year.
which means that not only are we using that much water for so few people (.002% of the population) we are also reserving all of that space for so few people. for courses out in the middle of nowhere that doesn't matter but for example are thousands of acres of courses inside and near cities. one in LA reported 117k visits one year and that required dawn to dusk 365 play only possible in a climate like LA.
A typical course is ~150 acres so we could fit 50 soccer fields into that same space and service 3.2 million players. Ill spare the envelope math but its pretty hard to come up with any other use for that space/water that could serve less people.
ok one more example that doesn't need water either. just up the street from the high volume golf course is runyon canyon. similar size 150 acres, serves 1.5m visits a year.
Basically golf is a major waste of resources for very little benefit in basically every way possible, and should be zoned out of existence.
Yes. And right after we kill golf lets pave over Central Park. And we can tear down sports arenas and concert venues after that. It's a huge waste when you think about how few people actually go there and what benefit they actually have. Sports overall are a waste of time. Everyone should just walk in the woods instead. Except soccer. You're right that we need thousands more soccer fields. Its way more efficient than any other kind of exercise. And there is so much demand for soccer in the US, you can't even find an empty field. In contrast to golf courses which are mostly empty like you noticed.
go ahead and reductio ad absurdum.
you aren't countering my core point that golf courses is the possibly the worst use of space and basically anything else would serve the same or better utility for way more people.
heck keep it golf but switch to driving ranges and its still like 10x more capacity for users than golf courses.
counter arguments that point out other things that are still 10x better uses of space than golf are not very compelling. a sports arena that serves 2.5m people/year is a terrible argument against a golf course that maxes out at 117k. (using sofi in LA numbers as its a similiar footprint as well 150acres)
im also not saying abolish golf, just only play it in places where the space/water use dont matter. like rural scotland. don't put golf courses in/near cities where we should use our space better.
[deleted]
I’m not commenting on whether generative AI is especially resource-wasteful or not, but a lot of people online seem to have newly discovered the existence of data centers.
Remind me what golf club comes close to the energy consumption of a data center?
Or do you literally mean the entire sport of golf vs 1 data center?
That's fair and I left it as is knowing that. But it was still factually correct.
The main point is people worrying about resources and how to allocate them. We are creatures who want joy and I will not apologize for that. Certainly in face of abstract things such as "the environment". Which environment? I am a humanist. That definition varies. Golf is not the problem sorry to pick on it!
I was curious and tried to come up with a crude estimate - all of the golf courses in the US use approximately the same amount of energy (mowing, fertilizer, etc) as one 1GW DC does.
K, but they're paying for that energy though. And the land and the building costs, etc. I really don't want to live in a country where the government just arbitrarily says "Oh, sorry we have 'enough' datacenters / Subway sandwich shops / golf courses / storage facilities so you aren't allowed to build one" -- even if it becomes a populist belief that they should.
The fact is, at this point there is so much demand for DC capacity that the prices are super high, and thus it's worth building them. Many people believe that it's a bubble or whatever. If they're right, well then a lot of people building DCs right now will lose their shirts. If they're wrong and the demand is sustained, DCs are exactly what we need to be building.
>where the government just arbitrarily says "Oh, sorry we have 'enough' datacenters / Subway sandwich shops / golf courses / storage facilities so you aren't allowed to build one" -- even if it becomes a populist belief that they should.
It's not really arbitrary though, is it? It's decided by duly-elected representatives who answer to the citizens of the municipality. A government "of the people", etc.
I feel like what you're arguing against is precisely why we have local governments in the first place.
Well, you could definitely do policies like that in a plain democracy or republic without any guaranteed rights of the citizens. In our particular constitutional systems, and many others, the founding documents say that the government can't do certain things, even if they pass a law saying they can. (Of course, the Constitution can be amended to remove rights, but there is an intentionally high bar to meet in order to do that.)
Congress could pass a law tomorrow, or California could pass a ballot measure with a 100% Yes vote, saying (for instance) that some random company now belongs to the government, or to Gavin Newsom personally, or to me, but the courts would still be obligated to strike it down since that grossly violates the Seventh Amendment. Ideas like this are commonly referred to as being checks against "mob rule" - the system recognizes that 'what the people want' isn't necessarily supreme when it runs up against other people's recognized rights.
Is there such a thing as a 1GW data center? Like a beefy gaming PC easily uses 1 KW According to Google, an individual GPU tower consumes anywhere between 10-100kW. An 1GW data center is can fit in a size of a mid size apartment.
Okay, but firstly golf is probably the most useless sport ever and secondly there is more than one AI datacenter.
Yes, some people have fun, but most of the "fun" nowadays is just endless mindless consumption.
Golf consumes more water than ALL datacenters on earth.
But they both consume too much water. That's why we need less of both. Much less golf courses though.
How people derive fun is different from person to person.
The fun of some is the horror of others.
Sure. That's not a useful insight.
What is the "horror" in this case?
Lol. Let's not pick on golf? All sports are pointless depending on how you look at them. But here we are thousands and thousands of years later and humans want what humans want.
I just ignore most of these people. 99.9% of the time they're losers who peddle misinformation.
to be fair you need to create a dozen images to get a decent one. Many times i need to get to almost what i want using paint.net and hope chatgpt removes the aliasing/cleans up the pixel boundaries without messing anything up - i'm not complaining though, still faster and cleaner than photoshop -- when it works
Must be early. Try the news.
Insert "Quit Having Fun" meme here.
Why?
May I ask why this is “most depressing thing I have read today”.
Concerns about misinformation? Compute/energy use? Creative work Replaced by AI?
Having a lot of fun with my son using an image of him and dressing him in all kinds of costumes. How is that depressing to you?
Didn't you hear? you must hire a real artist to do that for you or you're exploiting the industry
This is really depressing to me. I think this is abusive.
Your son has no concept of the implications of his body, face and identifying features being used here. He has no ability to opt-out, refuse consent, and avoid his data (biometric features) being swooped up in a data-centers for further training, further data collection & further monetization.
Your son didn't agree to the privacy policy of OpenAI. This is an extremely high level of disrespect to a person you created.
You're insane.
If you're not familiar of the pitfalls of this tech, and the fact that there are more ethical tools with which to do those things, I'm not sure what to tell you.
If you want to discuss a matter, it would be nice to spell out your point of view.
How is the person supposed to know your opposing point of view if you don't tell it?
What are some of the more ethical tools? I would like to check them out
Its a simple way to feel superior - to sneer at normal people having fun and using technology.
[dead]
[flagged]
4 minute old account, wonder what you should rename yours to?
It's poor Poster's Bushido to post like that on a throwaway account.
I think @conradludgate was alluding to that we have so much of AI slop these days.
But I see your point too. I too like to generate cute and funny images and share it with my friends and family.
But this being done at scale can lead to overall degradation of the online experience.
The slop is what upsets me the most right now.
Some people feel the need to attach an image to almost everything the post - even in chat rooms (including slack at work). These images serve no purpose, but the poster feels it is useful - though it was often reaction gifs in the past I'm seeing a lot more AI content than I ever saw reaction gifs, maybe because of novelty or maybe because it's ultra personalised.
Sometimes images help to visualise something or to get a point across, but I see so many people who think it's necessary to reply to a discord message with a cat with human limbs doing a dance, or a photo of "themselves" climbing a mountain with the Rust logo to show them mastering Rust... Ok?
Image models are useful and I'm thankful for much better visual reasoning but it really really frustrates me the constant need to burn money for all of this slop.
plus the energy cost - https://www.technologyreview.com/2023/12/01/1084189/making-a...
>Generating 1,000 images with a powerful AI model, such as Stable Diffusion XL, is responsible for roughly as much carbon dioxide as driving the equivalent of 4.1 miles in an average gasoline-powered car. In contrast, the least carbon-intensive text generation model they examined was responsible for as much CO2 as driving 0.0006 miles in a similar vehicle
That's the complaint? That if I make pictures for a few years, let's say on average 1 a day.. it will equal a small errand of carbon?
If anything now I feel LESS guilty about usage.
But 3 billion pictures per week is equivalent to 12.4 million miles or 639.6 million miles extra per year.* It's not about you. There are no 3 billion people each creating one picture per week. Its probably more like 10.000 assholes creating 2.5 billion pictures per week to satisfy some stupid online feed and the other users creating a picture per month on average.
* if we trust the numbers in the other comment
I mean that in a big number..
It equals around 0.004% of actual driven miles if my estimates are right.
That is with an estimated ~270 billion miles vehicle miles every week (I put commercial in there.. about ~200 billion miles if you only include passenger vehicles.)
That study is from 2023 when the cost to produce 1000 images was 2.91 Wh per image. More recent numbers for Stable Diffusion's 2025 models that puts it at 1.3 Wh per image and that number continues to decrease.
For context that means a days worth of image generation emits about the same amount of carbon dioxide as a single transatlantic flight (New York to London). There are approximately 1500 transatlantic flights per day.
So every day the people of Vermont emit more CO2 just commuting than all image generation through OpenAI per week. Nice. That does put it in perspective. It’s really efficient.
yep, what our climate really needs is another carbon emitter similar to an entire state's worth of vehicles just so people can not pay artists or make dumb images of themselves ripping off some artistic style like Studio Ghibli. that is so much more of a social good than people being able to make it to work and earn a living
Oh I think paying people to make 3 billion images would result in far more carbon emissions. Just think about how much CO2 a human emits while painting. The paint, the food. Just for the sake of the climate I would never pay an artist for something like this.
for sure, because every one of those 3 billion images rendered by human hand and not automated with a tool would exist. we wouldn't have a magnitude smaller number of focused designs, of course, that's definitely not how art and design processes happen
1000 images is worth 4miles?
Ours cars are atrocious. We can have the most powerful artificial minds imagine 1000 images from simple prompts, and that takes as much energy as moving a human being 4 miles.
And that will get more energy efficient. Gas cars have barely budged.
Gas cars are the problem, not AI.
This is an embarrassingly idiotic article. They're comparing against the least intensive text model, ie some million parameter model nobody uses. Stable Diffusion is a 3.5 billion parameter model, while the GPT models a billion people are using for text generation are over 10 trillion parameters. To say nothing of the differences in average context size usage. The actual ratio of SDXL to text generation pollution is probably literally reversed from what the article claims by lying with statistics.
(Note, however, that OpenAI's image model is much much larger than SDXL; however, we don't have precise numbers for it. Nonetheless, misinformation is misinformation.)
You're building a weak strawman here. The author didn't call you nor your son out specifically.
I wouldn't feel comfortable uploading a picture of my child to OpenAI servers (for fun nonetheless), but you do you.
I'm not following the news. Why should we not upload pictures of family members to OpenAI? Are any other cloud providers safe?
I wouldn't/don't really feel comfortable uploading private pictures to any other cloud provider either :). I find OpenAI especially untrustworthy because of their known business practices, unestablished business model and unknown future capabilities and role in society. You cannot take that data back once they have it. The condescending tone was unasked for, apologies to OP for that.
I wouldn't be uploading pics of my child to platforms that have historically used CSAM to train their models, but hey you do you
Typical HN luddite mindset at this point. This is an exciting time!
How many memes/gifs do you think have been created by hand or with generators over the last ~20 years?
Not sure you can compare huge ML clusters of GPUs running models in order to do these images, and the typical "meme generator" which is basically a call to imagemagick passing an image and some text.
My point was with respect to brainrot being depressing but existing for a long time before AI. Not resource use.
[dead]
[dead]
[flagged]