Khabar 24h SIMPLE EXPLAINERS ON WORLD AFFAIRS, SCIENCE, HEALTH AND MORE.

KHABAR 24H

Simple explainers on world affairs, science, health and more.

All news under one minute

Technology Read in one minute

How AI Image Generators Work: Diffusion Models Explained in Plain English

Type a sentence like a watercolour painting of a lighthouse at dawn and, seconds later, an AI image generator hands you exactly that, a picture that never existed before. Tools such as Midjourney, DALL-E and Stable Diffusion have put this power in everyone’s browser. The technology behind most of them is called a diffusion model, and while the mathematics is formidable, the core idea is surprisingly intuitive: teach a computer to remove noise from an image, and it learns to create images from noise. Here is how it works, in plain English.

The core idea: learning to undo noise

Imagine taking a photograph and gradually adding static until it becomes pure television snow. That is the forward process, and it is easy. The clever part is training a neural network to run it in reverse: start from pure noise and remove a little static at a time until a coherent image emerges. During training, the model is shown millions of images at every stage of noisiness and learns to predict what the slightly-less-noisy version should look like. After enough practice, you can hand it pure random noise and it will denoise its way, step by step, into a finished picture. Nothing is retrieved or copied from a database; the image is constructed from scratch by the denoising process.

How your words steer the picture

A diffusion model on its own would generate random images. The text prompt steers it. Alongside the image training, these systems learn the relationship between words and visuals, typically using a component that maps text and images into a shared mathematical space, so that the phrase golden retriever puppy sits near pictures of golden retriever puppies. When you type a prompt, that text encoding guides each denoising step, nudging the emerging image toward the described content, style and composition. Extra words like cinematic lighting or ukiyo-e style shift the guidance toward those aesthetics. This is why prompt craft matters: the model can only aim at what your words describe, and vague prompts produce vague, generic results.

From training data to finished image

The ability of these models rests on staggering training sets. Leading image models have been trained on billions of image-text pairs scraped from the public web, each teaching the model another tiny lesson in how the visual world looks: how light falls on water, how hands are shaped, what art deco means. Training such a model takes weeks on thousands of specialised chips and costs millions of dollars. Once trained, though, generating an image is cheap, a few seconds of computation on a single graphics card or even a phone. That asymmetry, enormously expensive to build and nearly free to use, is why the technology spread so fast once the first good models appeared.

What diffusion models still get wrong

Anyone who has used these tools knows their quirks, and the quirks reveal how the technology thinks.

  • Hands and fingers: the models famously mangle hands, because hands appear in training data in countless poses and the model learns average shapes rather than anatomy.
  • Text in images: signs and labels often come out as plausible-looking gibberish, since the model renders letter-like shapes without understanding language.
  • Counting and logic: ask for exactly seven apples and you may get six or eight; the model has a weak grasp of precise quantities.
  • Consistency: generating the same character twice in different poses remains hard, because each image starts from fresh noise.
  • Copyrighted styles: models can imitate living artists’ styles, raising unresolved questions about consent and compensation.

Each new generation of models shrinks these flaws, but none has eliminated them entirely.

The bigger questions around the technology

Beyond the clever engineering sit genuine debates. Artists argue, with reason, that models trained on their work without permission are industrial-scale appropriation; several lawsuits are working through courts in the US and Europe. The same tools that illustrate children’s books can also fabricate convincing fake photographs, and detection is an arms race the fakes are currently winning. On the other side, illustrators, designers, architects and filmmakers use these tools as sketchpads that compress days of concept work into an afternoon. Like every powerful creative technology before it, from photography to Photoshop, diffusion models are simultaneously a tool, a threat and a new medium, and society is still negotiating which it will be.

FAQs

Do AI image generators copy existing pictures? They do not retrieve stored images. They generate new pixels guided by patterns learned from training data, though the results can closely resemble works the model saw during training.

Why do I need to write detailed prompts? Because the text encoding is the only steering the model gets. Specific details about subject, style, lighting and composition give the denoising process a clearer target.

Can I use AI images commercially? It depends on the tool’s terms of service and your country’s evolving copyright law. Check the licence of the specific generator you use before selling the output.

Diffusion models turned image-making into a conversation: describe what you imagine, and mathematics sculpts it out of noise. The results can be breathtaking, and the flaws are instructive, each mangled hand a reminder that the machine understands patterns of pixels, not the world those pixels depict. Used with that understanding, these tools are less a replacement for human creativity than a strange and powerful new instrument for it.

Source: The Verge

Avatar photo
Written by
Khabar 24h Editorial Desk

Khabar 24h Editorial Desk — our explainers are prepared by the Khabar 24h editorial team using AI-assisted research tools, and every piece is reviewed by a human editor before publishing. We do not claim original reporting: our work is turning complex topics into simple, accurate summaries. Spotted an error? Write to contact@khabar24h.com — our corrections policy aims for same-day review.

More from this author →