Two families of model make nearly every generated image. Diffusion models start from noise and remove it a little at a time; autoregressive models write an image token by token, as a language model writes a sentence. Both are steered by an encoding of your prompt.