All posts

General/ Pinned

Image Generation for dummies

Fundamentally understand what different types of image generation is doing

Table of contents3 sections

To understand image generation, we must first understand probability.

What does probability mean for an image generator? What does p(x) mean?

Imagine we have our image generator model, named Picman. Picman is a very focused artist. He doesn't do commissions, only drawing what he feels like drawing. You go to Picman with a blank canvas (pure noise), and he draws you a picture.
You give Picman three noisy canvas for him to draw on each:

image
Random Noise of size 512x512
image
Random Noise of size 512x512
image
Random Noise of size 512x512

The first time, Picman gave you back this drawing:

image
1st picture he drew you a puppy
image
This time he drew a cat
image
"I'm gonna be the king of the pirates"

Picman drew on each canvas, he made a dog, a cat, and the man who will become king of the pirates, Monkey D. Luffy.

You walk away with three awesome pictures. As you walk away you see a group of AI researchers come up to Picman afterwards. They wanted to find out exactly the capabilities of Picman. They gave 100,000 noisy images to Picman and asked him to draw 100,000 times. (these freaks)

They researchers later organized the all the drawings Picman did into different categories. They found that Picman drew:
20,000 pictures of different dogs: golden retrievers, husky, border collie

image
20,000 different pictures of dogs

He drew 10,000 pictures of cats: Bald cat, orange cat, black cat:

image
10,000 different pictures of cats
image
5000 pepe memes
image
And he drew only 500 anime characters

Wow, you thought to yourself, given 100,000 random change, he only drew 500 anime characters, and he only drew Luffy 3 times! You sure hit the jackpot with the Luffy drawing you got, its a 0.00003% chance!

Now the team of AI researchers have figured Picman out completely. They know that, given a random noise, the chances of getting a picture of a dog is about 20%, since they got 20,000 images of dogs out of 100,000 total drawings Picman did. They called this:

$$p(dog) = 0.2$$

Similarly, they know that

$$p(cat)=0.1\newline p(anime) = 0.005$$

This is our $p(x)$: the probability density of a given image. Density here just means, how likely it is that Picman draws the picture x. If you want to see Picman draw a cute anime girl in a leather jacket holding a NVIDIA H100 GPU, your odds are very, very slim. Picman will almost never draw that.

$$p(\text{cute anime girl in a leather jacket holding a NVIDIA H100 GPU}) \approx 0$$

What is logp(x)? We want a better way to represent this number. Since $p(\text{luffy})=0.00003$, thats too many digits to write down. The value is too small. Especially if we want to compute with these numbers, like $p(\text{goku})=0.00002$, we can instead use the log function to change their representations. Remember, if we apply log to every number, then we didn't change any number, we just scaled everything up to be easier to read.

$$log(p(\text{luffy}))=log(0.00003) = -4.5 \newline log(p(\text{luffy}))=log(0.02) = -0.7$$

Now, a higher log value means more likely to generate image x.
If we take -log value, then, then a smaller log value means more likely to generate x.

$$-\log(p(\text{luffy})) = 4.5 \newline -\log(p(\text{dog}))= 0.7$$

0.7 < 4.5, which means we are more likely to get dog than luffy from Picman.

We can imagine all the 100,000 images Picman has drawn as a huge landscape.

image
Inside this landscape, each type of image Picman draws can be represented by a mountain. There is a dog mountain. This mountain is very wide and tall, with the top of the dog mountain being the MOST 'dog' image Picman draws. Here, Picman's favorite dog he likes to draw is a brown golden retriever. So the red tip would be something like this:

image
Tip of the mountain

And as you go further down the mountain, we see:

image
Golden Retriever facing left
image
Golden Retriever facing right
image
Golden Retriever facing backwards

As we go even further down the mountain, we see other kinds of dogs:

image
A Husky
image
A Border Collie
image
A German Shepherd

Similarly, the tip of Pepe looks like this, because this is Picman's favorite Pepe meme he likes to draw the most.

image
PepeCry

As we scale down the Pepe mountain, we will find images like this:

image
image
image

As we go even further down:

image
image
image

Now, suppose we want to know how Picman likes to draw. Each local maximum describes what Picman likes to do

image
Zooming in at dog distribution, wee see a local peak of Husky and global peak of Golden Retriever

If we show him a picture of a black golden retriever, $x_1$, Picman can tell us how to change the picture to make it look like the best golden retriever, aka the one he loves to draw the most. Because in Picman's mind, the best golden retriever should be Brown.

image
Picture at $x_1$, a golden retriever that doesn't quite look like Picman's favorite way of drawing a golden retriever.
image
Tip of the mountain, Picman's favorite way of drawing a golden retriever.

Picman can give us

$$\nabla_{x_1} p(x_1)$$

This equation tells us how to move towards a higher density. In this case, it tells us how to get to a higher mountaintop.

However, when we are around the dog distribution mountain, there will be multiple bumps. Each bump could represent how Picman likes to draw a certain breed.

Here, there is a bump for Husky. At the tip, it is the Husky that Picman loves to draw the most.

We present $x_2$, an image of a red Husky to Picman. Picman will tell us that this isn't what he thinks a Husky should look like. He can then tell you the exact way to change our drawing to make it look like the Husky.

$$\nabla_{x_2}p(x_2)$$
image
Picture at $x_2$, it is still a dog, but it looks more like a husky.
image
Top of the Husky mountain. This is the ideal Hukey in Picman's mind.

So what exactly does $\nabla_x p(x)$ look like? How do we get it?
Well, its quite simple. Picman will tell you, at every pixel in your drawing, how to make the current picture look more like his favorite picture.

For example, at the black golden retriever, he will tell you that at the black pixels, we want to move from

rgb(0,0,0) # black → rgb(204, 157, 89) # brown
In order for it to move to this, it will tell us the direction to update x:
$$ x_{\text{new}}=x+\eta\nabla_xp(x) $$
Here, we know we should add positive values to (r, g, b). something like (+5, +4, +3)

In general, it tells us how to move from our current location to a place with a higher probability density. How to make the image look more the closest thing that Picman likes to draw.

Similarly, we can also take the log of this value:

$$\nabla_x \log p(x)$$

This again is Picman telling us the direction to change each pixel, such that it looks like a picture he likes more that is closest to what you are giving him.

Now, we are ready to see how we can create our own Picman that can draw these images from scratch.

Energy based Models:

This is EnergyMan. The first way we can teach a noob artist how to draw images.
It is learning one function:

$$E_\theta(x)$$

This function takes an input of an image x.
It outputs a single single number y

lower energy = lower y value = 5: "this looks like a good picture"
high energy = higher y value = 999: "this looks wrong and bad"

Picman is very private, he can't tell us exactly what goes on inside his mind. We have to use an actual, fixed size dataset. Here, lets assume we have 100,000 different images we collected as our dataset. This includes pictures of fish, cars, cats, glasses, etc. It is a very diverse set of images.

However, we see that we don't have this continuous mountain map that we ideally would want. If we did, we can just say: make our mountain map the exact same as the dataset's mountain map. Instead, we have 100,000 different snapshots of different locations of this mountain. We can see that there is a bump of dog, where there are 8000 different pictures of dogs huddled around a hill.

The goal of energy-based models is to find a number that tells us how close we are to the tip of the mountain.

The set up for E:
We can for example use a ResNet backbone structure.
Given every image we can generate is a 32x32 image: $x\in\mathbb{R}^{3\times32\times32}$

Image unavailable: image

Set up for $E_\theta$

As input, we take every number in the image, (r,b,g) value at each pixel. Then, we make it smaller and smaller pixel wise, while adding more channels to the (r,g,b, c_1, c_2, c_2...c_509). Lastly, we make append a linear head to make it output just 1 energy score scalar.

Thats the set up of our $E_\theta$. How exactly do we train it?

When we first start:
Give it a picture of a cat: $E_\theta(x_{\text{cat}})=57$
and random noise: $E_\theta(x_{\text{garbage}})=22.$
That's not what we want, since we want the energy value of random noise to be higher, and evaluating a picture of real cat to be lower.
We want to teach it:$E_\theta(x_{\text{real}})\downarrow$
while $E_\theta(x_{\text{fake}})\uparrow.$
This gives the basic EBM training idea:

$$\boxed{L=E_\theta(x_{\text{real}})-E_\theta(x_{\text{fake}})}$$

We can minimize the loss. We need to correctly shape our function such that we will move the values in the direction that decreases.
For example, initially:
$$E(x_{\rm real})=5,\qquad E(x_{\rm fake})=2.$$Then: $$L=5-2=3.$$
How can we make L smaller? By decreasing the first term and increasing the second term.
$$E(x_{\rm real})=3, \qquad E(x_{\rm fake})=4.$$Then:
$$L=3-4=-1.$$

This way, we get lower loss. Conveniently, we lowered the first term, increased the second term.

As a contradicting example, since we are minimizing the loss, if we do:
$$E(x_{\rm real})=8,\qquad E(x_{\rm fake})=0.$$Then: $$L=8-0=8.$$ This won't happen as we are always moving our loss to be smaller.

1. We

2. Generate the image using $E_\theta(x_t)$

Start by passing in pure noise, we evaluate the image with our network, it outputs a scalar of how good it thinks it is. Then, we edit each pixel to lower the energy scalar until n times until we have our final image.
Image unavailable: image

This could be used for a later section

image
Picture we drew

We ask Picman, we want to find

$$\nabla_x p(x)$$

In this case, if we ask Picman how to make this drawing more like a dog, Picman will tell us $\nabla_{dog}p(dog)$
We can also take the same picture and ask Picman how to make this drawing more like Pepe, and he will instead tell us $\nabla_{pepe} p(pepe)$

image
Picman telling you how to make it look more like the dog distribution
image
Picman telling you how to make it look more like the pepe distribution