General/ Pinned
Image Generation for dummies
Fundamentally understand what different types of image generation is doing
Table of contents3 sections
To understand image generation, we must first understand probability.
What does probability mean for an image generator? What does p(x) mean?
Imagine we have our image generator model, named Picman. Picman is a very focused artist. He doesn't do commissions, only drawing what he feels like drawing. You go to Picman with a blank canvas (pure noise), and he draws you a picture.
You give Picman three noisy canvas for him to draw on each:
The first time, Picman gave you back this drawing:
Picman drew on each canvas, he made a dog, a cat, and the man who will become king of the pirates, Monkey D. Luffy.
You walk away with three awesome pictures. As you walk away you see a group of AI researchers come up to Picman afterwards. They wanted to find out exactly the capabilities of Picman. They gave 100,000 noisy images to Picman and asked him to draw 100,000 times. (these freaks)
They researchers later organized the all the drawings Picman did into different categories. They found that Picman drew:
20,000 pictures of different dogs: golden retrievers, husky, border collie
He drew 10,000 pictures of cats: Bald cat, orange cat, black cat:
Wow, you thought to yourself, given 100,000 random change, he only drew 500 anime characters, and he only drew Luffy 3 times! You sure hit the jackpot with the Luffy drawing you got, its a 0.00003% chance!
Now the team of AI researchers have figured Picman out completely. They know that, given a random noise, the chances of getting a picture of a dog is about 20%, since they got 20,000 images of dogs out of 100,000 total drawings Picman did. They called this:
$$p(dog) = 0.2$$Similarly, they know that
$$p(cat)=0.1\newline
p(anime) = 0.005$$This is our $p(x)$: the probability density of a given image. Density here just means, how likely it is that Picman draws the picture x. If you want to see Picman draw a cute anime girl in a leather jacket holding a NVIDIA H100 GPU, your odds are very, very slim. Picman will almost never draw that.
$$p(\text{cute anime girl in a leather jacket holding a NVIDIA H100 GPU}) \approx 0$$What is logp(x)? We want a better way to represent this number. Since $p(\text{luffy})=0.00003$, thats too many digits to write down. The value is too small. Especially if we want to compute with these numbers, like $p(\text{goku})=0.00002$, we can instead use the log function to change their representations. Remember, if we apply log to every number, then we didn't change any number, we just scaled everything up to be easier to read.
$$log(p(\text{luffy}))=log(0.00003) = -4.5
\newline
log(p(\text{luffy}))=log(0.02) = -0.7$$Now, a higher log value means more likely to generate image x.
If we take -log value, then, then a smaller log value means more likely to generate x.
$$-\log(p(\text{luffy})) = 4.5 \newline -\log(p(\text{dog}))= 0.7$$0.7 < 4.5, which means we are more likely to get dog than luffy from Picman.
We can imagine all the 100,000 images Picman has drawn as a huge landscape.
Inside this landscape, each type of image Picman draws can be represented by a mountain. There is a dog mountain. This mountain is very wide and tall, with the top of the dog mountain being the MOST 'dog' image Picman draws. Here, Picman's favorite dog he likes to draw is a brown golden retriever. So the red tip would be something like this:
And as you go further down the mountain, we see:
As we go even further down the mountain, we see other kinds of dogs:
Similarly, the tip of Pepe looks like this, because this is Picman's favorite Pepe meme he likes to draw the most.
As we scale down the Pepe mountain, we will find images like this:
As we go even further down:
Now, suppose we want to know how Picman likes to draw. Each local maximum describes what Picman likes to do
If we show him a picture of a black golden retriever, $x_1$, Picman can tell us how to change the picture to make it look like the best golden retriever, aka the one he loves to draw the most. Because in Picman's mind, the best golden retriever should be Brown.
$x_1$, a golden retriever that doesn't quite look like Picman's favorite way of drawing a golden retriever.Picman can give us
$$\nabla_{x_1} p(x_1)$$This equation tells us how to move towards a higher density. In this case, it tells us how to get to a higher mountaintop.
However, when we are around the dog distribution mountain, there will be multiple bumps. Each bump could represent how Picman likes to draw a certain breed.
Here, there is a bump for Husky. At the tip, it is the Husky that Picman loves to draw the most.
We present $x_2$, an image of a red Husky to Picman. Picman will tell us that this isn't what he thinks a Husky should look like. He can then tell you the exact way to change our drawing to make it look like the Husky.
$$\nabla_{x_2}p(x_2)$$$x_2$, it is still a dog, but it looks more like a husky.So what exactly does $\nabla_x p(x)$ look like? How do we get it?
Well, its quite simple. Picman will tell you, at every pixel in your drawing, how to make the current picture look more like his favorite picture.
For example, at the black golden retriever, he will tell you that at the black pixels, we want to move from
rgb(0,0,0) # black → rgb(204, 157, 89) # brown
In order for it to move to this, it will tell us the direction to update x:
$$
x_{\text{new}}=x+\eta\nabla_xp(x)
$$
Here, we know we should add positive values to (r, g, b). something like (+5, +4, +3)
In general, it tells us how to move from our current location to a place with a higher probability density. How to make the image look more the closest thing that Picman likes to draw.
Similarly, we can also take the log of this value:
$$\nabla_x \log p(x)$$This again is Picman telling us the direction to change each pixel, such that it looks like a picture he likes more that is closest to what you are giving him.
Now, we are ready to see how we can create our own Picman that can draw these images from scratch.
Energy based Models:
This is EnergyMan. The first way we can teach a noob artist how to draw images.
It is learning one function:
$$E_\theta(x)$$This function takes an input of an image x.
It outputs a single single number y
lower energy = lower y value = 5: "this looks like a good picture"
high energy = higher y value = 999: "this looks wrong and bad"
Picman is very private, he can't tell us exactly what goes on inside his mind. We have to use an actual, fixed size dataset. Here, lets assume we have 100,000 different images we collected as our dataset. This includes pictures of fish, cars, cats, glasses, etc. It is a very diverse set of images.
However, we see that we don't have this continuous mountain map that we ideally would want. If we did, we can just say: make our mountain map the exact same as the dataset's mountain map. Instead, we have 100,000 different snapshots of different locations of this mountain. We can see that there is a bump of dog, where there are 8000 different pictures of dogs huddled around a hill.
The goal of energy-based models is to find a number that tells us how close we are to the tip of the mountain.
The set up for E:
We can for example use a ResNet backbone structure.
Given every image we can generate is a 32x32 image: $x\in\mathbb{R}^{3\times32\times32}$
Image unavailable: image
Set up for $E_\theta$
As input, we take every number in the image, (r,b,g) value at each pixel. Then, we make it smaller and smaller pixel wise, while adding more channels to the (r,g,b, c_1, c_2, c_2...c_509). Lastly, we make append a linear head to make it output just 1 energy score scalar.
Thats the set up of our $E_\theta$. How exactly do we train it?
When we first start:
Give it a picture of a cat: $E_\theta(x_{\text{cat}})=57$
and random noise: $E_\theta(x_{\text{garbage}})=22.$
That's not what we want, since we want the energy value of random noise to be higher, and evaluating a picture of real cat to be lower.
We want to teach it:$E_\theta(x_{\text{real}})\downarrow$
while $E_\theta(x_{\text{fake}})\uparrow.$
This gives the basic EBM training idea:
$$\boxed{L=E_\theta(x_{\text{real}})-E_\theta(x_{\text{fake}})}$$We can minimize the loss. We need to correctly shape our function such that we will move the values in the direction that decreases.
For example, initially:
$$E(x_{\rm real})=5,\qquad E(x_{\rm fake})=2.$$Then: $$L=5-2=3.$$
How can we make L smaller? By decreasing the first term and increasing the second term.
$$E(x_{\rm real})=3, \qquad E(x_{\rm fake})=4.$$Then:
$$L=3-4=-1.$$
This way, we get lower loss. Conveniently, we lowered the first term, increased the second term.
As a contradicting example, since we are minimizing the loss, if we do:
$$E(x_{\rm real})=8,\qquad E(x_{\rm fake})=0.$$Then: $$L=8-0=8.$$ This won't happen as we are always moving our loss to be smaller.
1. We
2. Generate the image using $E_\theta(x_t)$
Start by passing in pure noise, we evaluate the image with our network, it outputs a scalar of how good it thinks it is. Then, we edit each pixel to lower the energy scalar until n times until we have our final image.
Image unavailable: image
This could be used for a later section
We ask Picman, we want to find
$$\nabla_x p(x)$$In this case, if we ask Picman how to make this drawing more like a dog, Picman will tell us $\nabla_{dog}p(dog)$
We can also take the same picture and ask Picman how to make this drawing more like Pepe, and he will instead tell us $\nabla_{pepe} p(pepe)$