Making something new: generative AI
Earlier we learned that a writing AI picks each next word by probability and strings them together. But is text all an AI can make? It draws a picture no one has ever seen, and composes a new melody too. It does not pull out something memorized; it makes something new from the patterns it learned. An AI that creates like this is called generative AI.
Not memorizing, but making new
Earlier we learned a writing AI.
It picked each next word by probability
and strung them one word at a time.
But here is one thing to note.
The AI does not store the sentences it saw
and then pull them out unchanged.
It learns patterns from countless examples,
and with those patterns it makes a brand-new
combination, one that never existed, on the spot.
Tap to recombine the learned pieces into something new. With the same pieces, each tap builds a different new combination. It is not pulling out a memory; it is making something new.
Same pieces, yet each tap
gave a different result.
If it were pulling out one memorized answer,
it would be the same every time.
By recombining pieces with learned patterns,
something new comes out each time.
Making something new like this
is called generation.
But is text all it can make?
Not just text: images and music too
We called a text-making AI an LLM.
But generative AI is broader than that.
An AI that learned many picture examples
makes new pictures.
One that learned songs writes new melodies,
and one that learned video makes short clips too.
What it can make
depends on what it learned.
An LLM is the one branch in charge of text.
Tap the text, image, music, and video buttons to switch. Generative AI can make all of these depending on what it learned, and text (an LLM) is just one of them.
Whichever you chose,
the idea underneath was the same.
Making new things from learned patterns.
Text, picture, or music,
it learns patterns from countless examples
and builds new combinations with them.
So how on earth does it make a picture?
It uses a method a little different
from stringing one word at a time.
From grainy noise to a picture
An AI that makes pictures, oddly enough,
starts not from a clean white page
but from grainy noise.
Like the dots that fill the screen
when a TV signal cuts out.
From there it cleans up the dots a little at a time.
As the fuzzy dots get tidied step by step,
some shape begins to emerge.
Repeat the cleanup and the picture grows clear.
Each tap of the clean-up button tidies the grainy noise grid one step. The fuzzy dots gradually emerge into the shape of a picture. (Fixed steps, so everyone gets the same result.)
The noise turned into a picture, right?
It did not pop out all at once;
it grew clear by cleaning up little by little.
The AI has learned, from countless examples,
which dots to tidy and how
to get closer to a real picture.
With that learned sense it cleans the noise
and makes a new picture.
But who decides what to draw?
A person gives directions: the prompt
The AI does not just make anything on its own.
A person tells it in words what to make.
This instruction written in words
is called a prompt.
Write cat, and the cleanup steers toward a cat;
write space, and it steers toward space.
The direction of the cleanup changes.
With the same AI, change the prompt
and the result that comes out changes.
Tap to change the prompt. Even starting from the same noise, the picture the cleanup aims for changes with the instruction you wrote. The person sets the direction.
Change the prompt and the result's
direction shifted sharply, right?
The AI has the skill to create,
and the person decides what to create.
The two work hand in hand.
The clearer the good instruction you write,
the closer you get to the result you want.
So what you write and how
becomes a skill in using generative AI.
Now let's wrap the whole thing up.
Let's wrap up
Gathered on one line, it is this.
Generative AI makes new things from learned patterns.
It does not copy something memorized.
What it makes is not only text.
It can make images, music, and video too.
So it is a broader word than an LLM.
An image starts from noise
and is cleaned up little by little into a picture.
And a person gives direction with a prompt.
Tap the key points in order to review. (make new from learned patterns -> not just text but images and music too -> clean up from noise into a picture -> direction by prompt)
Now you know what generative AI is.
With learned patterns it makes text, pictures,
and music anew, all as a tool.
It builds a picture by cleaning up noise
and takes direction from a person's prompt.
But when it makes things this well,
it can also make fakes
hard to tell from the real thing.
Next, let's look at how to use it responsibly.