Large language models

Future historians might point to the Fall 2022 release of OpenAI's ChatGPT large language model as the dawn of true AI.

What LLMs Are

Large language models (LLMs):

Their sole training goal: predict the next word in a sequence.

Emergent abilities

While learning to be expert text generators, LLMs also learn unexpected capabilities:

These abilities raise questions about:

AI research is moving rapidly to explore these areas.

How LLMs work

LLMs use transformer decoders (without encoders) trained in an unsupervised way using massive text datasets.

Pretraining:

Pretraining enables the model to learn language (including grammar and syntax), and, apparently, enough world knowledge to support emergent abilities.

After pretraining, the decoder generates text in response to the input prompt:

When used in chat mode:

Transformer models have a fixed input width or context window (around 32,000 tokens for GPT-4).

This large input window allows attention to go back to things that appeared far back in the input, which is something recurrent models cannot do.

Alignment

After pretraining, LLMs are ready for use, but often fine-tuned first on domain-specific data.

GPT-4's fine-tuning consisted of a step known as reinforcement learning from human feedback (RLHF):

Alignment is absolutely critical to ensure that powerful language models conform to human values and societal norms.

In-context learning

A remarkable property of LLMs is their in-context learning ability. LLMs can "learn on the fly" without changing their weights just by seeing examples in the prompt, rather than through formal retraining.

The attention mechanism built into the transformer architecture is the likely source of an LLM's in-context ability, but it isn't entirely clear.

Unintended new abilities

What in the data, training, and model architecture enables them to do what they do is still largely unknown.

From the "Sparks of Artificial General Intelligence" paper:

How does [GPT-4] reason, plan, and create? Why does it exhibit such general and flexible intelligence when it is at its core merely the combination of simple algorithmic components-gradient descent and large-scale transformers with extremely large amounts of data? These questions are part of the mystery and fascination of LLMs, which challenge our understanding of learning and cognition, fuel our curiosity, and motivate deeper research.

Simply put, researchers don't know why LLMs do what they do:

It was a happy accident that the transformer architecture evolved such abilities. This did not happen by design.

We can probably expect great things as more advanced transformer architectures come along.

Previous Generative adversarial networks All ⏎ Next Coding agents

A Kemar Joint