Large language models
Future historians might point to the Fall 2022 release of OpenAI's ChatGPT large language model as the dawn of true AI.
What LLMs Are
Large language models (LLMs):
- take a text prompt as input
- generate text token by token, using the prompt and all prior tokens as a guide
Their sole training goal: predict the next word in a sequence.
Emergent abilities
While learning to be expert text generators, LLMs also learn unexpected capabilities:
- answering questions
- mathematical reasoning
- high-quality computer programming
- logical reasoning
These abilities raise questions about:
- the nature of thought
- the meaning of consciousness
- the uniqueness of the human mind
AI research is moving rapidly to explore these areas.
How LLMs work
LLMs use transformer decoders (without encoders) trained in an unsupervised way using massive text datasets.
Pretraining:
- input text is split into tokens (words, subwords, or characters)
- tokens are mapped to a multidimensional embedding space, where similar tokens cluster together, capturing meaning and relationships
Pretraining enables the model to learn language (including grammar and syntax), and, apparently, enough world knowledge to support emergent abilities.
After pretraining, the decoder generates text in response to the input prompt:
- token by token until a
stoptoken appears - uses attention to weigh the importance of all tokens, capturing relationships even across distant positions in the text
When used in chat mode:
- LLMs appear conversational
- when in reality, each new prompt is passed to the model along with all the previous text (the user's prompts and the model's replies)
Transformer models have a fixed input width or context window (around 32,000 tokens for GPT-4).
This large input window allows attention to go back to things that appeared far back in the input, which is something recurrent models cannot do.
Alignment
After pretraining, LLMs are ready for use, but often fine-tuned first on domain-specific data.
GPT-4's fine-tuning consisted of a step known as reinforcement learning from human feedback (RLHF):
- the model is trained further using feedback from real human beings to align responses to human values and societal expectations
- this is necessary because LLMs are not conscious entities
- e.g., unaligned LLMs will respond with step-by-step instructions on how to make drugs or bombs
Alignment is absolutely critical to ensure that powerful language models conform to human values and societal norms.
In-context learning
A remarkable property of LLMs is their in-context learning ability. LLMs can "learn on the fly" without changing their weights just by seeing examples in the prompt, rather than through formal retraining.
The attention mechanism built into the transformer architecture is the likely source of an LLM's in-context ability, but it isn't entirely clear.
Unintended new abilities
What in the data, training, and model architecture enables them to do what they do is still largely unknown.
From the "Sparks of Artificial General Intelligence" paper:
How does [GPT-4] reason, plan, and create? Why does it exhibit such general and flexible intelligence when it is at its core merely the combination of simple algorithmic components-gradient descent and large-scale transformers with extremely large amounts of data? These questions are part of the mystery and fascination of LLMs, which challenge our understanding of learning and cognition, fuel our curiosity, and motivate deeper research.
Simply put, researchers don't know why LLMs do what they do:
- LLMs are trained on vast human-generated text, capturing language, grammar, and style
- increasing model size improves text prediction quality
- beyond a certain scale, emergent abilities appear and strengthen:
- this scale likely helped LLMs to learn a high-dimensional probabilistic representation of not only grammar and style but of the world in general
- including contextual relationships and simulations
It was a happy accident that the transformer architecture evolved such abilities. This did not happen by design.
We can probably expect great things as more advanced transformer architectures come along.