A large language model (LLM) is a neural network — usually a transformer — trained on vast amounts of text to predict what comes next.
Next-Token Prediction
Text is split into tokens (words or parts of words). During pretraining, the model repeatedly predicts the next token in real text and adjusts its billions of parameters to predict better. Doing this at huge scale teaches it grammar, facts, styles and a surprising amount of reasoning ability.
From Predictor to Assistant
A pretrained model continues text; it does not naturally follow instructions. Assistants are further trained with instruction tuning on examples of helpful responses, and often with reinforcement learning from human feedback to prefer answers people rate highly.
The Context Window
Everything the model can consider at once — instructions, documents, conversation history — must fit in its context window, measured in tokens. It has no memory between separate conversations unless an application provides one.
Strengths
Drafting and editing text, summarising, translating, answering questions about provided documents, classifying and extracting information, and writing and explaining code.
Limitations
- Hallucination: fluent but false statements, especially about specifics.
- Knowledge cut-off: no knowledge of events after training unless given in the prompt.
- Sensitivity to wording: small prompt changes can change results.
- Bias: patterns in training data can surface in outputs.
Using Them Well
Give clear instructions and the relevant information, ask for structured output when you need it, and verify anything that matters.