A quick reference to the vocabulary used across AI projects.
Core Concepts
- Algorithm — a procedure for learning from data, such as gradient boosting.
- Model — the trained result that makes predictions.
- Parameters / weights — numbers learned during training.
- Hyperparameters — settings chosen before training.
- Training / inference — learning from data / using the model on new inputs.
- Features / labels — model inputs / the answers it learns to predict.
Evaluation
- Accuracy — share of correct predictions.
- Precision / recall — of positive predictions, how many were right / of actual positives, how many were found.
- Overfitting — memorising training data instead of learning the pattern.
- Baseline — a simple reference result every model must beat.
- Benchmark — a standard test used to compare models.
Deep Learning
- Neural network — layers of connected units that learn representations.
- Transformer — the architecture behind modern language models, built on attention.
- Embedding — a vector of numbers representing meaning.
- Fine-tuning — further training a pretrained model on specific data.
- Transfer learning — reusing knowledge from one task for another.
Language Models
- LLM — large language model.
- Token — a chunk of text the model processes.
- Context window — the maximum tokens per request.
- Prompt — the input given to a model.
- Hallucination — confident but false output.
- RAG — retrieval-augmented generation: grounding answers in retrieved documents.
- Agent — an LLM that uses tools in a loop to complete tasks.
Responsible AI
- Bias — systematic unfairness in data or outputs.
- Explainability — understanding why a model made a prediction.
- Model card — documentation of a model's use, data and limitations.
- Drift — degradation as real-world data changes.