When a general-purpose model doesn't do what you need, there are three main levers. They solve different problems.
Prompting
Adjust instructions, provide examples and specify formats.
Best for: behaviour, tone, format and task definition. Pros: fast, cheap, easy to change. Cons: limited by context length; knowledge must be included each time.
Retrieval-Augmented Generation (RAG)
Retrieve relevant documents at query time and include them in the prompt.
Best for: answering from specific, changing or private knowledge, with citations. Pros: knowledge stays current by updating documents; answers can be traced to sources. Cons: quality depends on retrieval; adds infrastructure.
Fine-Tuning
Train the model further on your examples.
Best for: consistent specialised behaviour or style at scale, specific output formats, or making a smaller model perform a narrow task well. Pros: can reduce prompt length and cost; behaviour is built in. Cons: needs good training data; slower to change; not an efficient way to add frequently changing facts.
How to Choose
- Start with prompting and a solid evaluation set.
- If the model lacks knowledge, add RAG.
- If behaviour still isn't consistent enough, or you need a cheaper model to match a larger one, consider fine-tuning.
Combine Them
Many production systems use all three: a fine-tuned model, a carefully written prompt and retrieved context.