Language models don't remember anything between requests. Applications create the illusion of memory by sending relevant context each time.
Short-Term Memory: Conversation History
The simplest approach resends the conversation so far with each new message. It works until the history outgrows the context window or becomes expensive.
Managing Long Conversations
- Truncation: keep only the most recent messages.
- Summarisation: replace older messages with a running summary.
- Selective recall: keep key facts — the user's goal, decisions made — and drop small talk.
Long-Term Memory
To remember across sessions, applications store facts or past interactions and retrieve relevant ones later, often with embeddings and search. For example, remembering a user's preferred format or their project details.
Design Considerations
- Relevance: retrieve only memories that help the current request.
- Accuracy: stored memories can be wrong or outdated; let users correct them.
- Transparency: tell users what is remembered.
- Control: let users view and delete memories.
Privacy
Memory means storing personal data. Apply data minimisation, retention limits, access controls and your privacy obligations. Avoid storing sensitive information unless it's necessary and the user expects it.
Test Memory Behaviour
Include multi-turn scenarios in your evaluations: does the assistant use earlier context correctly, and does it avoid leaking one user's memory to another?