Retrieval systems give models access to organisational knowledge. Every layer needs protection.
Ingestion
- Control which sources are indexed; review new sources.
- Scan documents for secrets, malware and hidden instructions.
- Record provenance for every chunk.
Storage
- Vector databases hold embeddings and often original text. Protect them like the source data.
- Encrypt data and restrict network access.
- Embeddings can sometimes be inverted to recover approximate text; don't treat them as anonymised.
Retrieval
- Filter results by user permissions.
- Limit the number and size of retrieved chunks.
- Separate retrieved content from instructions in prompts.
Poisoning
Attackers who can add content — through shared drives, wikis or public sources — may plant misleading or malicious documents. Restrict write access to indexed sources and monitor changes.
Deletion
When documents are deleted or permissions change at the source, update the index promptly.
Auditing
Log queries and retrieved documents to investigate leaks and poisoning.