Multimodal models work with more than one type of data: text, images, audio, video or documents.
What They Can Do
- Describe and answer questions about images: charts, screenshots, photos, diagrams.
- Read documents: extract information from scanned forms, invoices and PDFs.
- Transcribe and understand speech, and respond with synthesised speech.
- Analyse video frames and events.
- Generate images or audio from text.
How They Work
Separate encoders turn each modality into representations the core model can process — for example, splitting an image into patches embedded like tokens. The model then reasons across them together.
Practical Uses
- Document processing: turning scans into structured data.
- Accessibility: describing images and transcribing audio.
- Quality inspection from photos.
- Customer support that handles screenshots.
- Analysing charts and dashboards.
Tips
- Provide images at adequate resolution, cropped to what matters.
- Ask specific questions rather than "describe this".
- For documents, ask for structured output and validate it.
- Check fine details — small text, exact numbers — which are common failure points.
Limitations and Risks
Models can misread text, miscount objects or invent details. Images can contain hidden prompt-injection text. Be careful with personal data such as faces and handwriting, and check data-handling terms before uploading sensitive material.