Named entity recognition (NER) identifies and classifies spans of text that refer to entities — people, organisations, locations, dates, amounts, products.
Example
"Acme Ltd hired Priya Shah in Perth on 3 March 2025" contains an organisation, a person, a location and a date.
Approaches
- Rules and dictionaries: regular expressions for dates and amounts, lists of known names. Precise and cheap, but brittle.
- Statistical models: conditional random fields over hand-crafted features.
- Transformer models: fine-tuned models label each token with an entity type, and are the standard for high accuracy.
- Large language models: can extract entities from instructions and a few examples, including custom entity types, without training data.
Custom Entities
Business problems often need entities beyond the standard set: contract clauses, part numbers, drug names. Fine-tuning on labelled examples or prompting an LLM with clear definitions both work.
Evaluation
Compare predicted entity spans with labelled ones and compute precision, recall and F1, usually requiring exact span and type matches.
Practical Tips
- Define entity types clearly with examples, especially for ambiguous cases.
- Combine rules for rigidly formatted entities with models for everything else.
- Normalise entities after extraction — link "Acme", "Acme Ltd" and "ACME Limited" to one record (entity linking).
- Handle personal information carefully; NER is often used to find and redact it.