What Is Named Entity Recognition vs Structured Extraction
NER as a machine learning task: BIO tagging, softmax, CRF and span models, evaluation and hard cases, and how structured extraction generalises it, on video.
By TubeExtract team · 12 min read
Named entity recognition (NER) is one of the oldest and most studied tasks in natural language processing: find the spans of text that name real-world things, and say what kind of thing each one is. Structured extraction generalises it: instead of a fixed set of labels, the output follows a schema you define, with typed fields for the entities and for everything said about them.
This post covers NER as a machine learning problem: how it is framed, encoded, modelled, trained and evaluated, and where it breaks. It then shows how structured extraction extends it, and what changes when the text is the caption track of a video. One sentence runs through all of it:
We drove from Austin to Lockhart on Saturday to eat at Rosie’s Smokehouse, where the brisket was $32 a pound and honestly the best I’ve had this year.
NER as a machine learning task
Formally, NER maps a sequence of tokens x = (x₁ … xₙ) to a set of typed spans: triples (start, end, type) with the type drawn from a label set fixed in advance. The classic label set comes from the CoNLL-2003 shared task: person, organization, location and miscellaneous. OntoNotes 5.0 extends it to 18 types, adding dates, times, money, percentages, quantities, products, events and more.
For the example sentence, the target output is five spans:
Sequence labeling and the BIO scheme
Predicting a set of spans directly is awkward for most architectures, so NER is usually recast as sequence labeling: exactly one tag per token. In the BIO scheme (also written IOB2), a token is tagged B-TYPE when it begins an entity, I-TYPE when it continues one, and O when it is outside every entity. With K entity types the tag set has 2K + 1 labels. Spans are recovered by reading each B tag and the run of same-type I tags after it.
| token | We | drove | from | Austin | to | Lockhart | on | Saturday | to | eat | at | Rosie | 's | Smokehouse | , | where | the | brisket | was | $ | 32 | a | pound |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| tag | O | O | O | B-LOC | O | B-LOC | O | B-DATE | O | O | O | B-ORG | I-ORG | I-ORG | O | O | O | O | O | B-MONEY | I-MONEY | O | O |
The encoding admits invalid sequences: an I-ORG directly after O has no entity to continue. Models either learn to avoid such transitions or are constrained when decoding. Variants such as BIOES add explicit end (E) and single-token (S) tags, giving the model a direct signal for where an entity stops.
Tokenisation and label alignment
Transformer encoders split words into subword pieces, so “Smokehouse” may become several tokens while its label is defined for the whole word. The usual fix is to predict and train on each word’s first subword only, masking the rest out of the loss. Getting this alignment wrong does not raise an error; it silently shifts labels and lowers the score.
The models, in equations
Token classification with a softmax
An encoder turns each token into a contextual vector hᵢ. A linear layer and a softmax turn that vector into a distribution over the 2K + 1 tags, independently for every position:
Because each position is predicted on its own, nothing stops the argmax sequence from containing O I-ORG. In practice a strong encoder rarely does this, but it is not prevented.
A linear-chain CRF on top
A conditional random field scores the whole tag sequence instead. Each position contributes an emission score from the encoder, and each pair of adjacent tags contributes a learned transition score A[yᵢ₋₁, yᵢ]:
The normaliser Z(x) sums over an exponential number of sequences, but the chain structure lets the forward algorithm compute it in O(n·T²) for n tokens and T tags. The best sequence is decoded with the Viterbi algorithm at the same cost. Transitions such as O → I-ORG learn strongly negative scores, or are masked outright, so decoded spans are well formed.
Span-based scoring
Span models skip tags. They enumerate every candidate span up to a maximum width L, which is O(n·L) candidates, build a representation from the boundary tokens and the width, and classify it into a type or “not an entity”:
Because spans are scored independently, overlapping and nested entities are representable, at the cost of many more candidates, most of them negative. Generalist span models extend this by embedding the type names themselves and matching spans against them, so types can be supplied as text at inference time.
Generative extraction
Sequence-to-sequence and decoder models treat NER as generation: given the text and an instruction or template, they emit the entities as a string or JSON, token by token. The output format can be enforced with constrained decoding, masking at each step every token that would break the grammar. This is the formulation that makes structured extraction possible, because the output can be any schema, not only spans and labels.
How NER models evolved
- Rules and gazetteers. Pattern rules and lists of known names. High precision on what they cover, no recall beyond it, and costly to maintain per domain.
- HMMs and CRFs. Hidden Markov models, then linear-chain CRFs, over hand-engineered features: capitalisation, word shape, prefixes and suffixes, part-of-speech tags, gazetteer membership.
- BiLSTM-CRF. Bidirectional LSTMs over word and character embeddings replaced most feature engineering; the CRF on top kept sequences valid. Character-level inputs made unseen names tractable.
- Transformer encoders. Pretrained models such as BERT, fine-tuned with a token classification head. Self-attention sees the whole sentence, which is what separates Apple the company from apple the fruit, or Lincoln the person from Lincoln the city.
- Span and generalist models. Span scoring for nested entities, and models that accept entity types as text instead of learning a fixed set.
- Instruction-following extraction. Generative models that take a schema and return structured records, which is where NER and structured extraction meet.
Training data and evaluation
Supervised NER needs text annotated at the token or span level, which is slow and costly to produce, and annotation guidelines differ between datasets: whether “Saturday” is an entity, whether a title belongs inside a person span, whether a product gets its own type. A model trained on one dataset’s conventions is penalised when scored against another’s.
The standard metric is entity-level precision, recall and F1 under exact match:
A predicted entity counts as correct only if its start, end and type all match the gold annotation. Predicting “Smokehouse” instead of “Rosie’s Smokehouse” is therefore scored twice: a false positive for the wrong span and a false negative for the missed one. Token-level accuracy is a poor substitute: most tokens are O, so a model that tags nothing at all still scores high on it. The same imbalance affects training, which is why class weighting or sampling of O tokens is sometimes used.
Where NER gets hard
- Ambiguity. The same string is a person, a place or an organisation depending on context: Jordan, Washington, Amazon.
- Boundaries. “The University of Texas at Austin” is one organisation; a model that stops at “Texas” gets no credit under exact match.
- Nesting. “Bank of America” is an organisation containing a location. A single BIO sequence must choose one.
- Domain shift. Models trained on news underperform on medical notes, legal text, code or chat, where entity types and naming conventions differ.
- Noisy text. Social media, speech transcripts and automatic captions drop capitalisation and punctuation, the cues feature-based models lean on hardest, and spell names the way they sound.
- Long documents. Encoders have a fixed context window, so long texts are split into windows, and the same entity found in several windows has to be reconciled.
From NER to structured extraction
NER answers which names appear, and of what type. Applications usually need more: the relation between entities (the brisket was bought at Rosie’s Smokehouse, which is in Lockhart, not Austin) and attributes of them (the price, the speaker’s opinion). In the research literature these are separate tasks: relation extraction, slot filling, event extraction and entity linking, each with its own models and datasets.
Structured extraction folds them into one: the output is a set of records conforming to a schema. Each field has a name, a natural-language description that tells the model what belongs in it, and a type. An entity is one field of the record; its relations and attributes are further fields. For the example sentence:
Three things differ from NER output. The schema, not a label set, decides what is extracted: Austin and Saturday are dropped because nothing asked for them. Relations are carried by the record itself: city and dish belong to this restaurant. And values are normalised to their types: "$32" becomes the number 32, and the verdict is forced into one of three options, so thousands of records aggregate cleanly.
Validation and grounding
A tagger can only label tokens that exist. A generator can produce values that are not in the text at all, so structured extraction pipelines check the output after generation:
- Type validation. Every record is checked against the schema: numbers parse as numbers, enum fields hold an allowed option, dates are dates.
- Grounding. A
verbatim-stringvalue must occur in the source text, and a number must be traceable to it, including numbers spoken in words (“thirty-two dollars”). Values that cannot be traced are dropped rather than returned. - Echo filtering. A value that merely repeats the field’s own description is a known failure of instruction-following models and is removed.
- Deduplication. The same entity extracted twice, often under slightly different spellings, is merged into one record.
| Named entity recognition | Structured extraction | |
|---|---|---|
| Output | Typed spans | Records with typed fields |
| Label space | Fixed at training time | Defined per request by the schema |
| Relations and attributes | Separate tasks | Fields of the same record |
| Value types | Text spans only | String, number, integer, boolean, date, enum, lists |
| Typical formulation | Token classification, CRF or span scoring | Schema-conditioned generation plus validation |
| Evaluation | Entity-level F1, exact match | Field-level accuracy per record |
NER is a special case: a schema with one field for the entity and an enum field for its type. That is the schema in the request further down.
NER and structured extraction on video
A video’s spoken words are available as its caption track, written by the creator or generated by speech recognition. Extracting from it meets several of the hard cases at once: automatic captions come in short timed segments without capitalisation or sentence punctuation, names are spelled by sound (“palanteer” for Palantir), and a 30-minute video runs to tens of thousands of characters, well past a single context window.
A pipeline for it:
- Assemble. Merge the caption segments into one document, so sentences split across segments are whole again.
- Chunk. Split the document into windows on sentence boundaries, sized so each fits the model with room for the schema and the output.
- Extract. Run a schema-conditioned extraction model on every chunk, in parallel.
- Validate and ground. Check each record against the schema types and against the chunk it came from.
- Merge. Deduplicate records found in more than one chunk and combine their field values into one row.
TubeExtract runs this pipeline as an API. You send YouTube URLs with a schema, up to 200 videos per request, and receive typed, validated rows per video, at 10 credits per started minute of video; every account starts with 3,000 free credits.
For the output at scale, see every restaurant visited in 500 food videos. To write your own schema, start with the quickstart and how to write a schema.
no_captions and is not charged.Frequently asked questions
What is named entity recognition in machine learning?
Named entity recognition (NER) is the task of locating spans of text that refer to named things and classifying each span into a predefined type such as person, organization, location, date or money. It is usually trained as sequence labeling: every token gets a tag, and a BIO scheme marks where each entity begins and continues.
What is the BIO tagging scheme?
BIO (also called IOB2) labels each token B-TYPE if it begins an entity, I-TYPE if it continues one, and O if it is outside any entity. It turns span detection into per-token classification. Variants such as BIOES add explicit end and single-token tags.
Why do NER models use a CRF layer?
A per-token softmax predicts each tag independently, so it can output invalid sequences such as I-ORG after O. A linear-chain conditional random field scores the whole tag sequence, adding learned transition scores between adjacent tags, and decodes the best sequence with the Viterbi algorithm, which keeps predicted spans well formed.
Which models are used for NER?
Historically hidden Markov models and conditional random fields over hand-built features, then BiLSTM-CRF networks, then fine-tuned transformer encoders such as BERT with a token classification head. Span-based models score candidate spans directly, which handles nested entities, and generative models produce the entities as text or JSON from an instruction.
How is NER evaluated?
Usually with entity-level precision, recall and F1 under exact match: a prediction counts only if both the span boundaries and the type are correct. Standard English benchmarks include CoNLL-2003 (person, organization, location, miscellaneous) and OntoNotes 5.0 (18 entity types).
What is the difference between NER and structured extraction?
NER returns typed spans from a fixed label set. Structured extraction fills a schema the user defines: each output is a record with named, typed fields, which can hold an entity but also its attributes and relations, such as a price, a rating or which city a restaurant is in. NER is a special case of structured extraction with two fields, the span and its type.
Can you run NER on YouTube videos?
Yes, on the words spoken in them. A video's caption track is text, so NER and structured extraction run on it once it is assembled into a document. Automatic captions have no capitalisation and little punctuation, which removes cues many NER models rely on, so models that read the whole context do better on them.