explainer
NLP: turn robot commands into token probabilities
Learn natural language processing through a small robot-command corpus. Compare tokenization, unigram and bigram counts, unseen contexts, embeddings, and evaluation.
What you will learn
- Explain what a tokenizer keeps and discards in a text representation.
- Calculate unigram and bigram probabilities from a command corpus.
- Preserve command boundaries and recognize an unseen context.
- Distinguish count models from learned embeddings and contextual representations.
Before you start
- Conditional probability and its denominator
- Counting, fractions, and percentages
“Move then stop” and “stop then move” contain the same words. Their order changes the requested behavior. A language system's representation determines whether it can even notice that difference.
Natural language processing, or NLP, studies computational methods for working with human language. We will build one small language model by counting robot-command tokens. Its predictions will be simple enough to inspect by hand.
Define the language task
Different language tasks ask for different outputs. A classifier might label a command's intent. A retrieval system might find a relevant manual passage. A language model estimates probabilities for token sequences or their next tokens.
Our task is narrow: given one previous token, estimate what follows it in a tiny command corpus. A corpus is a collection of text examples. Here it contains eight invented commands, including two copies of “move forward.”
Those counts describe patterns in the examples. They do not establish whether a command is feasible, safe, or appropriate for a robot's current state. The experiment does not control hardware.
The machine learning fields map places language processing alongside other application areas. One NLP application can combine several kinds of models.
Choose the units you will count
A tokenizer turns text into a sequence of units called tokens. A token can be a word, part of a word, a character, a byte-based unit, or a special boundary marker. Tokens and words are not interchangeable units.
Our educational tokenizer uses a small, explicit rule:
- Lowercase the English ASCII letters in each command.
- Keep a straight apostrophe inside a word, so
don'tremains one token. - Separate hyphenated parts, so
turn-leftbecomesturn,left. - Ignore one final period, question mark, or exclamation mark.
- Keep every word, including
doandnot.
Each corpus entry already holds one command. This helper does not find sentence boundaries in arbitrary documents. It rejects unsupported characters; it cannot process multilingual text as a general tokenizer would.
These choices fit the displayed examples. For another task, lowercasing or character normalization could remove useful distinctions. Removing frequent words can also erase negation: deleting “not” changes “do not move.”
Subword tokenizers split text into learned pieces, but the name alone does not guarantee complete coverage. Unknown tokens remain possible unless the vocabulary and fallback design cover the input. Hugging Face's tokenizer model reference explicitly exposes unknown-token and byte-fallback settings.
Decide which information to retain
A bag-of-words representation counts tokens while discarding their positions. “Move then stop” and “stop then move” therefore produce the same bag. Such counts can still help a topic classifier, but they cannot recover the order they discarded.
A unigram model estimates a next token from its overall count. It ignores the preceding tokens. A bigram model conditions on one previous token; a trigram model conditions on two.
An n-gram is a sequence of n tokens. An n-gram language model uses the preceding n − 1 tokens as its history. Jurafsky and Martin's language-model chapter develops these count estimates and their finite-history assumption.
This connects to the Markov property: the model limits how much history affects its next prediction. That is a modeling choice, not a claim that human language depends on only one preceding word.
Inspect the command corpus
We add START before each command and END after it. START supplies context for the first word. END is a prediction target that lets the sequence finish.
The eight commands contain 17 word tokens. Adding one END per command gives 25 prediction targets. The target vocabulary contains eight distinct words plus END, for nine possible output tokens; START is excluded.
The default table asks what follows move. Compare it with Unigram: no context, then switch back to the bigram model. The denominator changes because the reference group changes.
Choose Add two forward examples to repeat “move forward” twice more. This preset replaces the displayed corpus with ten commands; repeated clicks do not keep adding data. Watch the counts change while the tokenizer stays fixed.
Calculate the next-token probabilities
The bigram estimate divides a pair's count by the total number of observed transitions from its context:
P̂(w | h) = count(h, w) / Σᵥ count(h, v)
In the base corpus, move occurs four times. Two occurrences precede forward, one precedes left, and the last ends “do not move”:
| Next token after move | Count | Estimated probability |
|---|---|---|
| forward | 2 | 2 / 4 = 50% |
| left | 1 | 1 / 4 = 25% |
| END | 1 | 1 / 4 = 25% |
The END transition matters. Dropping it would change the denominator to three and overstate the probability of forward. We never count a transition from the end of one command to the beginning of another.
The unigram estimate asks a different question. There are two forward targets among 25 total targets, giving P̂(forward) = 2 / 25 = 8%. The bigram gives P̂(forward | move) = 50% because it restricts the reference group to transitions after move.
This is the denominator reasoning from conditional probability. Neither fraction is a score for the robot's understanding.
With the two extra examples, move has six outgoing transitions. Four lead to forward, so its probability becomes 4 / 6 ≈ 66.7%. The corpus's repetition directly affects the estimate.
Handle missing evidence honestly
Choose dance (unseen). The corpus contains no transitions from that context. The denominator is zero, so this unsmoothed conditional estimate is undefined. The experiment supplies no fallback distribution.
That differs from stop following move. The context move has observations, but none continue with stop. This model assigns that continuation probability zero. Zero records missing examples under the estimator; it does not prove that the phrase is impossible.
Choose END (command finished) for a third case. END deliberately terminates a command, so there is no next-token event to predict. Treating END as an unseen ordinary word would lose this boundary convention.
Practical count models can smooth probabilities or back off to a shorter context. Those methods add assumptions beyond raw counts. They should be explicit and evaluated on new data.
The bigram model also cannot distinguish the prefixes “move” and “do not move” after reading their final token. Both leave it with context move and the same next-token distribution. Its one-token memory discards the earlier negation.
Connect counts to embeddings and context
An embedding represents a token with a learned vector of numbers. An input embedding lookup returns the same vector for the same token ID. Later model layers can build a contextual representation that also depends on surrounding tokens and their positions.
Attention supplies one way to combine information from other positions. The Transformer described in Attention Is All You Need forms weighted combinations of value vectors using query-key compatibility scores. Its scaled dot-product version uses the dot product, followed by normalization into weights.
Those contextual vectors can retain information beyond one token. Which positions they can use depends on the model's attention mask; a causal next-token predictor must exclude future tokens. Longer context gives a model more information to work with, without guaranteeing a correct interpretation or action.
The explorer learns count tables. It does not train embeddings, attention layers, or a neural language model.
Evaluate the task you care about
Checking the training counts verifies the implementation. To test prediction quality, keep separate commands that did not contribute to training. Similar or duplicated commands across the split can make performance look better than it is on new requests.
For next-token prediction, a common measure averages the negative log probability assigned to the observed next tokens. Using natural logs, perplexity exponentiates that average. Compare it under the same tokenization, boundary convention, and evaluation corpus so the prediction units match.
An unseen continuation with probability zero gives infinite negative log loss. That exposes the unsmoothed model's limitation. Guessing a probability after seeing the test answer would invalidate the evaluation.
For a robot-command application, also test the intended task directly. Check negation, reordered actions, unfamiliar wording, and requests outside the supported command set. A model can predict common wording well while failing those distinctions.
Reproduce the counts in Python
This standard-library example recreates the base corpus and its boundaries. The regular expression implements the word-token rule for these supported commands.
import re
from collections import Counter, defaultdict
commands = [
"move forward", "move forward", "move left", "turn left",
"turn right", "stop", "do not move", "do not turn",
]
start, end = "<s>", "</s>"
unigram = Counter()
bigram = defaultdict(Counter)
for command in commands:
tokens = re.findall(r"[a-z]+(?:'[a-z]+)*", command.lower())
unigram.update(tokens + [end])
sequence = [start] + tokens + [end]
for previous, following in zip(sequence, sequence[1:]):
bigram[previous][following] += 1
counts = bigram["move"]
denominator = sum(counts.values())
print("Prediction targets:", sum(unigram.values()))
for token in ["forward", "left", end]:
label = "END" if token == end else token
count = counts[token]
print(f"{label}: {count}/{denominator} = {count / denominator:.1%}")
print(f"Unigram forward: {unigram['forward'] / sum(unigram.values()):.1%}")
assert "dance" not in bigram
assert end not in bigram
Expected output:
Prediction targets: 25
forward: 2/4 = 50.0%
left: 1/4 = 25.0%
END: 1/4 = 25.0%
Unigram forward: 8.0%
Check your understanding
Exercise 1. In the base corpus, find the next-token distribution after turn. Include command endings. How does it differ from the distribution after not?
Show solution: count both contexts
The three occurrences of turn precede left, right, and END once each. Each continuation has probability 1 / 3. The two occurrences of not precede move and turn once each, giving 1 / 2 to each. The contexts use different denominators.
Exercise 2. What probability does the base bigram model assign to move at the start of a command? Should the model count END → move between two separate commands?
Show solution: preserve command boundaries
Three of the eight commands begin with move, so P̂(move | START) = 3 / 8 = 37.5%. END has no outgoing transition. The next command starts with a fresh START marker, so joining the commands would introduce a pair that this model never observed within a command.
Sources and further study
- Jurafsky and Martin: N-gram Language Models, for count estimates, boundary markers, smoothing, and evaluation.
- Hugging Face Tokenizers: Models, for learned subword vocabularies, unknown tokens, and fallback settings.
- Vaswani and colleagues: Attention Is All You Need, for embeddings, attention, and causal masking in the Transformer architecture.