explainer
A map of machine learning: tasks, signals, and models
Understand how supervised learning, self-supervision, reinforcement learning, NLP, and deep learning fit together through practical robot and language examples.
What you will learn
- Describe a learning setup by its domain, data modality, task, learning signal, and model family.
- Distinguish supervised, unsupervised, self-supervised, and reinforcement learning.
- Explain how pretraining and fine-tuning can use different signals in one model's development.
Before you start
- Familiarity with tables, images, and text as forms of data
An image classifier can belong to computer vision, use supervised learning, and be built with a deep neural network. All three descriptions can be correct. They answer different questions about the same system.
This map gives you a way to describe a project before choosing a model. We will apply it to robot sensors, images, actions, and language.
Ask five different questions
Start with the intended behavior, then work through these questions:
| Aspect | Question | Example |
|---|---|---|
| Domain or field | What area does the problem belong to? | Robotics |
| Data modality | What form does the data take? | Sensor measurements in a table |
| Task | What should the system produce? | A fault category |
| Learning signal | What information guides learning? | Inspection labels |
| Model family | What kind of learned model represents the solution? | A decision-tree ensemble |
These aspects overlap and constrain each other. Images have spatial structure that can influence an architecture. Available labels constrain which supervised tasks you can train. A field such as robotics can combine several modalities and tasks.
Use the map to describe a setup, then check whether its parts work together.
Find the learning signal
Supervised learning uses input examples paired with target outputs. Classification predicts a category, such as fault or no fault. A classifier may also output numeric probabilities. Regression predicts a numeric quantity, such as a sensor's calibration offset.
Unsupervised learning looks for structure without supplied target labels for each example. Clustering groups observations using a chosen representation and similarity criterion. A cluster does not automatically correspond to a meaningful real-world category.
Self-supervised learning constructs targets from the data itself. For example, hide image patches and learn to reconstruct them. The original image supplies the target pixels. Self-supervision is often grouped under the broader heading of unsupervised learning; naming the target's origin is more precise than arguing over the boundary.
Reinforcement learning learns behavior using rewards associated with actions and their consequences. The objective concerns expected cumulative reward, which can include delayed effects. Learning may use live interaction or previously collected experience. An expert's action labels instead support supervised imitation learning.
Google's introduction provides examples of supervised, unsupervised, and reinforcement learning. These names describe learning setups; they do not prescribe a single model family.
Explore a concrete setup
Choose a scenario and read why each description fits. Each example proposes an approach that would need evaluation on real data.
The default sensor example uses supervised classification with a decision-tree ensemble. Compare it with grouping operating patterns, which uses unlabelled measurements and centroid-based clustering.
Then select Learn from robot images. Switch from Self-supervised pretraining to Supervised fine-tuning. The task and learning signal change while the image modality and neural-network family remain the same.
Describe a robot fault detector
Suppose a robot records a summary for each short operating window: return strength, temperature, and the fraction of missing sensor readings. Later inspection identifies whether a sensor fault was present.
One training row contains those three measurements and an inspection label. The input is a numeric feature vector. The target is the category fault present or fault absent.
We can now describe the setup precisely: robotics domain, tabular sensor data, binary classification task, supervised learning signal, decision-tree ensemble candidate.
A random forest can learn threshold rules and combine several trees. Logistic regression provides another candidate using a linear score. The vocabulary alone cannot establish which will work better on these measurements.
Check that every input will be available when a prediction is needed. Including a later technician diagnosis among the features would give the training model information the deployed detector lacks.
Now change the target to the calibration offset found during inspection. Predicting that numeric quantity becomes regression. The domain, input format, and supervised learning signal can stay the same. Changing the task may also call for a different model or evaluation measure.
Separate fields from data formats
Natural language processing, or NLP, studies computational work with language. Its tasks include classification, translation, retrieval, and generation. The field includes many training algorithms; the NLP lesson starts with tokens and a small language model.
Computer vision concerns visual information. Robotics also includes control, planning, physical interaction, and combinations of perception systems. A robot can use a camera image, joint measurements, and a spoken instruction together. That is a multimodal setup.
Modality names also depend on representation. A sensor stream has an order through time; summarizing it into table columns can discard timing information. Calling both versions “numeric data” does not make them interchangeable inputs.
Name the model family
Linear score models, decision trees, and neural networks make different assumptions about how inputs map to outputs. Tree ensembles combine trees. Deep learning uses neural networks with multiple layers that can learn successive representations.
A neural network can support classification, regression, self-supervision, or reinforcement learning. NLP can also use models built from counts or linear scores. Neither field nor signal uniquely determines the architecture.
Foundation model refers to a model trained on broad data at scale that can be adapted to a range of downstream tasks. This is the definition proposed by Stanford's foundation-model report. Self-supervised pretraining is common in such models, but the definition is not simply “a transformer” or “any large neural network.” Our small image example does not establish the breadth needed for that description.
Training also takes different forms. Neural networks commonly use gradient-based optimization. A tree can grow by selecting splits; k-means alternates assignments with updates to cluster centers. The scikit-learn documentation describes these k-means steps. Optimization does not always mean gradient descent.
Follow the training stages
A complete training history may involve several signals. In the explorer's image example, pretraining hides patches and learns to reconstruct them. That objective creates targets from existing images, so it is self-supervised. Masked Autoencoders demonstrates this kind of image pretraining.
Next, add an obstacle classifier and fine-tune the encoder using labelled images. Those labels supply a supervised signal. The resulting detector inherits parameters from self-supervision and receives further supervised training.
Language models can follow a related sequence. The original BERT paper describes pretraining on unlabelled text and fine-tuning for downstream tasks. The exact objectives differ from image reconstruction.
Always specify the stage. Saying only “this is a supervised model” can omit how its starting representation was learned. Pretraining also does not guarantee useful transfer to your data; that needs evaluation.
Decide what success means
The map helps formulate experiments. It does not replace them.
For the fault detector, measure missed faults and false alarms on separate robot runs. Randomly splitting overlapping windows from the same run can let near-duplicates cross the training and test boundary. A split should reflect the future conditions you want to predict.
For clustering, inspect whether groups are stable and useful for the intended investigation. Reducing within-cluster distance alone does not establish a meaningful operating category. For a reaching policy, check achieved goals and operating constraints as well as the reward used in training.
Begin with the data and an explicit evaluation question. The pandas lesson covers inspecting and transforming a table, including missing values. Those choices help determine what the model can learn from the examples you give it.
Try it yourself
Exercise 1. A system counts words in support messages and trains logistic regression on reviewed labels: hardware issue or other. Name its field, modality, task, signal, and model family. What changes if the target becomes the number of hours until resolution?
Show solution: separate the output from the input
The field is NLP, the original modality is text, the task is binary classification, the signal is supervised, and the family is a linear score model. Word counts are the numeric representation of the text.
Predicting resolution time is regression. The field, input representation, and supervised signal can remain the same, but logistic regression is a classifier and would need replacement or reformulation. These descriptions do not make either setup deep learning.
Exercise 2. You pretrain an image encoder by hiding patches, then fine-tune it on expert-labelled obstacle images. Name the signal at each stage. Has switching signals changed the modality? Does successful reconstruction prove the obstacle detector will work?
Show solution: follow where each target comes from
Pretraining is self-supervised because the original image supplies the hidden pixels. Fine-tuning is supervised because expert labels supply the obstacle targets. Both stages use images and a neural-network encoder.
Reconstruction performance does not establish obstacle-detection performance. Evaluate missed obstacles and false alarms on new scenes that represent the intended use.
Sources and further study
- Google: What is machine learning?, for introductory learning setups and examples.
- He et al.: Masked Autoencoders Are Scalable Vision Learners, for learning visual representations through masked image reconstruction.
- Devlin et al.: BERT, for language-model pretraining followed by adaptation to downstream tasks.
- Bommasani et al.: On the Opportunities and Risks of Foundation Models, for the proposed foundation-model definition and its scope.
- scikit-learn: K-means, for centroid-based clustering and its iterative optimization.