The dot product: angles, projections, and robot motion

Learn the dot product with an interactive vector diagram, a worked projection example, and a robot heading calculation. Includes exercises and solutions.

By 9 min read

What you will learn

  • Compute a dot product and explain its sign.
  • Project a vector onto a direction and check the result.
  • Choose between raw dot product and cosine similarity.

Before you start

A robot can face one direction while moving partly sideways. Its total speed does not tell you how fast it moves forward. To answer that, compare its velocity with the direction it faces.

The dot product measures how two vectors line up, while accounting for their lengths. It takes two vectors and returns one number. The same calculation helps measure motion, compare machine learning features, and choose a direction for an optimization step.

Start with arrows on a flat plane. Once the geometry makes sense, the arithmetic extends to any number of coordinates.

Multiply matching coordinates

A two-dimensional vector has an x component and a y component. For a = (3, 0), move three units right and zero units up. For b = (2, 2), move two units right and two up.

Multiply the matching components, then add:

a · b = axbx + ayby

Here, the result is 3 × 2 + 0 × 2 = 6. That six is a scalar: a single number with no direction attached.

For three dimensions, add the z-coordinate product. For a feature vector with 100 entries, add 100 matching products. Both vectors must have the same number of entries, and matching positions must refer to matching features or axes.

When you group many feature vectors or images, their axes need names and a consistent order. The tensors lesson shows how shape and indexing keep those values organized.

The dot product is symmetric: a · b = b · a. Swapping the inputs leaves the sum unchanged.

Read the angle and sign

For two nonzero vectors, the coordinate calculation also has a geometric form. The symbols ‖a‖ and ‖b‖ mean vector lengths; θ is the angle between them.

a · b = ‖a‖ ‖b‖ cos θ

This gives the sign a clear meaning:

  • Positive: the angle is below 90°. Part of one vector points along the other.
  • Zero: two nonzero vectors meet at 90°. They are perpendicular, or orthogonal.
  • Negative: the angle exceeds 90°. Part of one vector points against the other.

Change b in the experiment. Try “Perpendicular,” then “Opposed.” Next, compare “Worked example” with “Same angle, longer b”: the dot product doubles while the angle stays at 45°.

Interactive experiment

How much of b points along a?

Change either vector. The amber segment shows the projection of b onto a.

Two vectors and a perpendicular projectionVector a is (3, 0) and vector b is (2, 2). Their dot product is 6.00. The projection of b onto a is (2.00, 0.00).-4-4-2-22244xyab
Solid arrow: a. Dashed arrow: b. Amber segment: projection. Each grid square is one unit.

The projection points along a. The dot product is positive.

Dot product a · b
6.00
Angle
45.0°
Cosine similarity
0.707
Signed projection length
2.00
Projection vector
(2.00, 0.00)

Find the part along a direction

Imagine dropping a perpendicular from the tip of b onto the line through a. The meeting point marks the projection of b onto a. In the experiment, the amber segment runs from the origin to that point.

Two quantities describe this projection. The scalar projection gives its signed length along a:

compa(b) = (a · b) / ‖a‖

The vector projection gives both its length and its direction:

proja(b) = [(a · b) / (a · a)] a

Both formulas require a nonzero a. A negative scalar projection means the projected vector points against a. Projecting the zero vector onto a nonzero direction gives the zero vector.

Subtract the projection from b to get the residual. That leftover vector is perpendicular to a. This condition gives you a useful check for arithmetic errors. The projections and least squares lesson uses it to find the closest output a linear model can reach. MIT's lecture on projections develops that connection.

In linear regression, residuals measure the gaps between observed values and a model's predictions. Squaring and adding those gaps gives a loss that you can use to fit a line.

Check a complete example

Return to a = (3, 0) and b = (2, 2). Work through the result before checking the display:

  1. Dot product: 3 × 2 + 0 × 2 = 6.
  2. Lengths: ‖a‖ = 3 and ‖b‖ = √8 ≈ 2.828.
  3. Cosine: 6 / (3 × √8) ≈ 0.707.
  4. Angle: arccos(0.707…) = 45°.
  5. Scalar projection: 6 / 3 = 2.
  6. Vector projection: (6 / 9)(3, 0) = (2, 0).

The residual is (2, 2) − (2, 0) = (0, 2). Its dot product with a is zero, which confirms the right angle.

The geometry now matches the coordinates: two units of b point right, and two point up. Multiplying the forward part by the length of a gives the original dot product, 2 × 3 = 6.

Measure a robot's forward motion

Suppose a robot faces a direction described by h = (0.6, 0.8). This is a unit vector because √(0.6² + 0.8²) = 1. Its measured velocity is v = (2, 1) meters per second.

Use the same world coordinate frame for both vectors. The robot's forward velocity component is:

v · h = 2(0.6) + 1(0.8) = 2 m/s

Its total speed is √5 ≈ 2.236 m/s. Its forward speed component is 2 m/s. The difference comes from motion across its heading.

Choose l = (−0.8, 0.6) as the robot's left direction. It has unit length and sits perpendicular to h. Then v · l = −1 m/s, so the robot has a one-meter-per-second component toward its right.

You can reconstruct the original velocity: 2h − l = (2, 1). This works because h and l form an orthonormal pair: perpendicular vectors of unit length. Robotics uses these axes to describe orientation and change coordinate frames. Modern Robotics explains their role in rotation matrices.

Compare direction with cosine similarity

Raw dot products mix length with alignment. Cosine similarity divides out both lengths:

cosine(a, b) = (a · b) / (‖a‖ ‖b‖)

For nonzero vectors, it ranges from −1 to 1. A value of 1 means the same direction; −1 means opposite directions.

The lengths in this formula are L2 norms. Vector norms and normalization works through that choice, compares it with L1 and L∞, and explains why the zero vector has no normalized direction.

Take a query vector q = (1, 0) and two candidates: u = (1, 0) and v = (2, 1). Their dot products with q are 1 and 2, so v scores higher. Their cosines are 1 and about 0.894, so u has the closer direction.

For embeddings, use the similarity function the model supports and test it on your retrieval task. Unit-length embeddings make dot product equal cosine similarity. Normalization behavior depends on the model and settings; Sentence Transformers documents when that equivalence applies.

A cosine score is not a probability. There is no universal score that means two texts match. Choose a threshold using labeled examples from your application, and measure which useful results it misses or accepts incorrectly.

Use conditional probability when your question concerns an event within a defined reference group.

The natural language processing lesson follows text from tokens to counts and introduces vector representations. A shared vector space makes dot products possible; the representation and training objective determine what those comparisons mean.

Try it yourself

For each exercise, write your answer before opening the solution.

Exercise 1. Let a = (1, 2) and b = (4, 1). Find a · b, project b onto a, and check the residual.

Show solution: projection and residual

The dot product is 1 × 4 + 2 × 1 = 6. Since a · a = 5, the vector projection is (6/5)(1, 2) = (1.2, 2.4).

The residual is (4, 1) − (1.2, 2.4) = (2.8, −1.4). Check it against a: 1 × 2.8 + 2 × (−1.4) = 0. The residual is perpendicular to the projection direction.

Exercise 2. A robot faces h = (0, 1) and moves with velocity v = (3, −2) m/s. Find its forward velocity component and its total speed. Is it moving forward or backward relative to its heading?

Show solution: signed robot motion

The unit heading gives v · h = −2 m/s. The negative sign means backward motion along the heading. Its total speed is √(3² + (−2)²) = √13 ≈ 3.606 m/s, which includes its sideways motion.

The projected velocity vector is −2h = (0, −2) m/s. Subtract it from v to find the sideways component: (3, 0) m/s.

Next, use this geometry in gradient descent: a small step changes a function according to its dot product with the gradient.

Sources and further study