What makes a number system?
\(\mathbb{R} \quad\quad\quad \mathbb{C} \quad\quad\quad \mathbb{H} \quad\quad\quad \mathbb{O}\)

After complex numbers, the next number system to be found was the quaternions. They appear to be the third member of a sequence. The reals \(\mathbb{R}\), the complex numbers \(\mathbb{C}\), and then \(\mathbb{H}\), the quaternions. So what's the pattern  ? This is one of the more difficult spot the pattern puzzles in math, but the rule is surprisingly simple in the end: mutiplication by a fixed number should rotate and stretch space. We will start with this geometric requirement and go the other way, looking for algebraic rules that follow from it. It turns out that only four such algebras are possible. They exist in dimensions \( 1 \), \( 2\), \( 4 \) and \( 8 \).

A short history of the complex plane

It is worth mentioning how late the geometric interpretation of complex numbers arrived. Square roots of negative numbers show up in the sixteenth century as bookkeeping devices inside formulas for the roots of cubics. This is at a time when not even negative answers were accepted as valid solutions to problems. The first person to give them a systematic arithmetic was Rafael Bombelli around the year 1572. He could calculate with them, and he used them to obtain real answers, but they lacked any kind of form. Numbers represented quantities back then, not points in space.

Around 1685, John Wallis drew what is often considered the first number line, with negative numbers placed to the left of zero. This was roughly a century after Bombelli's discovery. The modern Cartesian plane builds on top of this idea and was standardized in Euler's work of 1748.

All this is a prerequisite to the interpretation of \(a+b\sqrt{-1}\) as a point on a plane. Caspar Wessel invented it near the end of the eighteenth century and was ignored. Jean Robert Argand published a version shortly after and was nearly ignored. The idea only became popularized by Gauss around 1830. Once this connection is made, it opens a new question. We live in three dimensional space. Is there a three dimensional number system ? Hamilton spent years on this problem, starting in the mid 1830s. At the time there were no vectors and no linear algebra, so an algebra of triples looked like the natural language in which three dimensional geometry and mechanics would eventually be written.

We will come back to Hamilton at the end. Let's try to think this through independently first.

What should be preserved ?

The reason to plot \(\mathbf{i}\) on the axis perpendicular to \(\mathbf{1}\), and not anywhere else, is that both multiplication and addition then get a very nice geometric structure. If each complex number is thought of as an oriented line segment, addition looks like tip to tail composition of the line segments.

Drag the tips of q and r.

The pattern for multiplication is more difficult to see from a single product \(\mathbf{q}\mathbf{r}\). Instead, fix \(\mathbf{q}\) and consider the map \(\mathbf{r}\longmapsto\mathbf{q}\mathbf{r}\). In other words, apply multiplication by \(\mathbf{q}\) to every point \(\mathbf{r}\) in the plane.

Drag q and r.
Drag q. The orange grid is the image of the plane under r → qr.

The origin stays fixed because \(\mathbf{q}\cdot\mathbf{0}=\mathbf{0}\) and the point \(\mathbf{1}\) is sent to \(\mathbf{q}\), since \(\mathbf{q}\cdot\mathbf{1}=\mathbf{q}\). The rest of the plane is carried along, rotating around the origin and uniformly stretching. The idea can also be thought of like this. If you want to find the product \(\mathbf{q}\mathbf{r}\), then apply to \(\mathbf{r}\) the same rotation and scaling that carries \(\mathbf{1}\) to \(\mathbf{q}\). The resulting vector is \(\mathbf{q}\mathbf{r}\).

Both of these transformations preserve shapes. Addition can translates the plane and multiplication rotates and uniformly scales it. Can this concept be extended at face value to higher dimensions ?

30°
Drag q over the sphere, then drag the twist slider to spin the space about the axis through q. 1 still lands on q either way. Drag anywhere else to turn the view.

Addition can easily behave in the same way, it's vector addition. Multiplication is different. In the plane, the requirement to send \(\mathbf{1}\) onto \(\mathbf{q}\) determines the transformation uniquely. In three dimensions, it doesn't. After turning \(\mathbf{1}\) toward \(\mathbf{q}\), we can still twist the whole space by any angle around the axis through \(\mathbf{q}\). The requirement only restricts what multiplication could be. We could keep adding conditions until this freedom is gone, but how would we decide which ones ?

The structure is already quite interesting on its own. Instead, what if we try to find all algebraic systems that fulfill it ? Anything, that doesn't look right can be thrown out at the end.

The goal. Find every space \(V \cong \mathbb{R}^n\), with lengths and angles as usual and a multiplication \(V \times V \to V\) such that

  1. there exists a multiplicative identity, \(\mathbf{1}\mathbf{r} = \mathbf{r}\mathbf{1} = \mathbf{r}\) for every \(\mathbf{r}\), and
  2. left and right multiplication by any vector rotates and stretches \(V\) about the origin.

Since \(V\) is a vector space, we borrow vector addition from it, and scaling comes along for free too. So, for instance, \(\mathbf{3} = \mathbf{1} + \mathbf{1} + \mathbf{1} = 3\cdot\mathbf{1} \). By convention, \(\mathbf{1}\) is just some vector in the space, and we'll keep it separate from \(1\). The number \(1\) is only a scalar or a distance. However, they need to be tied together through \(\|\mathbf{1}\| = 1\). That's nothing more than a choice of units in the space.

Nothing else is assumed. Commutativity, associativity and any other familiar algebraic properties can't be used without justification. With that, 2 solutions are already known: the real numbers in one dimension, and the complex numbers in two. Whether there are any others is the question.

The first clue

It might be daunting at first trying to make progress with the little we've given ourselves. Here's a clue that might help. If a planar object is rotating, how do you get the velocity of each point ? Easy, right. Take each position vector, rotate it by 90 degrees and stretch it by angular velocity. It feels like a lucky coincidence that this is precisely the job of multiplying by vectors that happen to be on the line orthogonal to the axis spanned by \( \mathbf{1}. \)

Each tip of \(\mathbf{r}\) carries \(\mathbf{q}\mathbf{r}\). Drag q onto the vertical axis and those vectors become the velocities.

So getting the velocities in 2 dimensions can be done by rotating and stretching. Can the same be said about 3d ? We will come back to this question later.

How multiplication and addition work together

To get started, let's try to build the complex numbers, but only using the two rules. We'll do it first for some example case. Take any complex number \(\mathbf{q} = \mathbf{1} - 2\mathbf{i}\) and another \(\mathbf{r} = \mathbf{1} + \mathbf{i}\). This is with \(\mathbf{i}\) chosen, as usual, to be \(\mathbf{i} \perp \mathbf{1}\) and \(\|\mathbf{i}\| = 1\). The product \(\mathbf{q}\mathbf{r}\) can be found geometrically in two ways based on rule 2. Either with \(\mathbf{r} \to \mathbf{q}\mathbf{r} \) or \(\mathbf{q} \to \mathbf{q}\mathbf{r} \).

Turn 1 onto q, and r is carried along.
Turn 1 onto r, and q is carried along.

Great, both methods give the same answer ! If they didn't, the instructions would have been contradictory.

Now we want to know how to go from the geometry to a calculation. First, look more closely on the case where \(\mathbf{1}\) was rotated to \(\mathbf{q}\).

r in the transformed basis.

Before the transformation, \(\mathbf{r}\) can be found by going one step in the \(\mathbf{1}\) direction and one step in the \(\mathbf{i}\) direction. After multiplication by \(\mathbf{q}\), those directions become \(\mathbf{q} \cdot \mathbf{1}\) and \(\mathbf{q} \cdot \mathbf{i}\). The coordinates of \(\mathbf{r}\) stay the same, only the basis changed. \[ \{\mathbf{1},\mathbf{i}\} \longmapsto \{\mathbf{q}\cdot \mathbf{1},\mathbf{q} \cdot \mathbf{i}\} \] Therefore

\[ \begin{aligned} \mathbf{q}\cdot\mathbf{r} &= \mathbf{q}\cdot(\mathbf{1} + \mathbf{i}) \\ &= \mathbf{q}\cdot\mathbf{1} + \mathbf{q}\cdot\mathbf{i} \end{aligned} \]

But look what happened. This is the distributive property ! And it follows only from multiplication behaving like rotating and stretching. Naturally, the same argument also applies when multiplying from the right.

\[ \begin{aligned} \mathbf{q}\cdot\mathbf{r} &= (\mathbf{1} - 2\mathbf{i})\cdot\mathbf{1} + (\mathbf{1} - 2\mathbf{i})\cdot\mathbf{i} \\ &= \mathbf{1}\cdot\mathbf{1} - 2(\mathbf{i}\cdot\mathbf{1}) + \mathbf{1}\cdot\mathbf{i} - 2(\mathbf{i}\cdot\mathbf{i}) \end{aligned} \]

Simplify multiplication by \(\mathbf{1}\) and we get

\[ \mathbf{q}\cdot\mathbf{r} = \mathbf{1} - \mathbf{i} - 2(\mathbf{i}\cdot\mathbf{i}). \]

The last question is what to do about \(\mathbf{i}^2\). Let's also find out visually.

Drag 1 onto i. Whatever i is carried to is i squared.

It's \(-\mathbf{1}\). Plug it in.

\[ \begin{aligned} \mathbf{q}\cdot\mathbf{r} &= \mathbf{1} - \mathbf{i} - 2(-\mathbf{1}) \\ &= 3 \cdot \mathbf{1} - \mathbf{i} \end{aligned} \]

Great, that's exactly the result the animation gave. This is just an example, but everything works the same for any 2 vectors.

\[ \begin{aligned} \mathbf{q}\cdot\mathbf{r} &= (q_0\mathbf{1} + q_1\mathbf{i})\cdot(r_0\mathbf{1} + r_1\mathbf{i}) \\ &= q_0(\mathbf{1}\cdot\mathbf{1})r_0 + q_0(\mathbf{1}\cdot\mathbf{i})r_1 + q_1(\mathbf{i}\cdot\mathbf{1})r_0 + q_1(\mathbf{i}\cdot\mathbf{i})r_1 \\ &= (q_0 r_0 - q_1 r_1)\mathbf{1} + (q_0 r_1 + q_1 r_0)\mathbf{i} \end{aligned} \] (1)

And we're done. This is how complex numbers work. The key was that multiplication is linear: brackets distribute and constants can be factored out, just like with ordinary numbers. This reduces the problem to figuring out how to multiply the chosen basis vectors.

The general case

We will borrow the fact that rotating and stretching is always linear for any dimension of \(V\). Brackets can therefore always be expanded exactly as before. Writing the components of \(\mathbf{a}\) and \(\mathbf{b}\) in some basis \(\{\mathbf{e}_i\}\), we get

\[ \begin{aligned} \mathbf{a}\cdot\mathbf{b} &= \Big(\sum_i a_i \mathbf{e}_i\Big) \cdot \Big(\sum_j b_j \mathbf{e}_j\Big) \\ &= \sum_i \sum_j a_i b_j\, (\mathbf{e}_i \cdot \mathbf{e}_j) \end{aligned} \]

And finding an algebraic system of this sort again comes down to finding the multiplication table \(\mathbf{e}_i \cdot \mathbf{e}_j\).

There is also one other important thing to notice here. The multiplication is differentiable with respect to \(a_i\) and \(b_j\) ! That's very good. We didn't know that before. Suppose more specifically that \(\mathbf{a}\) and \(\mathbf{b}\) are some differentiable functions of time \(t\), so that each traces out a smooth path through space. Then the product \(\mathbf{c}(t) = \mathbf{a}(t)\mathbf{b}(t)\) is also a differentiable path through space. Using the product rule, we have

\[ \begin{aligned} \frac{d}{dt}\big(\mathbf{a}(t)\mathbf{b}(t)\big) &= \sum_i \sum_j \Big(\frac{d}{dt} a_i(t) b_j(t)\Big)(\mathbf{e}_i \mathbf{e}_j) \\ &= \sum_i \sum_j \big(\dot{a}_i(t) b_j(t) + a_i(t)\dot{b}_j(t)\big)(\mathbf{e}_i \mathbf{e}_j) \\ &= \dot{\mathbf{a}}(t)\mathbf{b}(t) + \mathbf{a}(t)\dot{\mathbf{b}}(t). \end{aligned} \]

So the product rule works as usual. In the last step, linearity was used again to separate everything back out.

You might already see how we're closing in on the earlier remark. Choose a path \(\mathbf{q}(t)\), so that \(\|\mathbf{q}(t)\| = 1\). The stretching constant is therefore \(1\), and this is a pure rotation in the map \(\mathbf{r}_0 \to \mathbf{q}(t)\mathbf{r}_0 = \mathbf{r}(t) \). Here's how it could look like in 3 dimensions.

0.00
Move t to progress time. q(t) traces the dashed circle through 1, and r0 is carried to r(t). This is one possible way that the effect of multiplying by q(t) could look like in 3d.

Now differentiate the path.

\[ \begin{aligned} \mathbf{r}(t) &= \mathbf{q}(t)\mathbf{r}_0 \\ \dot{\mathbf{r}}(t) &= \dot{\mathbf{q}}(t)\mathbf{r}_0 \end{aligned} \]

The second equation gives \(\dot{\mathbf{q}}(t)\) an additional role. It has to generate the entire velocity field of the rotation by applying it to the initial position \(\mathbf{r}_0\) . This is a very restrictive additional requirement put on multiplication. Consider what does it mean for the example case in 3 dimensions.

0.00
The whole velocity field. Every arrow is perpendicular to its own position vector, and the two circled points on the axis never move.

Every nontrivial rotation in three dimensional space has an axis. A line of points that have \( \mathbf{0} \) velocity. For any point \(\mathbf{r}_0\) on this axis, \[ \mathbf{0} = \dot{\mathbf{r}} = \dot{\mathbf{q}}\mathbf{r}_0. \] Thus multiplication by \(\dot{\mathbf{q}}\) would have to send a nonzero vector to zero. But multiplication by \(\dot{\mathbf{q}}\) must also rotate and uniformly stretch space. Such a transformation is invertible and cannot send a nonzero vector to zero. The 2 requirements are incompatible. We have therefore reached a contradiction, showing that no such number system is possible in dim \( 3\).

Starting from 1830s, mathematicians spent decades trying to find a satisfactory way to "multiply triples", but it never worked out. This one way to understand why.

Return to the general case. The velocity is currently expressed in terms of the initial position \(\mathbf{r}_0\), rather than the current position, which is a little bit annoying. A simple work around will be choosing a path that passes through \(\mathbf{1}\). Let that happen at \(t=0\).

\[ \begin{aligned} \mathbf{r}(0) &= \mathbf{1}\cdot \mathbf{r}_0 \\ \dot{\mathbf{r}}(0) &= \dot{\mathbf{q}}(0)\mathbf{r}(0). \end{aligned} \]

Now the velocity is determined by the current position. There is another important consequence of this choice. Since \(\mathbf{q}(t)\) stays on the unit sphere, its velocity is always perpendicular to its position: \[ \dot{\mathbf{q}}(t)\perp\mathbf{q}(t). \] In particular, at \(t=0\) \[ \dot{\mathbf{q}}(0) \perp \mathbf{1}. \] You can see this geometrically in the previous plot. Moreover, nothing here depended on the particular path. Given any \(\mathbf{u}\perp\mathbf{1}\), we can choose a path with \(\mathbf{q}(0)=\mathbf{1}\) and \(\dot{\mathbf{q}}(0)=\mathbf{u}\), which means that \( \mathbf{u} \) has to generate the velocities for this rotation. What looked like a coincidence in the complex plane has become a requirement.

To see what this implies in general, first restrict to \(\|\mathbf{u}\|=1\). Multiplication by \(\mathbf{u}\) is then a pure rotation, with no stretching. The additional condition is that velocities are perpendicular to positions. \[ \mathbf{u}\mathbf{r}\perp\mathbf{r} \] (2) In two dimensions, this forces multiplication by \(\mathbf{i}\) to be a quarter turn. Applying it twice reverses every direction, \( \mathbf{i}(\mathbf{i}\mathbf{r})=-\mathbf{r} \).

Drag 1 onto i, then onto -1. The black arrow is where r is carried.

Does the same conclusion hold in higher dimensions? Let's apply \(\mathbf{u}\) twice. The vectors \(\mathbf{r}\), \(\mathbf{u}\mathbf{r}\), and \(\mathbf{u}(\mathbf{u}\mathbf{r})\) satisfy \[ \mathbf{r}\perp\mathbf{u}\mathbf{r}, \qquad \mathbf{u}\mathbf{r}\perp\mathbf{u}(\mathbf{u}\mathbf{r}). \] Since multiplication by \(\mathbf{u}\) preserves lengths, there is also \[ \|\mathbf{u}\mathbf{r}\|=\|\mathbf{r}\|, \qquad \|\mathbf{u}(\mathbf{u}\mathbf{r})\|=\|\mathbf{r}\|. \] These three vectors span at most a three dimensional subspace, so the entire situation can be visualized in three dimensions, regardless of the dimension of V.

-150°
r and ur are fixed. u(ur) has to be perpendicular to ur and the same length as r.

The configuration of the 3 vectors has to be similar to one of the shown possibilities. So which one is it ? To figure it out, we'll add a few more vectors. The key observation is that once we know where \(\mathbf{u}\mathbf{r}\) is sent after another mutliplication by \(\mathbf{u}\), we know how the entire plane spanned by \(\mathbf{r}, \mathbf{u}\mathbf{r}\) transforms. Suppose some vector has coordinates \( ( s_0, s_1 ) \) in the \( ( \mathbf{r}, \mathbf{ur} ) \) basis, then \[ \mathbf{u}( s_0 \mathbf{r} + s_1 \mathbf{u}\mathbf{r}) = s_0 \mathbf{u} \mathbf{r} + s_1 \mathbf{u} ( \mathbf{u}\mathbf{r}). \] Is the condition \(\mathbf{r} \perp \mathbf{u}\mathbf{r}\) satisfied for all points of the plane ? Look at the case for \( \mathbf{r} + \mathbf{u}\mathbf{r} \).

-150°
Now with r + ur and its image u(r + ur). The angle between them is only a right angle at one point of the circle.

Only one point on the circle gives \(\mathbf{u}(\mathbf{r} + \mathbf{u}\mathbf{r}) \perp \mathbf{r} + \mathbf{u}\mathbf{r}\). That's the one where \(\mathbf{u}(\mathbf{u}\mathbf{r}) = -\mathbf{r}\). Since nothing was assumed about \( \mathbf{r} \), this has to be true in general. Multiplying by \( \mathbf{u} \) must be some strange higher dimensional quarter rotation, which if applied twice makes it to a half rotation which is an inversion. Equivalently, we also have \((\mathbf{r}\mathbf{u})\mathbf{u} = -\mathbf{r}\).

The same fact also falls out quite cleanly from the algebra. If you are comfortable with inner products, here is that argument as well. We have already seen geometrically why the result must be true, so it can also be safely skiped.

Theorem. Any operator \(\mathbf{u}\) that is both a rotation and a generator of a rotation squares to minus the identity.

\[ \mathbf{u}(\mathbf{u}\mathbf{r}) = -\mathbf{r} \] (3)

Proof. Multiplication by \(\mathbf{u}\) is a rotation, so it preserves the inner product between any two vectors.

\[ \langle \mathbf{u}\mathbf{x}, \mathbf{u}\mathbf{y} \rangle = \langle \mathbf{x}, \mathbf{y} \rangle \tag{1} \]

If \(\mathbf{u}\) is a generator of some rotation, then it must be antisymmetric.

\[ \langle \mathbf{u}\mathbf{x}, \mathbf{y} \rangle = -\langle \mathbf{x}, \mathbf{u}\mathbf{y} \rangle \tag{2} \]

Now chain the two results. Apply (2) to (1), then

\[ -\langle \mathbf{x}, \mathbf{u}(\mathbf{u}\mathbf{y}) \rangle = \langle \mathbf{u}\mathbf{x}, \mathbf{u}\mathbf{y} \rangle = \langle \mathbf{x}, \mathbf{y} \rangle \]

or equivalently,

\[ \langle \mathbf{x}, \mathbf{u}(\mathbf{u}\mathbf{y}) + \mathbf{y} \rangle = 0 \]

for every \(\mathbf{x}\). So the vector \( \mathbf{u}(\mathbf{u}\mathbf{y}) + \mathbf{y} \) is perpendicular to every vector in the space. Hence

\[ \mathbf{u}(\mathbf{u}\mathbf{y}) = -\mathbf{y}. \]

Building the multiplication tables

Now we have all the tools needed to start looking for a basis \(\{\mathbf{e}_i\}\). Theorem (3) will help in finding the products \( \mathbf{e}_i \mathbf{e}_j \) and once those are known bilinearity can be used to multiply any 2 vectors in the space . The process works for all solutions \(V\) at once. First, take \(\mathbf{1}\) as the first basis vector, then pick any unit vector \(\mathbf{i}\) perpendicular to \(\mathbf{1}\). (3) gives \(\mathbf{i}^2 = -\mathbf{1}\), so the multiplication table is exactly the same as for complex numbers.

×1i
11i
ii-1

In fact, the whole subspace behaves exactly like complex numbers. The product of any two of its vectors is given by (1). Anything you know about them has to work here too, like Euler's formula, for instance.

\[ \mathbf{e}^{\mathbf{i}\theta} = \lim_{n\to\infty} \left(\mathbf{1}+\frac{\mathbf{i}\theta}{n}\right)^n = \mathbf{1}\cos\theta + \mathbf{i}\sin\theta \]

If the dimension of \( V \) is \( n = 2 \), we're done and we've found the complex numbers. If \(n > 2\), then continue. Pick any unit vector \(\mathbf{j}\) perpendicular to the plane of \(\mathbf{1}\) and \(\mathbf{i}\) and add it to the basis.

×1ij
11ij
ii-1i j
jjj i-1

Here comes the first challenge. What should be done about \(\mathbf{i}\mathbf{j}\) ? Some things are known about it immediately. It is orthogonal to \(\mathbf{j}\) because \(\mathbf{i}\mathbf{r}\) is orthogonal to all \(\mathbf{r}\). There's also a broader argument. If initially the set \(\{\mathbf{1}, \mathbf{i}, \mathbf{j}\}\) is pairwise orthogonal, then after a rotation it remains that way.

\[ \{\mathbf{1}, \mathbf{i}, \mathbf{j}\} \xrightarrow{\ \mathbf{i}\,\times\ } \{\mathbf{i}, \mathbf{i}^2, \mathbf{i}\mathbf{j}\} = \{\mathbf{i}, -\mathbf{1}, \mathbf{i}\mathbf{j}\} \]

This means that \(\mathbf{i}\mathbf{j}\) is also orthogonal to \(-\mathbf{1}\) and \(\mathbf{i}\). The way this manifests itself in the multiplication table is that all vectors in a row are orthogonal to each other. The same goes for columns, by multiplication from the right.

So we have proved that \(\mathbf{i}\mathbf{j}\) is orthogonal to the whole subspace spanned by \(\{\mathbf{1}, \mathbf{i}, \mathbf{j}\}\). It is also a unit vector, since multiplication by \(\mathbf{i}\) preserves lengths. If the dimension is \(n = 3\), this is a contradiction. No nonzero vector is orthogonal to the whole space, and \(n = 3\) is excluded once again. If \(n > 3\), then \(\mathbf{k} = \mathbf{i}\mathbf{j}\) serves as the next basis vector, no matter where it happens to land in the rest of the space. The remaining entries follow from iteratively using (3). See how much of the table you can complete yourself before revealing the hints.

After working through everything, it was possible to fill in the whole table. Multiplying any two vectors in this subspace produces another vector in the same subspace. Hooray ! We've found the quaternions . If \(n = 4\), then this is the only possible solution. The work is not complete, though. We've proved that this is the only possibility, but not that it works. Those are separate questions. That last step would have been to show that this table is in agreement with our definition and everything is consistent. We'll leave that part out.

William Rowan Hamilton discovered the quaternions in 1843, but his route was quite a bit different from ours. He arrived at them mostly through geometric intuition and trial and error. The day after the discovery, Hamilton wrote to his friend John T. Graves about the ideas that had led him there. The letter was later published in the Philosophical Magazine. You might find it extra interesting after reading through our derivation. Here's a link.

In Graves response, he was impressed by the boldness of the idea but uneasy about how freely one could invent new imaginaries and give them new properties. In his reply, he said “If with your alchemy you can make three pounds of gold, why should you stop there?”

We ask ourselves the same question. Let's keep on going in the case where \(n > 4\). First, the number of basis vectors will grow quickly here, so it's better to number them \(\{\mathbf{i}, \mathbf{j}, \mathbf{k}\} \to \{\mathbf{e}_1, \mathbf{e}_2, \mathbf{e}_3\}\). The process is otherwise the same. Without loss of generality, add a new basis vector randomly and see where the products lead. In the following interactive table, it's possible to add new columns as they are needed. Either step through the solution or try to come up with it on your own (it's not going to be as easy this time).

It was possible to exclude dimensions \(n = 5, 6\) and \(7\), but then we hit a wall. The simple deduction rule used up until now doesn't complete the table anymore. The system could be underdetermined for all we know. Still, looking at what was possible to resolve, there are clear patterns suggesting which numbers should go where. For instance, all the occurrences of \(\mathbf{e}_3\) seem to form 2 diagonals.

\[ \mathbf{e}_6 \mathbf{e}_5 \overset{?}{=} \pm \mathbf{e}_3 \]

Can this be justified ? Yes, but we need one more identity first. Take any \(\mathbf{u}, \mathbf{v} \perp \mathbf{1}\). Since \(\mathbf{u}+\mathbf{v} \perp \mathbf{1}\), we can normalize it and apply (3).

\[ \begin{aligned} (\mathbf{u}+\mathbf{v})\big((\mathbf{u}+\mathbf{v})\mathbf{r}\big) &= |\mathbf{u}+\mathbf{v}|^2 \left[\frac{\mathbf{u}+\mathbf{v}}{|\mathbf{u}+\mathbf{v}|} \left(\frac{\mathbf{u}+\mathbf{v}}{|\mathbf{u}+\mathbf{v}|}\mathbf{r}\right)\right] \\[4pt] &= -|\mathbf{u}+\mathbf{v}|^2\mathbf{r}. \end{aligned} \]

Expanding the left hand side gives

\[ \mathbf{u}(\mathbf{u}\mathbf{r}) + \mathbf{u}(\mathbf{v}\mathbf{r}) + \mathbf{v}(\mathbf{u}\mathbf{r}) + \mathbf{v}(\mathbf{v}\mathbf{r}) = -|\mathbf{u}+\mathbf{v}|^2\mathbf{r}. \]

Using the theorem (3) on the first and last terms,

\[ -|\mathbf{u}|^2\mathbf{r} + \mathbf{u}(\mathbf{v}\mathbf{r}) + \mathbf{v}(\mathbf{u}\mathbf{r}) - |\mathbf{v}|^2\mathbf{r} = -|\mathbf{u}+\mathbf{v}|^2\mathbf{r}. \]

Now expand \( |\mathbf{u}+\mathbf{v}|^2 \)

\[ -|\mathbf{u}|^2\mathbf{r} + \mathbf{u}(\mathbf{v}\mathbf{r}) + \mathbf{v}(\mathbf{u}\mathbf{r}) - |\mathbf{v}|^2\mathbf{r} = - (|\mathbf{u}|^2 + |\mathbf{v}|^2 + 2\langle \mathbf{u}, \mathbf{v} \rangle) \mathbf{r}, \]

so the \(|\mathbf{u}|^2\) and \(|\mathbf{v}|^2\) terms cancel, leaving

Clifford relations. For \(\mathbf{u},\mathbf{v} \perp \mathbf{1}\),

\[ \mathbf{u}(\mathbf{v}\mathbf{r}) + \mathbf{v}(\mathbf{u}\mathbf{r}) = -2\langle \mathbf{u}, \mathbf{v} \rangle\mathbf{r}. \]

If we plug any 2 orthogonal basis vectors \(\mathbf{i}, \mathbf{j}\) into it, the result will be

\[ \mathbf{i}(\mathbf{j}\mathbf{r}) + \mathbf{j}(\mathbf{i}\mathbf{r}) = -2\langle \mathbf{i}, \mathbf{j} \rangle \mathbf{r} = \mathbf{0} \]

This condition will fill in the empty spaces.

\[ \mathbf{i}(\mathbf{j}\mathbf{r}) = -\mathbf{j}(\mathbf{i}\mathbf{r}) \]

Awesome ! We have found a new member of the sequence, the octonions for \( n = 8 \).

The person to beat us to it this time was John T. Graves, Hamilton's friend from earlier. His question about why stop at 3 wasn't just rhetorical. By December 1843 he had constructed the same eight-dimensional algebra, which he called the octaves. Arthur Cayley independently discovered it and published it in 1845. The story is described in more detail at John Baez's article on the octonions.

If \(n > 8\), keep on going.

And there it is. Every step along the way was either a choice we were free to make or a forced consequence and at the end we find that

\[ \mathbf{e}_1 ( \mathbf{e}_{10} \mathbf{e}_{4} ) = \mathbf{e}_{15} \qquad \text{ and } \qquad \mathbf{e}_1 ( \mathbf{e}_{10} \mathbf{e}_{4} ) = -\mathbf{e}_{15} \]

Both can't be true at once and the construction therefore breaks in dimension \(n \ge 16\). There are no more members of the sequence. The octonions are the last.

Conclusion

We can now answer the question we began with. If multiplication by every nonzero vector is required to act as a rotation followed by a uniform scaling and there exists a multiplicative identity, then the dimension of the Euclidean space can only be \[ 1,\quad 2,\quad 4,\quad\text{or}\quad 8. \] The corresponding algebras are the real numbers \( \mathbb{R} \), the complex numbers \( \mathbb{C} \), the quaternions \( \mathbb{H} \) named after Hamilton, and the octonions \( \mathbb{O} \). This result is known as the Hurwitz theorem.

In standard terminology, these are called the finite dimensional real normed division algebras. Their definition usually looks like this.

Normed division algebra. A finite dimensional real vector space \(A\) carrying a bilinear product \(V \times V \to V\), a multiplicative identity \(\mathbf{1}\), and a Euclidean norm \(\|\cdot\|\), such that

\[ \|\mathbf{x}\mathbf{y}\| = \|\mathbf{x}\|\,\|\mathbf{y}\| \qquad\text{for every } \mathbf{x}, \mathbf{y} \in V. \]

To see how this is the same, fix a nonzero \(\mathbf{x}\) and consider the map \(\mathbf{y} \mapsto \mathbf{x}\mathbf{y}\). This map multiplies every length by the same factor \(\|\mathbf{x}\|\). And by bilinearity also every distance.

\[ \| \mathbf{x}\mathbf{y} - \mathbf{x}\mathbf{z} \| = \| \mathbf{x}(\mathbf{y}-\mathbf{z}) \| = \| \mathbf{x} \| \| (\mathbf{y}-\mathbf{z}) \| \]

All distances are scaled by the same factor and angles are preserved. In our derivation, we went the other way, getting linearity from orthogonal transformations. It's the same as what we had started with except for one detail. It doesn't exclude the possibility that this transformation reverses orientation. Reflections are allowed as well, not just rotations. Still, there are no reflections at the end. They can be excluded as a possibility by a simple continuity argument.

The name division algebras is because every nonzero vector has a multiplicative inverse. To see how the inverse is constructed, take a nonzero vector \(\mathbf{q}\) and decompose it in a plane which includes \(\mathbf{1}\) using a \( \mathbf{e} \perp \mathbf{1} \), \( | \mathbf{e} | = 1 \). \[ \mathbf{q} = q_0\mathbf{1} + q_1\mathbf{e}. \] Define the conjugate of \(\mathbf{q}\) by \[ \overline{\mathbf{q}} = q_0\mathbf{1} - q_1\mathbf{e}. \]

Since \(\mathbf{e}^2 = -\mathbf{1}\), multiplying \(\mathbf{q}\) by its conjugate gives \[ \overline{\mathbf{q}}\,\mathbf{q} = \mathbf{q}\,\overline{\mathbf{q}} = (q_0^2 + q_1^2)\mathbf{1} = \lVert\mathbf{q}\rVert^2\mathbf{1}. \] Therefore \[ \mathbf{q}^{-1} = \frac{\overline{\mathbf{q}}}{\lVert\mathbf{q}\rVert^2} = \frac{q_0\mathbf{1} - q_1\mathbf{e}} {\lVert\mathbf{q}\rVert^2}. \] and \[ \mathbf{q}^{-1} ( \mathbf{q} \mathbf{r} ) = \mathbf{r} \]