Lesson 11: Introducing Special Relativity
Introduction
Special relativity is the first place where the comfortable Euclidean geometry of everyday experience begins to come apart—and where the language of tensors we have been building becomes indispensable. In this lesson we take the leap from the familiar world of three-dimensional space and absolute time into the four-dimensional arena of Minkowski spacetime.
We begin with The Magic of the Speed of Light, the experimental fact that forces us to abandon the classical notions of absolute simultaneity and absolute length. From that single observational cornerstone flow The Postulates of Special Relativity, Einstein’s two deceptively simple statements that rewrite the rules of physics.
To keep the algebra clean we immediately adopt Natural Units, setting c=1 so that time and length are measured in the same units; this is not a mere convenience but a conceptual clarification. With that choice in hand we introduce the central geometric objects of the theory: Events, Spacetime, and the Spacetime Interval. The interval is the invariant that replaces the Euclidean distance, and it will be our constant guide.
Next we equip ourselves with the proper differential-geometric toolkit for this new arena: Tangent Vectors, One-Forms, and Tensors in Spacetime. Once those objects are under control, the Lorentz Transformations appear as the linear isometries that preserve the interval, and the Light-Cone Structure reveals the causal skeleton of spacetime.
We then organize the physical quantities of the theory into Four-Scalars and Four-Vectors, practicing the index gymnastics that will become second nature in general relativity. Finally, Doing this Stuff in Mathematica shows how the same formal manipulations can be performed symbolically, and a short Further Reading list points the way for those who wish to go deeper.
By the end of the lesson the stage will be set: we will have traded Newton’s absolute space and time for a single geometric arena in which the speed of light is the same for every inertial observer, and in which tensors speak the language of that arena with perfect clarity.
The Magic of the Speed of Light
The electric field vector
has the same function for electric fields that the gravitational field vector
has for the gravitational field. Similarly for the magnetic field vector
. Were we to begin with the equations governing the evolution of a wave in the electric and magnetic fields, we would be able to derive the Laplacian of the electric field for a place having no charge or current,
and we find the same combination of weird symbols,
, in the Laplacian of the magnetic field
It turns out that if we examine the ability of magnetic fields to pass through a vacuum, we encounter a number
, this is the quantity called the magnetic permeability, it has been given one of the symbols encountered above,
. There is also the quantity representing the ability for the electric field to pass through the vacuum, this is approximately
and is called the permittivity of space; and is given the other symbol from above,
. If we multiply them we have
The Ampere can be written as C/s
We then invert this product,
Strangely enough this has units of speed squared.
This makes sense since the coefficient of the second-order time derivative in the Laplacian should be the inverse wave propagation speed squared. Now, here is a remarkable thing, if we take the square root of this,
and we look at the right-hand result long enough we will realize that this is approximately the speed of light in a vacuum! Note that this speed is independent of the source or the observer—an experimental fact that Newtonian kinematics cannot accommodate. That observation is the true “magic” that forces the postulates of special relativity.
To put it another way, the square root of the inverse product between the ability of empty space to allow waves to propagate through the electric and magnetic fields is the speed of light in empty space. Looking at the Laplacian this tells us that the speed of light is the speed of an electromagnetic wave in empty space. This tells us that electricity, magnetism, and light are all aspects of the same thing.
Since the speed of light in a vacuum is determined by the use of fundamental constants, then it seems reasonable to postulate that the speed of light in a vacuum is itself a constant.
The numerical value of
was first obtained experimentally by Wilhelm Eduard Weber and Rudolf Kohlrausch in 1856, long before anyone suspected a connection with light. A decade later, Maxwell showed theoretically that the same combination of constants is precisely the speed that electromagnetic disturbances must propagate through empty space, and he identified that speed with the already-measured speed of light. Today, the logical order is reversed: the speed of light is defined to be exact, and the values of
and
are derived from it; the conceptual discovery, however, remains unchanged.
The Postulates of Special Relativity
Einstein formulated two postulates from which he derived the special theory of relativity:
Postulate 1: The laws of physics are the same in any inertial frame of reference.
This is just the principle of relativity due to Galileo, but Einstein’s version is stronger: it applies to all the laws of physics (including electromagnetism), not merely mechanics. This is equivalent to the statement that the physical relationships between quantities are geometric relationships between geometric objects.
Postulate 2: the speed of light in vacuum is the same in every inertial frame, independent of the motion of the source.
We have already seen a strong physical argument for this constancy: c is fixed by the electromagnetic properties of empty space itself. Once the postulate is accepted, a simple causality requirement—that an effect cannot precede its cause—implies that no signal or material object can travel faster than c. We adopt that implication as part of the working framework of the theory, but we do not elevate it to a third independent postulate.
Let the weirdness follow.
Natural Units
The first ramification of Postulate 2 is the ability to avoid having to write out the speed of light in SI units every time we need to do a calculation. The speed of light can be approximated as
to make it easier to write in SI. This is still unwieldy.
Now, we have to realize that SI units are not natural units. Nature did not assign the meter or the second as fundamental. We can choose whatever units we want so long as they are consistent. We can choose our unit of time to be the time it takes light to travel one meter and call this unit a meter of time.
The speed of light is then,
So, the speed of light is 1. This may seem startling, but it is a very natural way of looking at it. In fact, that is the name for this system of units, they are called natural units. There are other units in this system, but we will not concern ourselves with them for now.
Special Relativity, Events, Spacetime, and the Interval
We shall call any occurrence in the physical world an event. Events have no extent in either space nor time. As such they can be represented by a point.
The set of all events that will happen, have happened, and are happening now is what we mean by spacetime. We can denote this spacetime by the label M. In this way an event in spacetime is sometimes called a point in M. Here, of course, M is the state space in relativity.
It is important to note that spacetime in not the fiber-bundle of space and time from classical mechanics,
, that we have been dealing with up to now. It is much more complicated and subtle. We will use a label for this kind of situation and then give you a feel for it, and only later will we get to its true nature. We will call such a thing a spacetime manifold, for now we can think of this manifold as the space of special relativity.
We use a neat device for studying spacetime in relativity, it is called a spacetime diagram. It looks similar to a normal coordinate frame except the time axis is vertical and the spatial axes are perpendicular to it and each other (for three spatial dimensions this becomes impossible to visualize).
Figure 11.1. Spacetime diagram.
To identify the location of an event in spacetime we need four coordinates. We can write these
,
,
, and
. We can also write this as
where we understand that μ is allowed to be summed from 0 to 3. If we restrict such coordinate systems to be in uniform linear motion, we say that that coordinate system is an inertial frame. Each observer in uniform linear motion is at the origin of their own reference frame within a region of spacetime, and thus we can call them inertial observers, or just observers.
A single point in spacetime, P, is what we have called an event. In a spacetime diagram we have,
Figure 11.2. An event in spacetime.
Say that we have two inertial frames in the same region of M. Let’s say that we label one system as
and the other as
. For each such system we can measure every event in M. Since we can observe the same event on different spacetime diagrams, we should be able to use that event to place one spacetime diagram into another by drawing the relative coordinate axes. We can represent this on a spacetime diagram, assuming that the
system is moving with some velocity v with respect to the
system. In a spacetime diagram the
axis is the locus of constant
. Since
is moving at some velocity v, this axis will not longer be vertical, but will be tilted in the
direction.
Figure 11.3. Comparing the vertical axis of a moving frame and a relatively stationary frame.
In order for the
system to look at the event located at P at some time
, we must shine a light on it. This light originates from a point
.
Figure 11.4. An observer in the
system shining a light on an event..
Since the y system determines the coordinates of an event, as does the x system, we can write the y coordinates as functions of the x coordinates,
,
,
, and
. Clearly we cannot use the same general index for both x and y. We can use μ for x,
, and ν for y,
. Thus we have
. It is reasonable to assume that such a function is smooth,
, if we are careful in our choice of coordinates. If the y values are smooth functions of the x coordinates, then the inverse functions should also be smooth and we can write,
.
Exercise 11.1: Show this last sentence to be true.
From this we can see that we have the possibility of a large number of coordinate systems, any two of which can be smoothly related to one another. This leads us to the notion of a specific mathematical object, a four-dimensional manifold. It is important to understand that such an application of mathematical structure can never be proven in a rigorous way, it can only ever be assumed. There is never any guarantee that such a choice of structure will gain any advantage. Once it is made, though, we need to stick with it—that is how a theory is developed. We will be certain to assume that all statements about physical phenomena in spacetime will be considered about points in the spacetime manifold M.
If we allow an event to take some small time over a small distance to occur,
, called the spacetime interval. If we then choose units of distance and time so that the speed of light c does not change, we then write
(11.1)
and we say that the spacetime interval of an event is invariant. No matter what your inertial frame, the spacetime interval between events will be the same. Space and time will distort to make sure this interval is invariant.
If we take the square root of the spacetime interval,
(11.2)
then the quantity τ is what we call the proper time.
Let’s say that we are watching a particle move through M. Were we to follow the motion of this particle through its evolution of states in spacetime, we would see each point at some given proper time P(τ) evolve into a smooth curve. We call such a curve in spacetime a world line. A world line completely defines the past, present, and future of a particle for its existence.
We can extend this idea to higher dimensional objects. If we follow a string, instead of a particle, it evolves into a world sheet. A ring would evolve into a world tube. Similarly a membrane would sweep out a world volume. Particles that collide would have their world lines intersect where the point of intersection is the location of the event of their collision.
The important thing to realize is that in this view of spacetime nothing ever really happens, it is already laid out in its future, present, and past. There is no dynamics, it is all geometry.
There is an illusion of motion in spacetime. This illusion occurs when we consider a stack of spatial surfaces where each surface is encoded with a specific proper time parameter. As we allow proper time to evolve from one surface to the next, world lines are traced out in spacetime. This is very similar to the fiber bundle structure of space-time.
Thus we can recover dynamics by slicing spacetime into a stack of three-dimensional surfaces. There is no natural process of slicing spacetime into such surfaces. Since these are surfaces of constant proper time, what we call isochronous surfaces, there can be no natural concept of simultaneity. We will see that this mechanism leads the way to numerical relativity.
Mathematical Interlude: Tangent Vectors, One-Forms, and Tensors in Spacetime
Any set of four quantities
that transform under a change of coordinates in the same way as the spacetime interval
according to (11.1) form what is called a tangent vector in spacetime. Geometrically we think of a tangent vector as an arrow connecting two events, one at its tale and one at its head. The invariant quantity
(11.3)
may be called the squared norm of the tangent vector in spacetime. With a second tangent vector
, we have the scalar product invariant in spacetime
(11.4)
In order to get a convenient way of writing such invariants we introduce the technique of lowering indices in spacetime. Define
(11.5)
Then the expression on the left-hand side of (11.3) may be written
, it is understood that a summation is to be taken over the four values of μ. It is important to note that
is a scalar, the indices contract and we are left with the scalar product invariant of v and v. With the same notation we can write (11.4) as either
or else
. Here we are also left with a scalar, again we say that the indices are contracted. The scalar product invariant <v w> remains.
The four quantities
introduced by (11.5) may also be considered as the components of a one-form in spacetime.
From the two tangent vectors
and
we may form the sixteen quantities
. These sixteen quantities form the components of a tensor of the second rank. This is sometimes called the outer product of the vectors
and
, as distinct from the scalar product (11.3), which is also called the inner product. The outer product, as we have seen, is sometimes called the tensor product and is denoted
.
The tensor formed by the tensor product
is a rather special tensor because there are special relations between its components. But we can add together several tensors constructed in this way to get a general tensor of the second rank,
(11.6)
The important thing about the general tensor is that under a transformation of coordinates its components transform in the same way as the quantities
.
We may lower one of the indices in
by applying the lowering process we used above on each of the terms on the right-hand side of each expression in (11.5). Thus we may form
or
. We may lower both indices to get
.
Exercise 11.2: What happens when we lower the indices of
to get either
,
, or
?
In
we may set ν=μ and get
. We will always sum over the four values of μ when an index appears twice in a term . Thus
is a scalar. It is equal to
.
Exercise 11.3: Show that
is a scalar T.
Enough mathematical formalism for now. Let’s get back to physics.
Lorentz Transformations
Getting back to our example from above, we see that the light beam forms a world line that crosses the
axis at the point
.
Figure 11.5. The worldline of our observer’s light beam heading to the event.
Once the light beam reflects off the event P it returns to the
axis.
Figure 11.6. The worldline of our observer’s reflected light beam heading.
The reflected light beam crosses the
worldline at
.
If we assume that the light beam leaves at time
and is reflected back at
we can fix the coordinates of our event,
We can make this more useful by stating that the initial time occurs at some unit τ,
, the event P itself occurs at some factor of the initial time later,
, and
. We now have
Recall from elementary kinematics,
We can solve this for k in terms of v,
We can see that
will occur a factor, k, later than
, so
and similarly,
If we solve this system of equations we get,
(11.7)
and
(11.8)
These are the famous Lorentz transformations, and they tell us how to look at one coordinate system from another.
Problem 11.1: Explain how the Lorentz transformations combine with Postulate 1 to state that the laws of physics must be Lorentz invariant.
If we make the definition
(11.9)
The Lorentz transformations then become,
(11.10)
From this we see that,
if and only if
Thus we have the ordered pair
for the
coordinate.
Similarly,
if and only if
This gives us the ordered pair
for the spatial
coordinate.
We see that the spacetime diagram now looks something like this,
Figure 11.7. The relationship between the
and the
fames..
This is called a boost in the
plane. There is a change in velocity, but no rotation.
The collection of all boosts and all spatial rotations form the Lorentz group, but we will not go into the details here. The group of transformations satisfying the equation,
(11.11)
is called the Poincaré group. Thus the Poincaré transformations consist of the Lorentz transformations followed by a spacetime translation.
We will get into this in more detail in another lesson.
The Light-Cone Structure
Nothing moves faster than light. This is one of the assumed facts of special relativity. We can use it to examine yet another piece of the structure of a spacetime manifold.
Faster implies speed. Speed is defined as distance traveled—displacement—per unit of time. This idea is fundamentally incompatible with the idea of spacetime.
Exercise 11.4: Explain this incompatibility.
So the idea of moving faster than anything requires a reformulation in order to fit into the concept of spacetime. Suppose that an event occurs in spacetime, at which point a spherical pulse of light is emitted. No particle whose world line passes through this point in spacetime (the event) can ever escape from the spherical pulse. In this way we can say that the particle cannot exceed the speed of light. Were we to label the event p then we can represent the expanding pulse as a cone in spacetime whose vertex is located at p. We can see in such a case that the world line of the particle passing through the event p lies inside the cone.
Figure 11.8. The worldline of a particle passing through the vertex of a light cone is within that light cone.
From this we can see that there is, for every event in M a cone whose vertex is that event. The world line of a particle that passes through the vertex event lies inside the cone. We will call such a cone a future light cone.
Say that we have two future light cones that are close to each other in spacetime. If an event q lies inside the future light cone of the event p, then we say that q is timelike with respect to p. In this case we say that q lies in the future of p.
Figure 11.8. A timelike relation of events in the future.
If p lies in the future of q, then we say that q is timelike related to p and is in its past.
Figure 11.9. A timelike relation of events in the past.
If q lies on the future light cone (or the past light cone) of p we say that q is null related to p in the future (or the past).
Figure 11.10. A null relation of events.
If neither event is within nor on the light cone of the other, then we say that the events are spacelike related to the other.
Figure 11.11 A spacelike relation of events.
In flat spacetime, (another name for special relativity), you can arrange things so that all light cones look like normal cones—they all have the same opening angle and the timelike axes are all parallel. There is no reason to make such an assumption in a generic way. In general relativity, where we will admit curved spacetime, this assumption is no longer justifiable. Indeed, avoiding this assumption is equivalent to assuming such curvature.
Four-Scalars (aka Lorentz Scalars)
Thus far, for the sake of simplicity, we have been considering spacetime as having only one spatial dimension. Of course, we know that in the real world there are three apparent spatial dimensions. This gives us a total of four dimensions, so the spacetime interval is,
Adopting the convention that a Latin index is summed from 1 to 3, we can rewrite this
(11.12)
This is a four-dimensional quantity that is invariant. Such a quantity is called a four-scalar or a Lorentz scalar.
Four-Vectors and Index Gymnastics
A rank-1 tensor is a tangent vector if it has one contravariant index,
, and a one-form if it has one covariant index,
. These are different geometric objects. Components are a vector or a one-form only if they transform as we will see below.
We can consider the four-dimensional Lorentz transformations
We can write the matrix representation of the Lorentz transformation in spacetime
(11.13)
So we write the Lorentz transformations
.
In inertial coordinates the metric is
In a manner similar to
, we can make a transformation to a new set of coordinate axes in spacetime, where each of the
of (11.1) becomes a linear function of
of the new set of axes so that the quadratic form (11.1) become the general quadratic form,
(11.14)
Again we state that the metric tensor is symmetric.
If we have a four-dimensional vector that undergoes a Lorentz transformation, we call it a four-vector. A generic four-vector,
, has the form
(11.15)
Let’s say we have another four-vector,
. We can then can define another four vector as the linear combination
when λ is given a value. Its squared length is
(11.16)
This must be an invariant (four-scalar) for all values of λ.
It then follows that each term is separately an invariant.
The coefficients of λ are then
(11.17)
we can interchange the indices in the second term,
(11.18)
so that we can rewrite (11.17)
(11.19)
From this we see that the second term in (11.17) is an invariant, it is the inner product of
and
.
We can define g as the determinant of
. If this determinant were to vanish (g=0), then the four axes would not provide independent directions in spacetime. We thus, again, assert that the determinant must not vanish. If we have orthogonal axes, the diagonal elements of
become 1, -1, -1, -1 and the off-diagonal elements are all 0. From this we can calculate g=-1.
Exercise 11.5: Perform this calculation.
For oblique axes, similar to those given by a Lorentz transformation, g must still be negative.
Exercise 11.6: Why is this true?
We now define a one-form
such that,
(11.20)
Since g≠0, we can solve these equations for
,
(11.21)
We calculate each
as the cofactor of the corresponding
in its corresponding determinant, divided by the determinant itself.
If we substitute the value of
from (11.21) with that in (11.20)
(11.22)
This equation must be true for any four quantities
we can make the inference,
(11.23)
We can use (11.21) to lower any index in a tensor in spacetime. We can use (11.22) to raise any index in a tensor in spacetime.
For example, a specific four-vector is the 4-velocity
(11.24)
We can also define another 4-vector, the 4-momentum,
(11.25)
We examine the components of the 4-momentum in spacetime
(11.26)
(11.27)
(11.28)
(11.29)
We can use this change of variables,
(11.30)
When we apply this to the components of the 4-momentum we get
(11.31)
(11.32)
(11.33)
(11.34)
We can state the principle of conservation of 4-momentum, given m particles before an interaction and n particles after the interaction
(11.35)
Doing this Stuff in Mathematica
The Constancy of the Speed of Light
We make the substitution 4 π=a,
Here we apply the unit conversion for the Ampere,
Note that Mathematica does not evaluate the units correctly. We can apply the unit conversions for the Ampere by hand and include the units
Then we invert this,
Then we take the square root,
which is approximately the speed of light.
An Inelastic Collision
We will now examine the four-momentum of a simple collision of two particles. Particle 1 of mass
is in its rest frame and is struck by particle 2 of mass
and speed
. We will use the ratio of the speed of light, β=v/c. Say that the two particles become a single particle of mass
now moving at speed
relative to an observer.
We begin by writing out the 4-momentum of particle 1. The four-momentum is written
, where the energy is written E=m γ. Here we assume that we can line up the particles in the
direction
Then we write the 4-momentum of the moving particle
Then we write the 4-momentum of the conglomerate particle
We will state the conservation of 4-momentum according to (11.35), and we will ignore the 0==0 components
We can then solve this equation in terms of
and
Of course we can make this nicer by writing
Particle Decay
What happens if we allow a particle to decay into two other particles? Say the initial particle is in its rest frame. We can then write
If we assume that the decay leads to one particle with positive velocity and one with negative velocity. So, we can write
We again apply the conservation of 4-momentum,
We can write the mass of particle 3 as
, where we write the metric
We can apply a transformation to this,
We can solve this for
or,
We can write the kinetic energy for particle 2 is then just
,
Compton Scattering
This is a traditional problem that resulted in a Noble prize for Arthur Compton. We have a photon of wavelength
whose energy is determined by
. Thus its 4-momentum will be
This photon strikes an electron at rest
The result is the scattering of the photon and the electron. The scattered photon will have wavelength
and scattering angle
.
The scattered electron will have a gamma factor and velocity, with scattering angle
.
We define the metric tensor in the normal way.
We can then find the scalar invariant of
. By conservation of 4-momentum we will have
, or
.
We can solve this for
.
We can find the shift in wavelength.
We can explicitly write the conservation of 4-momentum
The first equation listed is the energy. We can solve this for γ,
We can use ex15 to remove γ.
Then we use ex12 to remove
We attempt to introduce an auxiliary expression
We now reverse the auxiliary expression
Returning to the equations from the conservation of 4-momentum, the other two give us the equations for the scattering angles
We can eliminate γ from the system.
Once again we can use ex12 to eliminate
.
We introduce an auxiliary term to trick Mathematica
We solve this for the dummy tan
,
We now replace the dummy tan
.
Further Reading
Kip S. Thorne, Roger D. Blandford, (2017), Modern Classical Physics, Princeton University Press
Charles W. Misner, Kip S. Thorne, John Archibald Wheeler, (1973), Gravitation, W. H. Freeman and Company.
Edwin F. Taylor, John Archibald Wheeler, (1992), Spacetime Physics, 2nd Edition, W. H. Freeman and Company.
Robert L. Zimmerman, Frederick I Olness, (2002), Mathematica for Physicists, Addison-Wesley Publishing Company Inc.