Lesson 8: More Mechanics and Fields
Introduction
In the electrodynamics lessons you mastered fields, tensors, and the continuum description of matter. In the first steps of continuum mechanics you began to treat bodies as continuous regions that deform and move. Now we pause to strengthen the underlying mechanical and mathematical language that unifies all of it.
This lesson gathers the essential tools of classical mechanics expressed in modern tensor form. You will see how orthogonal transformations change the components of vectors and tensors when you rotate your basis, how systems of particles and rigid bodies are described with equal ease, and how differentiation and integration are written cleanly in index notation. From these foundations the action principle emerges as the single, elegant statement from which the equations of motion follow. Lagrangian mechanics then becomes the practical calculus that turns that principle into differential equations, and the conservation laws appear as the beautiful, automatic consequences of symmetry.
Why does this matter? Because every continuum theory, every field theory, and every discrete mechanical system rests on the same scaffolding. Once you can move freely between bases, write rates of change in tensor language, and derive equations from an action, the later developments of continuum mechanics, continuum electrodynamics, and even relativistic field theories become natural extensions rather than new subjects.
Systems and Rigid Bodies
Matter is made of smaller particles, mostly atoms and ions. The main problem of developing a theory of matter is that if we consider each component particle separately, then we have so many particles in even the smallest object that to calculate the motion of each would tax the most powerful supercomputers. How do we solve this problem?
The decisive simplification is to treat a large collection of particles as a single continuum, or, at the simplest level, as a single particle located at a special point—the center of mass—while still respecting the distances that separate the particles. Consider a perfect cube where we may imagine all of its mass concentrated at the geometric center, yet we keep the edge lengths fixed so that the shape never changes. The same idea works for any object. When the distances between every pair of particles remain constant no matter how the object moves, we call the object a rigid body.
Center of Mass
Take a concrete example—a triangle formed by three particles with masses
,
, and
. The total mass is simply the sum
(8.1)
We can then locate the position of the location of this total mass (the center of mass). In Cartesian coordinates this turns out to be
(8.2)
In more generalized coordinates, this becomes,
(8.3)
This forms a position vector, this vector points to the location of the center of mass.
How can we use this? One way is to consider the total momentum of the triangle (add up the momenta of each particle),
(8.4)
To represent the motion of a rigid body we need a concrete model. The simplest non-trivial example is a system of three particles that remain at fixed distances from one another. Label the particles O, P, and Q. From O we draw the two vectors
and
; their vector product
then supplies a third direction, giving us a complete local basis in which any other point of the rigid body can be located.
Figure 8.1 A simple rigid body.
Although nine coordinates would be required to locate three free particles, the three fixed distances imposed by rigidity reduce the number of independent quantities to six: three that locate the center of mass (or any convenient reference point) and three that describe the orientation of the body.
Now suppose the whole rigid body is moved. The simplest motion is a pure translation: every point is shifted by the same displacement vector
Figure 8.2 Translating our simple rigid body.
Finding the positions of the three points after the displacement is just a matter of adding the vector
to the origin and endpoint of each position vector.
In general we can distinguish three elementary types of displacement of a rigid body:
1. Planar Displacement: All points move parallel to a fixed plane.
2. Translation: Every point experiences the same displacement vector.
3. Rotation: There exists a line (the axis) whose points remain fixed while all other points move on circles about that axis.
Any rigid motion in three-dimensional space can be composed of a translation of the center of mass together with a rotation about an axis through the center of mass. This decomposition, together with the six degrees of freedom already counted, gives a complete and practical description of the kinematics of a rigid body.
With the notions of center of mass, total momentum, and the elementary rigid displacements in hand, you are ready to write the kinetic energy of a rigid body, to introduce the inertia tensor, and to derive the rotational equations of motion that govern everything from a spinning top to a satellite in orbit.
The Rotation of a Rigid Body
You have already seen how a rigid body can translate so that every point experiences the same displacement. The second elementary motion is rotation. To describe it cleanly we begin with pure geometry and then recover the classical Rodrigues formula that tells you exactly where every point moves.
Setting Up the Axis
Choose any two distinct points X and Y that belong to the rigid body and draw the straight segment that joins them. Along that segment we place a unit vector
and we select an origin O that lies somewhere on the same line. The line itself will become the axis of rotation.
Figure 8.3 Establishing a segment between two points in a rigid body.
Next we pick an arbitrary point P that is not on the axis and we introduce the position vector that runs from the origin to that point
(8.5)
The next goal is then to identify a point, say P, on the rigid body and follow this point as it moves. We must then establish a position vector
that we will denote as
.
Figure 8.4 Establishing a point within a rigid body to follow.
We can also extend a line perpendicular to XY from the point N to the point P.
Figure 5.5 Establishing a normal line to our point of interest.
If we rotate the figure about XY through the angle θ we will change the position of our point from P to P'.
Figure 8.6 Rotating our rigid body.
This gives a new position vector
and we will denote this
. Since this is a rigid body the path of the point describes a circle.
Recall that the axis of rotation can be viewed as the vector product of two vectors. If we take the vector product of
and
then we get a new vector in the plane NPP’ that is perpendicular to
. Now, since the length of these vectors never changes (it is a rigid body),
(8.6)
It must also be true that,
(8.7)
We will then have something that looks like this
Figure 8.7 A top-down view of our rotating rigid body.
Constructing the New Position Vector
By simple vector addition we can always write
,
(8.8)
Since the vector
is simply the projection of
on the segment ON,
(8.9)
The vector
is the vector
less the vector
and from Figure 8.7 above it is clear that M is located at the cosine of
,
(8.10)
It is also clear that
is the sine of the vector product
(8.11)
Putting these together gives us,
(8.12)
Exercise 8.1: Derive this result.
The Finite Rotation Tensor
We can rewrite this, where we introduce the notation 1 for the Kronecker delta in component-free notation.
(8.13)
We can then make the following definition
(8.14)
This is called the finite rotation tensor.
Consequently the rotation is simply the linear mapping
(8.15)
We can also rewrite the finite rotation tensor in component notation,
(8.16)
Further Vector Identities
Now, if we take the vector product
(8.17)
Exercise 8.2: Derive this result.
Since a vector product of a vector with itself vanishes, this simplifies to
(8.18)
Once again we recall that we can always introduce an auxiliary term that leaves the equation unchanged,
(8.19)
We can recombine terms,
(8.20)
now we can write,
(8.21)
We now multiply the right-hand side by an auxiliary term, (sin θ/sin θ) to get,
(8.22)
Exercise 8.3: Derive this result.
We can now rewrite and reorder (8.22),
(8.23)
So,
(8.24)
Exercise 8.4: Derive this result.
Rodrigues’ Formula
We can rewrite (8.24) to get the displacement
(8.25)
If we make the assignment,
(8.26)
We can think of this as a vector along the axis of rotation with magnitude
. Our displacement then becomes,
(8.27)
This is called Rodrigues’s formula and was invented in 1840 by the French banker (and amateur mathematician) Olinde Rodrigues. It expresses the finite rotation of every point of a rigid body in terms of a single vector
that points along the axis and whose magnitude is determined by the angle of rotation.
You now possess a complete geometric and algebraic description of pure rotation. Combined with the translational motion examined earlier, it gives you the most general rigid-body displacement in three-dimensional space. The same finite-rotation tensor Φ will reappear when you study the attitude of spacecraft, the rolling of a rigid wheel, and the continuum theory of large deformations.
Euler’s Theorem
Imagine a rigid body that is free to move in any way you like, except that one of its points is permanently fixed in space—think of a gyroscope spinning about a fixed pivot, or a rigid satellite tumbling about its center of mass. You take a snapshot of the body in some initial orientation and then, after an arbitrary motion that keeps the fixed point in place, you take a second snapshot.
At first glance the two orientations may appear completely unrelated; the body could have tumbled through a complicated sequence of twists. Yet a remarkable geometric fact intervenes, that there always exists a single fixed axis, passing through the fixed point, and a single angle such that a pure rotation about that axis carries the body exactly from the initial orientation to the final one.
This is Euler’s theorem. It tells you that the most general displacement of a rigid body with one point held fixed is nothing more than a rotation. All of the elaborate intermediate motion can be replaced, for the purpose of relating the two end configurations, by one clean turn about a uniquely determined axis.
The axis itself is called the Euler axis (or the instantaneous axis of the equivalent rotation), and the angle is the Euler angle of that rotation. Once you know them, the finite-rotation tensor you constructed earlier becomes an explicit, practical tool where you simply insert the Euler axis and the Euler angle into Rodrigues’ formula and the entire displacement is recovered.
Euler’s theorem is the reason a rigid body’s orientation in three-dimensional space can always be described by only three parameters—the direction of an axis and the angle of rotation about it—and it is the geometric foundation upon which the whole subsequent theory of rigid-body kinematics rests.
The Composition of Finite Rotations
Once you can rotate a rigid body from an initial position
to a new position
, a natural question arises, “What happens if you follow that rotation by a second, independent rotation that carries the body onward to a third position
?”
In the language of the finite-rotation tensor the first step is simply
(8.28)
The second rotation acts on the intermediate result
(8.29)
Substituting the expression for
immediately yields the composite mapping
(8.30)
The single tensor
therefore encodes the net effect of the two successive rotations.
A crucial geometric fact must be kept in mind, that the order in which the rotations are performed matters. In general the composition performed in one sequence is not the same as the composition performed in the reverse sequence,
(8.31)
Exercise 8.5: Show this to be true in general.
Finite rotations do not commute. This non-commutativity is the reason that the orientation of a rigid body cannot be treated as an ordinary vector quantity and is the source of many of the subtleties that appear in the kinematics of spinning tops, spacecraft, and robotic manipulators.
Euler’s Theorem Revisited
You have already met Euler’s theorem in its geometric form, where any displacement of a rigid body that leaves one point fixed is equivalent to a pure rotation about some axis through that point. We now recover the same result from the algebraic properties of the finite-rotation tensor itself.
Recall that the Cartesian components of the finite-rotation tensor are given by Rodrigues’ formula
So we can write the matrix in
,
(8.32)
The eigenvalue problem for this matrix is simply
(8.33)
Because Φ represents a pure rotation, its three column vectors are orthonormal, so the sum of the squares of the entries in each column equals unity, and the inner product of any two distinct columns vanishes. In other words Φ is an orthogonal matrix,
(8.34)
(8.35)
(8.36)
(8.37)
where the dummy indices go from 1 to 3.
A fundamental property of any orthogonal matrix is that it preserves the determinant of a set of vectors
(8.38)
Substituting the eigenvalue equation then yields
(8.39)
Provided the determinant does not vanish we conclude at once that
(8.40)
Thus every eigenvalue of a rotation matrix lies on the unit circle. In three dimensions an orthogonal matrix with determinant +1 (a proper rotation) must possess at least one real eigenvalue equal to +1. The corresponding eigenvector is precisely the axis of the equivalent rotation. This is Euler’s theorem expressed in the language of linear algebra.
To see the eigenvalue +1 appear explicitly we form the characteristic (or secular) equation
(8.41)
Writing the matrix out in full gives the homogeneous system
(8.42)
A non-trivial solution exists only when the determinant of the coefficient matrix vanishes,
(8.43)
One root of this cubic is necessarily λ=1, confirming that a real axis of rotation always exists.
Exercise 8.6: Explain why the vanishing of the characteristic determinant is equivalent to the earlier statement that at least one eigenvalue satisfies ∣λ∣=1.
You have therefore recovered Euler’s theorem from the intrinsic algebraic structure of the finite-rotation tensor. The same matrix that rotates every vector of the body also possesses a privileged eigenvector that remains unmoved—the Euler axis itself.
The General Displacement of a Rigid Body
A natural question now arises, “What is the most general displacement of a rigid body?” Intuition suggests that it must be some combination of a pure translation and a pure rotation. We shall make that intuition precise.
Choose an arbitrary point of the rigid body and call it the base point B. Choose a second point and call it the reference point R. We first translate the entire body so that B moves to a new location B’ while R moves to R’.
Figure 8.8 A translation of a rigid body.
We then rotate the body about the new position of the base point so that B' is carried to a final location B''.
Figure 8.9 A rotation fo0llowing the translation of a rigid body.
If we regard the segment
as a position vector, we may write
(8.44)
The pure translation simply adds the same vector to every point, so the intermediate position is
(8.45)
After the subsequent rotation about B’ the final position vector relative to the new reference point is
(8.46)
Because the rotation is performed by the finite-rotation tensor Φ, we have
(8.47)
Combining these relations (and carefully tracking the sense of each segment) yields the compact expression for the most general displacement of a rigid body
(8.48)
Exercise 8.7: Show this to be true.
Exercise 8.8: Show that the rotational part of the displacement is independent of the particular choice of reference point.
In the special case of planar motion the same reasoning shows that every displacement of a rigid body can be realized by a pure rotation about a suitably chosen axis perpendicular to the plane—the instantaneous center of rotation. In three dimensions the combination of an arbitrary translation of a base point with an arbitrary rotation about that point exhausts all possible rigid motions.
You now possess a complete kinematic description of the most general displacement of a rigid body allowing you to translate any convenient base point, then rotate about the new position of that point. This decomposition is the foundation upon which the dynamics of rigid bodies is built.
Screw Displacements
Among all possible rigid-body displacements there is a particularly elegant subclass. Suppose the translation of the base point happens to lie exactly parallel to the axis of the accompanying rotation. The body then advances along the axis while simultaneously turning about it—precisely the motion of a nut on a bolt. Such a combined motion is called a screw displacement.
A remarkable geometric theorem asserts that every rigid displacement in three-dimensional space can be realized as a screw displacement, provided one is free to choose the base point appropriately. In other words, there always exists a unique axis (the screw axis) such that the given displacement consists of a pure translation along that axis together with a pure rotation about the same axis. This is Chasles’s theorem. This is named for the French mathematician and geometer Michael Chasles (1793-1880).
The theorem is the natural three-dimensional completion of the planar result you met earlier, so that every planar rigid motion reduces to a pure rotation about an instantaneous center, every spatial rigid motion reduces to a pure screw about an instantaneous screw axis. Once the screw axis and the associated pitch (the ratio of translational to rotational advance) are known, the entire displacement is completely characterized by two scalar parameters and the direction of a single line.
Chasles’s theorem therefore supplies the most economical description of a general rigid-body displacement and is the kinematic foundation for the theory of screws that appears throughout rigid-body dynamics and mechanism design.
The Velocity of a Rigid Body
Recall the geometry of a pure rotation about a fixed axis (Figure 8.6). A point P of the rigid body is carried along a circular arc to a new location (P’’). The infinitesimal version of that motion will give us the instantaneous velocity of every point of the body.
The finite displacement of the point is simply the difference of the two position vectors
(8.49)
Rodrigues’ formula expresses that displacement in terms of the auxiliary vector
(8.50)
Writing
, and rearranging yields the equivalent form
(8.51)
Exercise 8.9: Derive this result.
To obtain velocities we divide by a small time interval Δ t and pass to the limit Δ t→0
(8.52)
The product rule applied to the first term on the right-hand side gives
(8.53)
where the second equality follows because
itself changes only by the displacement
. Recall that the Rodrigues vector is related to the angle of rotation by
(8.54)
Consequently
(8.55)
and the time derivative becomes
(8.56)
When the angular increment Δθ is vanishingly small we have sec Δ θ≈1, so
(8.57)
The vector
(8.58)
is the angular velocity of the rigid body. Its direction is the instantaneous axis of rotation and its magnitude is the instantaneous rate of rotation about that axis.
Returning to the displacement formula and using (8.58) together with the observation that
as Δ t→0, we find that the second term on the right-hand side of (8.53) vanishes. The velocity of the point whose position vector is
therefore reduces to the simple vector product
(8.59)
This is Poisson’s formula (also known to Euler). One immediate and useful consequence is that angular velocities add as ordinary vectors
(8.60)
Finally, consider a general point of the rigid body whose position relative to a fixed origin is
. The most general rigid displacement consists of a translation of a base point together with a rotation about that base point. Differentiating with respect to time therefore yields the velocity of an arbitrary point P
(8.61)
Since
is the position of P relative to the base point. This is the fundamental velocity formula for a rigid body.
You now possess the instantaneous kinematic description that complements the finite-displacement theory developed earlier. The angular-velocity vector
and Poisson’s formula will be the starting point for the kinetic energy, the angular momentum, and the rotational equations of motion of any rigid body.
Momentum of a System
You already know the momentum of a collection of discrete particles, where the i-th particle has mass
and velocity ![]()
, the total linear momentum is simply the sum
(8.62)
When the number of particles becomes indefinitely large we pass to a continuous distribution of matter. The sum is replaced by an integral over the volume occupied by the body,
(8.63)
where the integral is understood to run over every infinitesimal volume element that makes up the rigid body. In other words, you add up the momentum contributions of all the “little boxes” that together constitute the continuous mass distribution.
(The quantity m that appears under the integral sign is the local mass density; in more conventional notation one writes ρ dV for the mass of each volume element.)
This continuum expression for the total momentum is the starting point for the dynamics of both rigid bodies and deformable continua. Once you can evaluate the integral—by exploiting the rigidity constraint or by introducing the center-of-mass velocity—you recover the familiar result that the total momentum of a rigid body is exactly the same as that of a single particle whose mass is the total mass and whose velocity is the velocity of the center of mass.
Angular Momentum of a System
You already possess an expression for the total linear momentum of a rigid body. That momentum, however, knows nothing about the body’s rotation; it only tracks the motion of the center of mass. To capture the rotational contribution we must form the moment of the momentum—the vector product of a radius vector with the linear-momentum density. The resulting quantity is the angular momentum. We denote the angular momentum of the body by the vector
. For a continuous mass distribution it is defined by the integral
(8.64)
or, equivalently, by the continuum limit of the discrete sum
(8.65)
When the velocity of each mass element is decomposed into the translational velocity of a base point plus the rotational contribution
given by Poisson’s formula, the angular momentum expands into the explicit form
(8.66)
Exercise 8.10: Derive this result.
(The same identity holds under the integral sign for a continuous body; the factor m is then understood as the local mass density.)
This is the angular momentum of the rigid body relative to the chosen origin. The first term involves the motion of the reference point itself; the remaining terms are quadratic in the position and linear in the angular velocity.
The Idea of a Continuous System
Until now we have spoken of systems composed of a finite number of individual particles, though we have alluded to continuous bodies. We now make that precise. There is another, complementary way of thinking about matter. Imagine that no matter how close two particles may be, you can always insert still another particle between them. When a system admits this possibility at every location, we say the system is continuous: it has no gaps, and the number of particles it contains is infinite.
A rigid body is a special continuous system where the mutual distances between all material points remain forever fixed; the distribution of mass never changes as the body moves. By contrast, a system that contains only a finite number of particles is called discrete.
The continuous idealization lets us replace discrete sums by integrals over volume and is the conceptual bridge that carries us from the mechanics of particles to the continuum theories of rigid bodies, deformable solids, and fluids.
The Kinetic Energy of a Continuous System
When a system is continuous, each “particle” is only an infinitesimal mass element dm. The discrete sum that defined the kinetic energy of a finite collection of particles is therefore replaced by an integral over the mass distribution of the body
(8.67)
It is almost always convenient to refer the motion to the center of mass. Let
be the position of the center of mass and
its velocity. The position of a generic mass element may then be decomposed as
(8.68)
and thus, for the velocity of the center of mass
,
(8.69)
Substitution into the kinetic-energy integral produces three terms
(8.70)
Exercise 8.11: Derive this result.
Because
is independent of the integration variables one may factor it out, rewriting the middle integral
(8.71)
where
is the total linear momentum and M is the total mass. By the very definition of the center of mass,
, so the middle term is zero. The kinetic energy therefore reduces to the elegant sum of two contributions
(8.72)
In words: the kinetic energy of a continuous body is equal to the kinetic energy of a single point mass M moving with the velocity of the center of mass, plus the kinetic energy of the motion relative to the center of mass. This decomposition was first obtained in 1751 and is known as König’s theorem (after the German mathematician Johann Samuel König).
For a rigid body, the relative-velocity field is completely determined by the angular velocity
, and the first integral becomes the rotational kinetic energy. The same separation remains valid for deformable bodies; only the evaluation of the relative-motion integral changes.
Another Look at Angular Momentum
You have already met the angular momentum of a system of particles. Now that we are working with continuous bodies it is natural to revisit the same idea in integral form. Relative to an arbitrary point P the angular momentum of a continuous mass distribution is simply the first moment of the linear-momentum density
(8.73)
For a rigid body the velocity of every mass element relative to P is given by Poisson’s formula
(when P is fixed in the body or is the center of mass). Substituting that expression and using the vector identity
converts the integral into
(8.74)
This last integrand is interesting, we can write it as the tensor
(8.75)
This is the inertia tensor about the point P. Consequently the angular momentum takes the compact and fundamental form
(8.76)
The inertia tensor is a symmetric, positive-definite linear map that encodes how the mass of the body is distributed relative to the chosen point. Once
is known, every rotational property of the rigid body—angular momentum, rotational kinetic energy, and the equations of motion—follows at once by ordinary matrix algebra.
This relation
is the rotational analogue of the linear relation
and is the starting point for the entire dynamics of rigid bodies.
Exercise 8.12: Write the component form of inertia tensor.
Kinetic Energy and the Inertia Tensor
The portion of the kinetic energy that arises from rotation about a fixed point P (or about the center of mass) is precisely the relative-motion integral that appeared in König’s theorem:
(8.77)
Because the body is rigid, the relative velocity of every mass element is given by Poisson’s formula. With the position measured from the center of mass we have
(8.78)
Exercise 8.13: Derive this result.
Substitution into the rotational kinetic energy then yields
(8.79)
Exercise 8.14: Explain this result.
The remaining integral is exactly the angular momentum relative to the center of mass, so
(8.80)
We can rewrite this in terms of the inertia tensor
(8.81)
Exercise 8.15: Explain these steps.
Returning to the full kinetic energy of the rigid body and restoring the translational term given by König’s theorem, we obtain the complete decomposition
(8.82)
The first term is the energy stored in the rotational motion; the second is the energy of the center-of-mass translation. Together they account for the entire kinetic energy of any rigid body.
Exercise 8.16: Write (8.85) in component form.
Components of the Inertia Tensor
If we adopt a three dimensional coordinate system over
, we can express the inertia tensor as a matrix of components,
(8.83)
We call this a matrix representation of the inertia tensor. The diagonal components will have the form,
(8.84)
These are called the moments of inertia. The off-diagonal components have the form
(8.85)
These components are called the products of inertia.
They measure the imbalance of the mass distribution with respect to the coordinate planes. Because the inertia tensor is symmetric,
, so only six numbers are required to specify it completely in any given frame.
These nine (or six independent) components are the concrete numbers that appear in every practical calculation of angular momentum, rotational kinetic energy, and the torque equation for a rigid body.
Moment of Inertia
A continuous body is characterized not only by its total mass but by the way that mass is distributed in space. The simplest local measure of the distribution is the mass density
(8.86)
the mass per unit volume. (You have already met densities of charge, polarization, and energy in the electrodynamics lessons; mass density plays the analogous role for continuum mechanics.)
When the density is known, the scalar moment of inertia of the body about a chosen axis is defined by the integral
(8.87)
where
is the perpendicular distance from the axis to the mass element dm. This single number tells you how resistant the body is to angular acceleration about that particular axis.
In the tensor language developed earlier the same information is contained in the full inertia tensor
evaluated at a convenient origin P. If a new orthonormal frame
is introduced, the components of the tensor transform in the usual way for a second-rank tensor. Defining the orthogonal transformation matrix by
(8.88)
one has
(8.89)
Once the inertia tensor is known in any frame, the moment of inertia about an arbitrary axis whose direction is the unit vector
is recovered by the simple quadratic form
(8.90)
Thus the entire family of scalar moments of inertia—about every possible axis through P—is encoded in the six independent components of the inertia tensor. To compute any one of them you evaluate (or look up) the tensor and then form the indicated contraction with the unit vector of the desired axis.
The Parallel Axis Theorem (Steiner’s Theorem)
It is often easiest to compute the inertia tensor about the center of mass and then shift the result to some other convenient point P. To do so we begin with the elementary decomposition of the position vector,
Then,
(8.91)
Exercise 8.17: Explain these steps.
Because
is the center of mass, the first moment of the relative position vanishes identically
(8.92)
All of the cross terms therefore disappear, leaving the clean relation
(8.93)
This is called the parallel-axis theorem. It is also called Steiner’s theorem, or the tennis racket theorem. In words: the inertia tensor about an arbitrary point P is equal to the inertia tensor about the center of mass plus a simple correction that depends only on the total mass and on the vector that joins the two points.
If one is interested solely in the scalar moment of inertia about an axis whose direction is the unit vector
, the same theorem reduces to
(8.94)
The quantity
is precisely the square of the perpendicular distance between the axis through P and the parallel axis through the center of mass. Thus the theorem recovers the elementary statement familiar from introductory mechanics: the moment of inertia about any axis is the moment about a parallel axis through the center of mass plus
.
The parallel-axis theorem is one of the most frequently used practical tools in rigid-body dynamics; it lets you move the inertia tensor from the center of mass (where it is usually simplest) to any other point required by the problem at hand.
The Resultant Force
A rigid body is ordinarily acted upon by a distributed system of forces—body forces that act throughout its volume and contact forces that act on its surface. The net effect of all these forces on the translational motion of the body is completely captured by a single vector, the resultant force,
(8.95)
In other words, one simply adds (integrates) every infinitesimal force contribution that acts on every mass element of the body. Once the resultant is known, the center-of-mass motion obeys Newton’s second law in its most familiar form,
(8.96)
exactly as though the entire mass were concentrated at the center of mass and acted upon by the single force
.
All of the detailed information about how the forces are distributed is needed only when one turns to the rotational motion; for the translation of the center of mass the resultant alone is sufficient.
Torque
Just as the resultant force is the first moment of the force distribution with respect to mass, the torque (or moment of force) is the first moment of that same distribution with respect to a chosen origin. If
is the position vector of a mass element relative to the origin and
is the infinitesimal force acting on it, the total torque is
(8.97)
When the origin is fixed in an inertial frame (or is the center of mass), the fundamental relation between torque and angular momentum becomes simply
(8.98)
In other words, the net torque acting on the body is equal to the time rate of change of its angular momentum. This is the rotational analogue of Newton’s second law and is the starting point for every dynamical calculation that involves the rotational motion of a rigid body.
Euler’s Equations
The two fundamental balance laws you have already met—the resultant-force equation for the center of mass and the torque equation
—are collectively known as Euler’s equations of motion. A natural question is, “About which points may the torque equation be written in that simple form?”
Consider an arbitrary point B that is fixed in the rigid body. The position vector relative to a fixed origin may be decomposed as
The total torque about the fixed origin then splits at once into
(8.99)
Whenever the second term vanishes (in particular when B is the center of mass, or when B is a fixed point of the motion), the torque equation reduces to the pure statement
. Thus any point that maintains a fixed position relative to the rigid body may be used, provided the appropriate extra term is retained when that point is not the center of mass.
A practical difficulty arises because the inertia tensor, when referred to axes fixed in space, changes as the body rotates. The difficulty is removed by introducing a coordinate frame that is itself fixed in the rotating body. In that body-fixed frame the components of I are constant. If an asterisk denotes the time derivative computed in the body frame, the ordinary inertial-frame derivative of the angular momentum is related to the body-frame derivative by the transport theorem
(8.100)
Because
and I is now constant, one has
. Substitution therefore yields Euler’s equation in its classic form
(8.101)
When principal axes of inertia are chosen as the body frame, the matrix I becomes diagonal and the vector equation separates into the three scalar Euler equations that govern the rotational dynamics of every rigid body.
The Orthogonal Transformations of Bases
Before we apply the full machinery of calculus to tensors, it is worth pausing to examine a particularly important class of transformations—those that take one Cartesian basis into another. You have already met the essential ideas while discussing rigid bodies; here we make them precise.
Imagine two Cartesian coordinate systems that describe the same Euclidean space. One system uses the unprimed coordinates
with orthonormal basis vectors
; the other uses the primed coordinates
with orthonormal basis vectors
. Because both sets span the same three-dimensional space, each basis vector of one set can be expanded in terms of the other
(8.102)
The coefficients
and
are simply the elements of a pair of transformation matrices; they are not the components of a tensor. One matrix is the transpose of the other, and the two matrices are inverses of each other
(8.103)
Because the bases are Cartesian they are orthonormal, so
(8.104)
(The last step uses the orthonormality of the primed basis.) The final equality states that the transpose of the matrix
is also its inverse. In other words, the transformation matrices are orthogonal.
Exercise 8.18: Explain each step to get (8.104).
An orthogonal matrix represents either a pure rotation or a rotation combined with a reflection. In the great majority of physical applications we deal with proper rotations (determinant +1), so the matrices Λ are simply rotation matrices. They allow you to change from one Cartesian frame to another while preserving lengths, angles, and the orientation of the space—precisely the transformations that leave the underlying Euclidean geometry unchanged.
With this clear understanding of how bases transform, you are ready to differentiate tensors, to integrate over volumes, and to build the variational principles that govern the mechanics of particles, rigid bodies, and continuous media.
Transforming a Vector
You have already seen how orthonormal bases are related by orthogonal transformation matrices. It is now worth pausing to synthesize what those transformations do to the components of a vector, while leaving the geometric vector itself untouched.
A tangent vector is a geometric object that exists independently of any particular coordinate frame. Its components, however, are always expressed relative to a chosen basis. Suppose the same vector
is written in two different Cartesian bases
(8.105)
The vector
itself has not changed; only the numerical labels that describe it with respect to a new set of axes have changed. In other words, the geometric object is basis-independent, while its component representation transforms in a definite, linear way under a change of basis.
This elementary observation is the prototype for the transformation law of every tensor: the underlying geometric (or physical) quantity remains the same; only the array of numbers that represents it in a particular frame is altered according to a precise rule involving the matrices Λ.
Transforming a Tensor
What you have just seen for a vector extends, without surprise, to tensors of any rank. A tensor is a geometric (or physical) object that exists independently of the coordinate basis in which it is expressed. Only its components change when the basis changes.
For a third-rank tensor, for example, the components in the primed Cartesian frame are related to those in the unprimed frame by the orthogonal transformation matrices Λ
(8.106)
Each index of the tensor transforms exactly as the components of a vector transform. The same pattern holds for a tensor of arbitrary rank, where one factor of Λ appears for every index.
Exercise 8.19: Verify this result.
Once this rule is understood, the transformation properties of the inertia tensor, the stress tensor, the finite-rotation tensor, and every other tensor that appears in continuum mechanics become immediate applications of a single, uniform principle.
Transforming a Point
It is tempting to treat the coordinates
of a point as though they were the components of a vector, but the two notions are not identical.
Exercise 8.20: What is the difference between point a vector?
When two Cartesian frames share a common origin, the numerical coordinates of any given point are related by precisely the same orthogonal matrices that transform vector components
(8.107)
The geometric point itself remains fixed in space; only the labels that locate it relative to each basis change. The transformation law looks identical to that of a vector because, once an origin has been chosen, the position of the point relative to that origin is a vector. If the origins of the two frames do not coincide, an additional constant translation appears and the simple homogeneous rule above no longer holds.
This distinction—between a true geometric vector and the coordinates of a point—will become important as soon as we begin to differentiate tensor fields and to change variables in volume integrals.
The Action Principle
The central aim of classical mechanics is to predict the future state of a system once its present state and the laws that govern its evolution are known. When those predictions can be made with no uncertainty other than the practical limits of computation, we say that we understand the system.
What information is required to specify a state? In Newtonian mechanics the answer is the set of all positions together with the set of all velocities. It may seem puzzling that accelerations are not also needed. The reason is simple, Newton’s equations of motion are second-order differential equations for the positions. A unique solution of a second-order equation is fixed by exactly two pieces of data at an initial instant—the position and the velocity of each particle.
If we are assuming that given some state and the equations of motion, then we can predict the future state of the system. The equations of motion as determined by Newton are second-order differential equations that define the positions of particles given some future time. However, such equations require two initial conditions to be uniquely solved. Those are the positions and velocities.
Those equations may be written, for a single particle in a potential, as
(8.108)
From the same ingredients—kinetic energy and potential energy—one can construct a single scalar functional of the entire path,
(8.109)
This functional is called the action. Every conceivable path that the particle might follow between the fixed end-points
and
yields its own numerical value of S.
The action principle asserts that the path actually taken by the particle is the one that renders the action stationary under small variations of the path, denoted by δ
(8.110)
Exercise 8.21: Can you derive this result?
This means that we must choose
so that it vanishes at the end-points
and
. If we write
, then we can integrate by parts,
(8.111)
Exercise 8.22: Can you derive this result?
Because the variation
is otherwise arbitrary, the integrand itself must vanish, and Newton’s second law is recovered. Thus, every true dynamical trajectory is a stationary path of the action. This is the principle of stationary action (often called the principle of least action). It elevates a single scalar functional into the fundamental law from which all of the equations of motion follow, and it remains the most economical and powerful starting point for the whole of classical mechanics.
Differentiation in Tensor Notation
Transforming a Point
Let us return for a moment to the geometric idea of a fiber bundle. Suppose we are given a base space B. At every point of B we attach a vector space (or a more general algebraic space) whose elements are tensors of a definite type—say the space
of mixed tensors with one contravariant and one covariant index. The total collection of all these attached spaces is called a tensor bundle. It is the natural geometric home for tensor fields, where a single continuous object that assigns to each point of the base space a tensor of the chosen type.
Figure 8.10 A tensor field on a tangent vector, forming a tensor bundle.
The individual tensors that sit in the fibers, once they are assigned smoothly from point to point, constitute a tensor field on the base space. In the language of the xAct suite the same geometric structure is precisely what is created when you issue the command DefManifold, where you are declaring the base canvas together with the tensor bundles that will live over it. That is why xAct speaks of tensor bundles whenever you define the arena in which your subsequent tensor calculus will take place.
With this picture in mind—the base space as the stage and the fibers as the linear spaces of tensors attached to each point—you are ready to differentiate tensor fields, to form covariant derivatives, and to write the coordinate-free equations of continuum mechanics and field theory.
Continuity
A tensor field is more than a collection of algebraic objects attached to each point of the base space; those objects must also vary from point to point in a controlled way. When a tensor is expressed in a coordinate basis and every one of its component functions is of class
(that is, r times continuously differentiable) with respect to the coordinates, we say that the tensor field itself is of class
.
In particular, a tensor field of class
is merely continuous, while a tensor field of class
is smooth. Most of the fields that appear in continuum mechanics and classical field theory are assumed to be at least
, so that their first derivatives exist and are continuous; many theoretical developments quietly strengthen the assumption to
in order to avoid technical distractions.
This simple regularity requirement guarantees that the algebraic operations of tensor calculus—contraction, symmetrization, covariant differentiation—produce new tensor fields that remain inside the same smoothness class, and it is the starting point for every existence and uniqueness theorem that follows.
A First Application of Differentiation
Consider a vector field
(with components
) defined on ordinary Euclidean space
. Suppose we are given two Cartesian coordinate systems whose coordinates are related by a smooth, invertible transformation
(8.112)
Each unprimed coordinate is therefore a function of all the primed coordinates, and vice versa. Differentiating the transformation immediately supplies the linear map that takes the components of the vector from one frame to the other:
(8.113)
The same reasoning extends, index by index, to a tensor of arbitrary rank. If T has components
in the unprimed frame, its components in the primed frame are
(8.114)
Observe the characteristic reversal of the partial derivatives that act on the covariant indices where each upper index transforms with a factor ∂ x'/∂x, while each lower index transforms with a factor ∂x/∂ x'. This is precisely the transformation law that guarantees the tensorial character of T.
General Orthogonal Coordinates
Cartesian coordinates
are the simplest possible system, but many problems become far more transparent when expressed in a different orthogonal curvilinear system
}. Let
be the underlying Cartesian orthonormal basis and let
be the orthonormal basis associated with the new coordinates. The two bases are related by the scale factors
(there is no sum on i) according to
(8.115)
Then
(8.116)
Exercise 8.23: Write cylindrical coordinates according to (8.119)
The Directional Derivative Operator
Imagine a smooth curve that starts at a fixed origin O and passes through a variable point P(t). We may parametrize the curve so that the displacement from the origin is simply proportional to a chosen vector
(8.117)
The directional derivative operator along
is then defined by ordinary differentiation with respect to the parameter t, evaluated at the initial instant
(8.118)
When this operator acts on a scalar field φ one recovers the familiar directional derivative
(8.119)
The same definition extends at once to an arbitrary tensor field T. In coordinate-free language the directional derivative is the limit
(8.120)
Note that the operation raises the rank of the tensor by one, since it produces a new tensor that possesses an extra linear slot ready to accept the vector
.
Because the directional derivative is essentially an ordinary derivative along a line, it obeys the usual product rules. For the tensor product of two fields one has
(8.121)
while for a scalar multiple the Leibniz rule reads
(8.122)
These two identities, together with linearity, are the only algebraic properties needed to differentiate every tensor expression that appears in continuum mechanics and classical field theory. The directional derivative is therefore the concrete differential operator from which the gradient, the covariant derivative, and all subsequent calculus of tensors will be built.
The Nabla Operator
In a Cartesian coordinate system the fundamental differential operator of vector calculus is simply the collection of partial derivatives
(8.123)
When the same operator is written with respect to a general (possibly non-orthogonal) coordinate basis
(8.124)
where the
are now the dual basis vectors associated with the chosen coordinates.
In an orthogonal curvilinear system
with scale factors
the expression acquires the familiar weighting by the reciprocal scale factors
(8.126)
Exercise 8.24: Write the Nabla operator for cylindrical coordinates.
With these coordinate formulae in hand you can compute gradients, divergences and curls in any orthogonal system by purely algebraic manipulation of the scale factors and the unit vectors
The Gradient
When the directional-derivative operator is applied to a scalar field φ, the result is called the gradient of φ. In any coordinate system one has
(8.126)
The object ∇ φ (often written d φ when one wishes to emphasize its character as a differential form) is a covector, or 1-form, where it is a linear machine that accepts a vector
and returns the rate of change of φ in the direction of
. Geometrically it counts how many level-surface elements of φ are pierced by
. We reserve the special notation d φ for the gradient of a scalar; for all higher-rank fields we continue to write ∇. In a general orthogonal system with scale factors
the gradient takes the form
(8.127)
The same operator can be applied to fields of higher rank. For a vector field
one defines the tensor gradient (or covariant derivative)
(8.128)
or, in components
(8.129)
(The precise arrangement of unit vectors may be written in any consistent order; the essential point is that each partial derivative is accompanied by the appropriate scale-factor corrections that arise from differentiating the coordinate basis.)
Exercise 8.25: Derive this result.
Exercise 8.26: Write the tensor gradient of a vector in cylindrical coordinates.
Thus the gradient operation, first introduced for scalars, extends systematically to every tensor field and supplies the raw material from which divergence, curl, and the covariant derivative will be built.
Christoffel Symbols
In any coordinate system that is not Cartesian the basis vectors themselves change from point to point. The quantities that measure that change are the Christoffel symbols, named for the German mathematician Elwin Bruno Christoffel (1829–1900).
Given a metric tensor
and its associated coordinate basis, the Christoffel symbols of the first kind are defined by
(8.130)
(They are symmetric in the last two indices.) Raising the first index with the inverse metric produces the Christoffel symbols of the second kind,
(8.131)
These symbols are not the components of a tensor; they compensate for the variation of the basis so that the covariant derivative of a tensor transforms correctly. In Cartesian coordinates they vanish identically. In every curvilinear system they supply the extra terms that appear when one differentiates the unit vectors
or when one converts ordinary partial derivatives into covariant derivatives.
Once the Christoffel symbols of a given coordinate system are known, the covariant derivative, the geodesic equation, and the curvature tensor can all be written by purely algebraic combinations of the
and their derivatives.
The Covariant Derivative
Ordinary partial differentiation is not sufficient once the basis vectors themselves vary from point to point. The covariant derivative is the unique extension of the directional derivative that accounts for that variation and still produces a tensor. In a Cartesian frame the extra terms vanish and the covariant derivative reduces to the ordinary partial derivative. We denote the covariant derivative by a semicolon and the ordinary partial derivative by a comma:
versus
.
Begin with a contravariant vector field
. Differentiating with respect to the coordinate
gives
(8.132)
Exercise 8.27: Explain each step.
The derivatives of the basis vectors are linear combinations of the basis vectors themselves; the coefficients are precisely the Christoffel symbols of the second kind,
For a covariant vector (1-form)
the same reasoning produces a minus sign in front of the Christoffel term
(8.133)
Exercise 8.28: Explain each step.
Problem 8.1: Explore the notion of a contravariant derivative, where
or
.
Problem 8.2: Extend these ideas to rank two tensors.
Problem 8.3: What happens when we take the gradient of a vector followed by a contraction. This is the divergence. Note the Voss-Weyl formula. Extend this to a general rank two tensor.
Problem 8.4: Explore the curl of a tangent vector, 1-form, and tensor.
Problem 8.5: Explore the Laplacian of a scalar field, a tangent vector, and a 1-form.
The covariant derivative is the fundamental differential operator of tensor analysis on a manifold. Once it is available, the notions of parallel transport, geodesic, curvature, and divergence all follow by purely algebraic combinations of the semicolon derivative.
Lagrangian Mechanics
Look back at the integrand that appeared inside the action integral (8.111). That combination of kinetic energy minus potential energy is called the Lagrangian
(8.134)
(The precise historical and variational reasons for this particular combination lie outside the scope of the present discussion.)
Problem 8.6: How do we get this form of the Lagrangian?
For a system of n particles one introduces 3 n generalized coordinates
(α=1,…,3n) together with their associated generalized velocities
. The Lagrangian then becomes a function of these 6 n variables and of time,
(8.135)
Exercise 8.29: Determine the Lagrangian of a free particle.
Exercise 8.30: Determine the Lagrangian of a falling body.
Exercise 8.31: Determine the Lagrangian of a simple harmonic oscillator.
Once the Lagrangian is known, the stationarity of the action yields the Euler–Lagrange equations (also called simply Lagrange’s equations)
(8.136)
Exercise 8.32: Derive this result.
The quantity
is called the conjugate momentum, where we write
(8.137)
The quantity
is called the generalized force. When these identifications are inserted into the Euler–Lagrange equations one recovers Newton’s second law, demonstrating the complete equivalence of the two formalisms for unconstrained systems.
So, how do we use these? We write out the Lagrangian of our system. Then we take the derivative with respect to the generalized coordinates. In another step we take the derivative of the Lagrangian with respect to the velocities, then we take the derivative of that with respect to time. When we equate these two expressions we have a set of equations of motion.
Exercise 5.20: Determine the Lagrange equations of a free particle.
Exercise 5.20: Determine the Lagrange equations of a falling body.
Exercise 5.20: Determine the Lagrange equations of a simple harmonic oscillator.
The practical procedure is therefore straightforward:
Construct the Lagrangian T−V in any convenient set of generalized coordinates.
Compute the partial derivatives
and
.
Form the total time derivative of the latter and set it equal to the former.
Why prefer the Lagrangian formulation? There is no deeper theoretical superiority—the two sets of equations are equivalent—but there are decisive practical advantages. All of the physics is encoded in a single scalar function. Changing coordinates requires only the substitution of the new expressions for (T) and (V); one never has to transform a whole system of vector equations. Constraints are handled by the simple device of choosing generalized coordinates that already incorporate them, thereby reducing the number of degrees of freedom at the outset. Moreover, every continuous symmetry of the Lagrangian immediately supplies a conserved quantity (Noether’s theorem); translational invariance, for example, yields conservation of linear momentum.
The principal arena in which the Newtonian approach retains an advantage is the presence of non-conservative forces such as friction. Those forces are most naturally written as explicit terms on the right-hand side of Newton’s laws; they will reappear when we study the continuum mechanics of viscous fluids.
With the Lagrangian in hand you possess a compact, coordinate-flexible, and symmetry-transparent route from the geometry of a system to its differential equations of motion.
Volumes
In the Euclidean plane
the “volume” of a region is simply its area. When the underlying basis is orthonormal, the signed area of the parallelogram spanned by two vectors
and
is given by the contraction with the Levi-Civita tensor,
(8.138)
The same construction extends at once to three dimensions. The signed volume of the parallelepiped spanned by
,
, and
is
(8.139)
The Surface and Volume Elements
In Cartesian coordinates the elementary area and volume elements are simply the products of the coordinate differentials,
(8.140)
(The corresponding expressions in orthogonal curvilinear coordinates acquire the appropriate products of scale factors.)
The Surface Area Vector
Given a parallelogram that lies on a surface and is spanned by the two vectors
and
, one defines the surface-area vector
(8.141)
Its magnitude equals the ordinary area of the parallelogram, while its direction is normal to the surface (the sense being fixed by the right-hand rule). This oriented area element is the natural quantity that appears in flux integrals and in the integral theorems of vector calculus.
With these elementary geometric notions in hand you are ready to change variables in multiple integrals, to transform volume forms under coordinate transformations, and to write the integral balance laws of continuum mechanics in a fully coordinate-free language.
Conservation Laws
Every conservation law has the same fundamental structure: the total amount of a conserved quantity contained inside an arbitrary volume can change only by flowing across the boundary of that volume. Consequently the rate of change of the integrated quantity, plus the net outward flux through the closed surface, must vanish.
Consider, as a prototype, the conservation of particle number. Let n be the number density of particles and let
be the particle flux vector. Then, for any fixed volume V with boundary surface S,
(8.142)
Because the volume is arbitrary, the integrand itself must vanish, and one obtains the local continuity equation
(8.143)
Exactly the same reasoning produces the local conservation laws for mass, momentum, energy, charge, and every other quantity that is neither created nor destroyed. In each case one writes
(8.144)
(or, when sources are present, equal to the appropriate source density). The integral statement guarantees global conservation; the differential statement is the form that enters the field equations of continuum mechanics and classical field theory.
Doing This Stuff in Mathematica
Say that we have a mass, m, on an inclined plane. The inclined plane has mass, M, and is allowed to move in the x direction, the plane has height, h, and the length of the incline is, ℓ, with an angle of incline, θ.
Figure 8.11. A mass on a moving inclined plane.
The first task is to write the Lagrangian. To do this, we have to find the kinetic and potential energies.
To make things simpler we can change variables to a coordinate based on the incline, we will symbolize this as s[t],
So we can write the Lagrangian.
In[13]:=
L=T-V /.cov //Simplify
We will then load the add-on package VariationalMethods.
We can reduce these to a different form.
In[16]:=
TableForm[eom=Solve[eq1, {xb''[t],s''[t]}] //Flatten]
We can find the constants of motion using the FirstIntegrals command.
What are these values? Let’s examine the first one, the First Integral for xb. If this is based on a translational invariance, then it is the momentum. We can write the momentum.
Check and make sure this is conserved.
The second of the constants is a time translation symmetry, thus the energy. We can write the Hamiltonian
So these are the constants of the motion.
The solution of the differential equations is found.
The square of the time to reach the bottom of the incline.
We will present with the tensor product command
We now load xAct.
Then we define the underlying space,
.
Notice that this command establishes the tensor bundle called TangentE3. This is the space that all of the tensors will live in.
We can define the underlying metric of our space, where we are defining the determinant of the metric, the label of the metric, the programming symbol for the covariant derivative and its labels, and the symbol for the metric.
There is lots of stuff that we will not explain here. All of it will be explained by the end of the course.
We now define a tensor
How do we differentiate tensor in Mathematica using xAct?
We can specify a directional derivative. We use the Dir[α] command to indicate a directional index where we can replace α with a vector or 1-form. We establish the vector
We can use the comma symbol as a normal coordinate derivative (partial derivative)
| T |
|
We can define the Christoffel symbol for the derivatives,
| Γ[∇] |
|
Say we have
| T |
|
We can change this to partial derivatives
This is not readable by normal people. As a rule we will from this point on make the following statement
Now we execute our command
There is another system in xAct that is not abstract indices, but allows for calculations in specific coordinate bases. We begin by loading xCoba
We will use cylindrical coordinates. We begin by defining the coordinate system, we label it eb1 for Euclidean Basis #1, the space it is sitting in
, and the coordinate functions
We can define the coordinates for tensor slots, using CIndexForm[label number, coordinate basis label already defined]
We can define the metric in our coordinate system
Here we display the metric in terms of cylindrical coordinates,
We can see how the symmetries of the metric work,
The Christoffel symbols can also be written
| Γ[∇,D] |
|
Further Reading
Leonard Susskind, George Hrabovsky, (2011), The Theoretical Minimum. Basic Books.
Jorge V. José, Eugene J. Saletan, (1998), Classical Dynamics A Contemporary Approach. Cambridge University Press.
Ian D. Lawrie, (2012), A Unified Grand Tour of Theoretical Physics. Taylor & Francis Group, LLC under the CRC imprint in its 3rd Edition.
Edward A. Desloge, (1982), Classical Mechanics, vol 1. John Wiley and Sons
John Michael Finn, (2008), Classical Mechanics, Infinity Science Press LLC.
E. A. Fox (1967), Mechanics, Harper and Row.
Kip S. Thorne, Roger D. Blandford, (2017), Modern Classical Physics, Princeton University Press
Charles W. Misner, Kip S. Thorne, John Archibald Wheeler, (1973), Gravitation, W. H. Freeman and Company.
Robert C. Wrede, (1963), Introduction to Vector and Tensor Analysis, John Wiley and Sons, reprinted in 1972 by Dover Publications