Lesson 7 Applications of Electrodynamics II

Introduction

In Lesson 6 you watched electromagnetic waves enter conducting media, feel the damping of the skin effect, and respond to the ordered moments of magnetic materials. You saw how free charges and microscopic currents reshape propagation, how tensors capture anisotropy, and how collective order appears below the Curie temperature. Those insights remain with you; they form the foundation on which the present lesson is built.

Now the focus widens from the interior of materials to the larger stage on which fields act such as the guided channels that carry energy, the accelerating charges that radiate, the distant patterns that antennas paint across the sky, and the subtle momentum and angular momentum that light itself carries. You will move from the macroscopic Maxwell equations—already prepared by the constitutive relations of Lesson 6—to the concrete geometries of waveguides and cavities, then outward to the radiation fields of accelerating charges, multipoles, and scatterers. Along the way you will meet the practical design of antennas and radiation patterns, the mechanical consequences of electromagnetic momentum, and the optical phenomena that ultimately rest on the same equations.

Why does this matter physically? Because every radio signal, every laser pulse, every photon that reaches your eye, and every force exerted by a focused beam is an application of the same principles. The mathematics you are about to explore is nature’s way of converting the motion of charges into traveling fields that can transport energy, information, and even linear and angular momentum across empty space. It is the bridge between the microscopic world of electrons and the macroscopic technologies that define modern life—and between those technologies and the quiet beauty of light itself.

Macroscopic Maxwell Equations

In the preceding lessons you watched electromagnetic fields interact with conductors and magnetic materials, seeing how free charges rearrange and how microscopic moments align. Those descriptions already employed averaged quantities—conductivity tensors, permeability tensors, magnetization l7_1.png—yet the underlying Maxwell equations themselves remained microscopic. Now you will see, carefully and from first principles, how the rapid, chaotic fields that exist between atoms can be smoothed into the macroscopic fields l7_2.png and l7_3.png that govern every practical device, every waveguide, and every optical instrument.

The procedure is one of nature’s quiet masterpieces where by averaging over volumes that are large compared to atomic distances yet small compared with the scale where the fields themselves vary, the frantic microscopic fluctuations disappear and a clean set of equations emerges. The price of that cleanliness is the introduction of two new fields—polarization l7_4.png and magnetization l7_5.png—that encode the averaged response of the material. Once those fields are defined, the familiar macroscopic Maxwell equations follow at once.

Averaging over Microscopic Fields to Obtain Equations in Matter l7_6.png)

We start with a physical idea first. At the atomic scale the electric and magnetic fields are violently irregular. Electrons orbit nuclei, spin magnetic moments precess, and bound charges oscillate. Maxwell’s equations remain exact for these microscopic fields l7_7.png and l7_8.png and the Maxwell equations become

l7_9.png

(7.1)

l7_10.png

(7.2)

l7_11.png

(7.3)

l7_12.png

(7.4)

Here l7_13.png and l7_14.png include every charge and every current—free and bound alike. For most macroscopic purposes, however, you never need this microscopic detail. What you measure with laboratory probes, what enters circuit equations, and what propagates through a waveguide is an averaged field.

Definition 7.1 Macroscopic Fields: The macroscopic electric and magnetic fields are the spatial averages

l7_15.png

(7.5)

and

l7_16.png

(7.6)

where the average is taken over a volume that is large compared with interatomic distances yet small compared with the distances over which the macroscopic fields change appreciably.

When you perform this average on the microscopic Maxwell equations, the free charges and currents survive as the macroscopic sources you control, while the bound charges and currents are absorbed into two new material fields.

The bound charge density arising from the displacement of charges within atoms and molecules can be written, after averaging,

l7_17.png

(7.7)

The vector field l7_18.png is called the polarization and it is the electric dipole moment per unit volume. In component form the same relation reads l7_19.png.

Exercise 7.1: A long dielectric cylinder of radius (a) carries a uniform polarization l7_20.png parallel to its axis. Calculate the bound volume charge density l7_21.png and the bound surface charge density l7_22.png. Show that the total bound charge on any finite length of the cylinder is zero, and interpret the result physically.

Likewise, the averaged effect of all microscopic circulating currents (orbital and spin) appears as an effective current density

l7_23.png

(7.8)

where l7_24.png is the magnetization—the magnetic moment per unit volume. (In tensor notation the same statement is l7_25.png.)

Exercise 7.2: A long cylinder of magnetic material of radius (a) has a uniform magnetization l7_26.png. Write the bound volume current density l7_27.png and the bound surface current density l7_28.png. Using Ampère’s law for the macroscopic l7_29.png field (with no free currents), find l7_30.png and l7_31.png both inside and outside the cylinder.

It is now convenient to define two new macroscopic fields that absorb the material contributions

l7_32.png

(7.9)

and

l7_33.png

(7.10)

Exercise 7.3: In a linear isotropic dielectric with permittivity ε, a free point charge l7_34.png is placed at the origin. Write the macroscopic equation for l7_35.png and solve for l7_36.png
. Then obtain l7_37.png and the bound charge density. Explain why the total charge (free + bound) that appears in Gauss’s law for l7_38.png is smaller than l7_39.png.

With these definitions the averaged Maxwell equations take their macroscopic form

l7_40.png

(7.11)

l7_41.png

(7.12)

l7_42.png

(7.13)

and

l7_43.png

(7.14)

Exercise 7.4: Show this to be true.

Only free charges and free currents appear on the right-hand sides. All of the complicated microscopic physics has been neatly packaged inside l7_44.png and l7_45.png, or equivalently inside the constitutive relations that link l7_46.png and l7_47.png to l7_48.png and l7_49.png.

Principle 7.1: Averaging does not alter the fundamental character of Maxwell’s equations, instead it merely separates the sources you can control from the sources that the material itself produces. The fields l7_50.png and l7_51.png are bookkeeping devices of extraordinary power where they let you write macroscopic boundary-value problems whose solutions describe every practical electromagnetic device.

Exercise 7.5: Inside a linear isotropic medium the constitutive relations are l7_52.png and l7_53.png. Starting from the macroscopic Maxwell equations, derive the wave equation for l7_54.png in a source-free region. Identify the phase velocity and compare it with the vacuum value c. What happens to the wave equation if the medium is weakly anisotropic, so that l7_55.png is a tensor?

For example,  inside a dielectric the free charge may be zero, yet l7_56.png does not imply l7_57.png. The polarization charge l7_58.png still shapes the microscopic field, and the macroscopic l7_59.png is reduced accordingly—precisely the origin of the dielectric constant you already used in earlier wave-propagation calculations.

The averaging procedure assumes that the macroscopic fields vary slowly on the scale of the averaging volume. When the wavelength becomes comparable to atomic distances (with X-rays, for example) the separation into free and bound contributions must be revisited, and a more microscopic treatment is required.

You now possess the clean macroscopic platform on which the remainder of Lesson 7 rests. Waveguides, radiation, scattering, and optics will all be written in terms of l7_60.png,
l7_61.png, l7_62.png, and l7_63.png. The microscopic storm has been tamed so that the elegant macroscopic equations remain.

Exercise 7.6: At a planar interface between vacuum and a linear dielectric there is no free surface charge or free surface current. State which components of l7_64.png, l7_65.png, l7_66.png, and l7_67.png
are continuous across the interface. Sketch the field lines of l7_68.png and l7_69.png for an incident field that is neither purely parallel nor purely perpendicular to the surface, and explain the difference in terms of bound surface charge.

Exercise 7.7: The averaging procedure that produces l7_70.png and l7_71.png discards an enormous amount of microscopic detail, yet the resulting macroscopic equations remain exact for all practical purposes at ordinary wavelengths. Reflect on why this is possible: what physical scale separation justifies the procedure, and what would happen if the wavelength became comparable to atomic distances? How does the clean separation into free and bound sources mirror the separation you already used between free and bound charges in electrostatics? What larger sense of wonder does this averaging evoke about the way nature allows simple macroscopic laws to emerge from complex microscopic motion?

Waveguides and Cavities

In the preceding section you saw how spatial averaging yields the macroscopic Maxwell equations written in terms of l7_72.png, l7_73.png, l7_74.png, and l7_75.png. Those equations govern the fields inside ordinary materials; they also govern the fields that can be confined by conducting boundaries. When the boundaries form a long hollow tube, the fields become guided waves that travel along the tube. When the tube is closed at both ends, the fields form standing-wave resonances. These two geometries—waveguides and resonant cavities—are among the most important practical applications of the macroscopic theory.

You already know that plane waves in free space or in unbounded media can propagate at any frequency. Boundaries change that freedom. They quantize the allowed transverse patterns and introduce a lowest frequency below which no energy can travel. The mathematics is a direct application of the macroscopic equations together with the boundary conditions on perfect (or good) conductors that you met in Lesson 6. The physical reward is the ability to transport microwave power, store electromagnetic energy, and filter frequencies with exquisite precision.

Propagation in Waveguides

Imagine a hollow metal pipe of rectangular or circular cross-section. An electromagnetic wave that tries to travel down the pipe must satisfy the macroscopic Maxwell equations inside the empty interior and must also obey the boundary conditions on the walls, where the tangential component of l7_76.png vanishes and the normal component of l7_77.png vanishes (for perfect conductors). These constraints force the fields to arrange themselves into discrete modes. Each mode is labeled by a pair of integers that count the number of half-wavelength variations across the two transverse directions.

For a rectangular waveguide with cross-section 0xa, 0yb and propagation along z, the fields can be written in the form

l7_78.png

(7.15)

and likewise for l7_79.png.

Definition 7.2 Waveguide Mode: A waveguide mode is a solution of the macroscopic Maxwell equations that satisfies the boundary conditions on the guide walls and propagates (or decays) along the guide axis. Two families appear: TE modes l7_80.png) and TM modes l7_81.png). A pure TEM mode (l7_82.png) cannot exist inside a hollow single-conductor waveguide; it requires at least two separate conductors (as in a coaxial line).

The transverse wave numbers are fixed by the boundary conditions

l7_83.png

(7.16)

(with the restriction that m and n cannot both be zero for TE or TM modes). The longitudinal wave number k then follows from the dispersion relation obtained by substituting into the macroscopic wave equation

l7_84.png

(7.17)

In a standard X-band rectangular waveguide (a2.3 cm, b1.0 cm) the dominant l7_85.png mode carries power efficiently near 10 GHz while higher modes remain cutoff. This is why microwave links, radar systems, and particle-accelerator RF structures rely on carefully dimensioned waveguides.

Cutoff Frequencies

The expression under the square root for (k) must be positive for the wave to propagate without exponential decay. Consequently every mode possesses a cutoff frequency

l7_86.png

(7.18)

Below l7_87.png the wave number becomes imaginary and the mode is evanescent, where the fields decay exponentially along the guide and no time-averaged power is transported.

Principle 7.2: Cutoff is a direct geometric consequence of the boundary conditions. The walls force a minimum transverse curvature on the fields; that curvature costs a minimum frequency before the fields can also vary along the propagation direction.

For the dominant l7_88.png mode the cutoff reduces to the simple result

l7_89.png

(7.19)

(in vacuum). All higher modes have higher cutoffs, so a frequency band exists in which only the dominant mode propagates—an essential feature for single-mode operation.

Real walls have finite conductivity, so even above cutoff there is a small attenuation. The skin-effect analysis of Lesson 6 supplies the correction; the ideal cutoff frequency itself remains an excellent first approximation.

Resonant Cavities

Close both ends of a waveguide with conducting plates and the traveling waves become standing waves. The structure is now a resonant cavity. The longitudinal wave number is likewise quantized

l7_90.png

(7.20)

(where d is the cavity length). The resonant frequencies follow at once

l7_91.png

(7.21)

Each triplet ((m,n,p)) labels a discrete resonant mode. At resonance the cavity stores electromagnetic energy that oscillates between electric and magnetic forms, exactly as a lumped LC circuit does, but now distributed throughout the volume.

Definition 7.3 Resonant Cavity: A closed conducting enclosure that supports discrete electromagnetic eigenmodes at specific frequencies.

The quality factor q measures how long the energy remains stored before it is dissipated in the walls or coupled out through an aperture. High-q cavities are the heart of microwave oscillators, particle accelerators, and precision frequency standards.

A cylindrical cavity operating in the l7_92.png mode can reach q values of tens of thousands; the same geometry appears in hydrogen masers and in the accelerating structures of modern linacs.

Principle 7.3: Cavities quantize the field in all three dimensions. The continuous spectrum of free-space waves collapses into a discrete set of sharp resonances—nature’s way of turning a continuous electromagnetic continuum into a set of precise, usable tones.

You now possess the essential language of guided and confined electromagnetic fields. The macroscopic Maxwell equations, together with simple boundary conditions, have produced cutoff frequencies, discrete modes, and high-q storage of energy. These ideas will reappear when you study radiation, antennas, and the optical cavities that underlie lasers.

Radiation from Accelerated Charges

In the preceding sections you watched electromagnetic fields confined by conducting walls, quantized into waveguide modes, and stored as sharp resonances inside cavities. Those fields were already present and the boundaries merely shaped them. Now the question reverses, “How are electromagnetic waves created in the first place?”

The answer is both simple and profound. Steady currents and static charges produce only static or quasi-static fields. Whenever a charge accelerates, however, the electromagnetic field cannot adjust instantaneously across all space. The mismatch between the old field configuration and the new one propagates outward at the speed of light as a genuine electromagnetic wave. Radiation is therefore the inevitable consequence of an accelerated charge—the mechanism by which the motion of every electron, every oscillating dipole, and every antenna ultimately becomes light, radio waves, or X-rays.

You will see this process first through the total power carried away by the wave (the Larmor formula), then through the simplest and most common source (the oscillating electric dipole), and finally through the detailed structure of the fields in the radiation zone. Throughout, the same macroscopic Maxwell equations you have already mastered remain the governing framework; only the sources are now time-dependent and localized.

Larmor Formula

Consider a single non-relativistic point charge Q whose acceleration is l7_93.png. At each instant the charge is “trying” to update the Coulomb field that surrounds it. Because the update can travel only at finite speed c, a transverse electromagnetic disturbance is left behind. The energy carried away by that disturbance per unit time is the radiated power.

A detailed calculation from the retarded potentials (or from the Poynting vector integrated over a large sphere) yields the celebrated Larmor formula

l7_94.png

(7.22)

where a=∣a∣ is the instantaneous acceleration. The power depends only on the square of the acceleration and is independent of velocity (in the non-relativistic limit).

Definition 7.4 The Larmor Formula: (7.22) gives the total instantaneous power radiated by a non-relativistic accelerated charge.

The factor l7_95.png tells you that uniform motion produces no radiation—exactly as we will see from special relativity later. Only a change in velocity creates the mismatch that becomes a wave. The same formula explains why electrons circling in a synchrotron emit intense beams and why a charge falling in a gravitational field (in the equivalence-principle sense) must radiate when observed from an inertial frame.

When the speed approaches c, the Larmor formula is replaced by the relativistic Liénard result, which contains an extra factor l7_96.png and angular dependence. For the present non-relativistic discussion the original Larmor expression is sufficient and exact.

Dipole Radiation

Most practical radiators—atoms, molecules, radio antennas—are not single free charges but collections of charges whose net monopole moment is constant. The leading time-varying multipole is then the electric dipole moment

l7_97.png

(7.23)

When p(t) oscillates, the system radiates. In the long-wavelength (dipole) approximation the radiated power, averaged over a cycle, becomes

l7_98.png

(7.24)

where l7_99.png is the amplitude of the oscillating dipole and ω its angular frequency.  

Definition 7.5 Dipole Radiation: The electromagnetic radiation emitted by a time-varying electric dipole moment in the limit that the source size is much smaller than the wavelength.

The l7_100.png dependence is decisive, allowing higher frequencies to radiate far more efficiently. This is why the sky is blue (due to Rayleigh scattering scales as l7_101.png) and why ultraviolet transitions in atoms are typically stronger than infrared ones.

For example, a classical model of an atom consists of an electron bound by a spring-like restoring force. Driven by an external field, the electron’s dipole moment oscillates and re-radiates; the process is ordinary dipole radiation. The same mathematics describes a short center-fed linear antenna when its length is much less than λ.

Principle 7.4: Nature prefers to radiate through the lowest-order multipole that is allowed. When the dipole moment vanishes by symmetry, quadrupole or magnetic-dipole radiation takes over—still higher powers of ω/c suppress the emission still further.

Radiation Fields

Close to an accelerating charge the fields are a complicated mixture of Coulomb-like near fields (falling as l7_102.png) and intermediate induction fields (also falling as l7_103.png). Far from the source—in the radiation zone where rλ and r≫ source size—only the true radiation fields survive. They fall as 1/r, are purely transverse, and carry a time-averaged power that remains constant through any large sphere.

For a non-relativistic charge the radiation electric field is

l7_104.png

(7.25)

where l7_105.png is the unit vector from the retarded position of the charge to the observation point, and the subscript “ret” means the quantity is evaluated at the retarded time l7_106.png, is the earlier moment at which a signal traveling at the speed of light must have left a distant source in order to reach the observer at the present time t. It appears throughout radiation theory because electromagnetic influences cannot propagate instantaneously; every field you measure now was generated by charges and currents at their positions and velocities at the corresponding retarded time. The accompanying magnetic field is simply

l7_107.png

(7.26)

Definition 7.6 The Radiation Fields: The pieces of l7_108.png and l7_109.png that decay as 1/r and transport energy to infinity.

These fields are transverse, l7_110.png and l7_111.png. Their Poynting vector points radially outward and, when integrated over a distant sphere, recovers exactly the Larmor power. The angular distribution is the familiar l7_112.pngθ pattern of dipole radiation—maximum broadside to the acceleration, zero along the acceleration axis.

Because the radiation fields fall so slowly, an arbitrarily small acceleration can still produce a detectable wave at enormous distances. Every photon that reaches your eye, every radio signal from a distant spacecraft, and every X-ray from an astronomical source began as a radiation field of this kind.

You have now traced the path from an accelerating charge to the electromagnetic wave that travels across the universe. The Larmor formula quantifies the total power, the dipole approximation organizes the most common sources, and the radiation fields themselves reveal the transverse, 1/r structure that carries energy, momentum, and information. These results will reappear when you study multipole expansions, antennas, and the scattering of light.

Multipole Radiation

In the preceding section you saw that any accelerated charge radiates, that the total power is given by the Larmor formula, and that the simplest oscillating system—an electric dipole—produces the familiar l7_113.png pattern. Most real sources, however, are neither single free charges nor pure dipoles. They are localized collections of charges and currents whose radiation can be organized, term by term, according to the multipole moments of the source.

The multipole expansion is nature’s systematic bookkeeping. When the source is small compared with the wavelength, the radiated fields can be written as a series ordered by increasing powers of the small parameter k a (where a is the source size and k=ω/c). The leading term is usually electric-dipole radiation; if that moment vanishes by symmetry, the next terms—magnetic dipole or electric quadrupole—take over. Each successive multipole radiates more weakly at long wavelengths and carries a distinctive angular pattern and parity.  

We will first examine the electric multipole series, then the magnetic multipole series. Throughout, the radiation-zone fields you already know remain the foundation; only the source description becomes more refined.

Electric Multipole Expansion of Radiation Fields

Any localized distribution of charge can be characterized by its multipole moments where the total charge (monopole), the electric dipole moment l7_114.png, the electric quadrupole tensor l7_115.png and higher moments. Because a static monopole does not radiate, the first radiating term is the time-varying electric dipole. When that vanishes, the electric quadrupole becomes the leading electric contribution.

In the radiation zone the electric-dipole fields are exactly those you met earlier

l7_116.png

(7.27)

and

l7_117.png

(7.28)

(The real part is understood.) The time-averaged power is the dipole result already derived,

l7_118.png

Definition 7.7 The Electric Multipole Expansion of the Radiation Field: The systematic series in which successive terms are generated by the electric dipole, quadrupole, octupole, ... moments of the source.

When the dipole moment is identically zero (for example, by symmetry in a pure quadrupole oscillator), the next electric term is the electric quadrupole. Its radiation fields involve the second time derivative of the quadrupole tensor and fall with an extra factor of k. The radiated power then scales as l7_119.png, two powers higher than the dipole, confirming that quadrupole radiation is suppressed at long wavelengths.

For example, atomic transitions are classified as E1, E2, ... according to the multipole that connects the initial and final states. Ordinary optical lines are overwhelmingly electric-dipole (E1). Forbidden lines that appear in planetary nebulae are often electric-quadrupole (E2) or magnetic-dipole (M1) transitions; their weakness is precisely the higher power of k a predicted by the multipole expansion.

Principle 7.5: Each additional electric multipole order brings an extra factor of k a in the amplitude and therefore an extra l7_120.png in the power. Nature therefore radiates most efficiently through the lowest allowed electric multipole.

Magnetic Multipole Expansion of Radiation Fields

Currents as well as charges produce radiation. A localized current distribution possesses a magnetic dipole moment

l7_121.png

(7.29)

and higher magnetic multipoles. Even if all electric multipoles vanish, a time-varying magnetic moment still radiates.

In the radiation zone the magnetic-dipole fields are

l7_122.png

(7.30)

and

l7_123.png

(7.31)

(The roles of l7_124.png and l7_125.png are interchanged relative to the electric-dipole case, and the overall parity is opposite.) The time-averaged power is

l7_126.png

identical in form to the electric-dipole power except for the replacement p->m/c.

Definition 7.9 The Magnetic Multipole Expansion: This organizes the radiation according to the magnetic dipole, quadrupole, ... moments of the current distribution.

Higher magnetic multipoles again carry extra factors of k a. In practice the magnetic dipole (M1) is the most important magnetic term; magnetic quadrupole radiation is rarely observed at optical frequencies.

For example, the 21 cm hydrogen line is a magnetic-dipole transition. Its extreme weakness (lifetime of order l7_127.png years) follows directly from the smallness of k a for a nuclear-sized source at radio wavelengths. The same magnetic-dipole radiation appears in the spin-flip transitions of electrons in atoms and in certain engineered microwave antennas.

Principle 7.6: Magnetic multipole radiation is generated by the circulating currents of the source. Because a magnetic moment has different parity and transformation properties from an electric moment, the angular patterns and selection rules are distinct—allowing experimental separation of the two families.

In a complete multipole treatment the total radiation field is the coherent sum of all electric and magnetic multipoles. For a source small compared with the wavelength the hierarchy is clear:

E1 (the electric dipole)

M1 (the magnetic dipole) and E2 (the electric quadrupole)

Higher orders

Each successive order is weaker by a factor of order l7_128.png in power. This ordered expansion is the reason a short antenna can be treated as a pure electric dipole, why certain atomic lines are “forbidden,” and why the design of efficient radiators always begins with the lowest allowed multipole.

You now possess a systematic language for any localized radiating system. The electric and magnetic multipole expansions together convert an arbitrary charge–current distribution into a hierarchy of simple, universal radiation patterns. These patterns will reappear when you study scattering, antennas, and the optical transitions that fill the spectrum of the universe.

Scattering of Electromagnetic Waves

In the preceding sections you watched accelerated charges produce radiation fields and saw how those fields can be organized into a multipole series. Scattering is the natural next step where an already-existing electromagnetic wave encounters a charge, an atom, or a small particle, sets it into motion, and thereby generates new radiation that travels outward in directions different from the original beam. The incident wave is partially extinguished; the difference appears as scattered light.

Scattering is therefore radiation driven by an external field. All of the multipole machinery you have just acquired applies at once. When the scatterer is a free electron the process is especially simple (Thomson scattering). When the electron is bound by a restoring force the frequency dependence becomes dramatic (Rayleigh scattering). When the driving frequency approaches a natural resonance of the bound system the scattered power can become enormous (resonance scattering). These three classical regimes explain everything from the blue sky to the sharp colors of atomic vapors and the behavior of free electrons in a plasma or in the solar corona.

Thomson Scattering

A free electron in an electromagnetic wave experiences the Lorentz force and accelerates. That acceleration, according to the Larmor formula, produces radiation. Because the electron is free, its response is independent of frequency (provided the motion remains non-relativistic and the wavelength is long compared with the classical electron radius). The scattered radiation has the same frequency as the incident wave and a characteristic dipole angular pattern.

The time-averaged power radiated by the electron is most conveniently expressed through the Thomson cross-section

l7_129.png

(7.32)

where l7_130.png is the classical electron radius. The differential cross-section has the familiar l7_131.png dipole shape

Definition 7.10 Thomson Scattering: The classical scattering of electromagnetic waves by free electrons in the low-intensity, non-relativistic limit.

For example, X-rays scattered by the electrons in the solar corona, or by the free electrons in a laboratory plasma, follow the Thomson formula (7.32). The process is also the classical limit of Compton scattering when the photon energy is much smaller than the electron rest energy.

Principle 7.7: Because a free electron has no intrinsic frequency scale, the Thomson cross-section is independent of wavelength. The scattered light is polarized in the plane of the incident polarization, a fact used routinely in laboratory diagnostics.

Rayleigh Scattering

When the electron is bound—by the restoring force of an atom, a molecule, or a small dielectric particle—the acceleration is no longer simply l7_132.png. In the long-wavelength limit the induced dipole moment is proportional to the incident field, l7_133.png, and the radiated power is ordinary dipole radiation. Because the polarizability α is roughly frequency-independent far below any resonance, the scattered power inherits the l7_134.png factor of the dipole formula.

The Rayleigh cross-section for a single scatterer therefore reads

l7_135.png

(7.33)

Definition 7.11 Rayleigh Scattering: The scattering of electromagnetic waves by bound charges or small particles in the long-wavelength (dipole) limit, characterized by a steep l7_136.png frequency dependence.

For example, Sunlight passing through the atmosphere is Rayleigh-scattered by N₂ and O₂ molecules. Blue light, having roughly twice the frequency of red light, is scattered roughly sixteen times more strongly—hence the blue sky. At sunset the long path length removes most of the blue, leaving the reds and oranges that reach your eye.

Principle 7.8: The same multipole hierarchy you studied earlier appears here, where the induced dipole is the leading term, and the l7_137.png dependence is simply the dipole radiation factor evaluated for a driven oscillator. When the particle size becomes comparable to the wavelength the Rayleigh approximation fails and one must move to Mie theory. Where Mie theory is the exact analytical solution of Maxwell’s equations for the scattering of a plane electromagnetic wave by a homogeneous sphere of arbitrary size and refractive index. When the sphere is much smaller than the wavelength it reduces to Rayleigh scattering, but for particles comparable to or larger than the wavelength it fully accounts for the complicated angular patterns, polarization, and extinction that appear in clouds, colloidal suspensions, and atmospheric aerosols.

Resonance Scattering

If the bound electron is modeled as a classical harmonic oscillator with natural frequency l7_138.png and damping constant γ, its polarizability becomes strongly frequency-dependent

l7_139.png

(7.34)

When the driving frequency ω approaches l7_140.png, ∣α∣ rises dramatically. The scattered power therefore exhibits a sharp peak—the resonance scattering (or resonant fluorescence) line.

Near resonance the cross-section can exceed the Thomson value by many orders of magnitude; in the classical oscillator model its peak value is of order l7_141.png, a geometric area set by the wavelength itself rather than by the tiny classical electron radius.

Definition 7.12 Resonance Scattering: This occurs when the frequency of the incident wave lies near a natural oscillation frequency of the scatterer, producing a large, sharply peaked cross-section.

For example, the brilliant yellow D-lines of sodium vapor, the absorption and re-emission of sunlight in the solar chromosphere, and the operation of a resonant optical cavity all rely on this enormous enhancement. In the laboratory a dilute gas of atoms illuminated near an atomic resonance can scatter light with astonishing efficiency.

The classical oscillator model captures the essential resonance structure but omits quantum selection rules, saturation, and spontaneous-emission details. Nevertheless it already explains why certain narrow spectral lines dominate the scattering spectrum and why lasers can interact so strongly with atomic vapors.

You have now traced three classical regimes of scattering—free-electron, bound but non-resonant, and resonant—each built directly on the radiation and multipole ideas of the preceding sections. Together they account for the color of the sky, the visibility of free electrons in plasmas, and the brilliant selectivity of atomic resonances. These processes will reappear when you study antennas, optical devices, and the propagation of light through continuous media.

Antennas and Radiation Patterns

In the preceding sections you watched accelerated charges and oscillating multipoles produce radiation fields, and you saw how an incident wave can drive a scatterer to re-radiate. An antenna is simply a practical device that deliberately engineers those same processes: oscillating currents are arranged on a conducting structure so that the resulting radiation is launched efficiently into free space (or into a waveguide) with a desired pattern of intensity and polarization.

Everything you have already learned about dipole radiation, multipole expansions, and radiation-zone fields applies directly. The new elements are the engineering language that quantifies how much power is actually radiated, how that power is distributed in angle, and how the antenna appears as a circuit element to its driving source. You will first examine the basic theoretical framework, then the concept of radiation resistance, and finally the measure of directional concentration known as directivity.

Basic Antenna Theory

Any antenna is a localized distribution of time-varying current. In the radiation zone the fields are completely determined by the multipole moments of that current distribution—most often the electric-dipole moment for short antennas, or a coherent superposition of multipoles for longer or more complex structures. The far-zone electric field retains the universal transverse, 1/r character you met earlier

l7_142.png

(7.35)

where l7_143.png is the impedance of free space and l7_144.png is the vector pattern function fixed by the antenna geometry and current distribution.

Definition 7.13 Antenna: A structure that converts guided electromagnetic energy into radiating waves (or vice versa) by supporting a controlled distribution of oscillating currents.

Because Maxwell’s equations are linear and time-reversal invariant, antennas obey the reciprocity theorem: the pattern of radiation when the antenna is used as a transmitter is identical to its pattern of sensitivity when used as a receiver. This single principle lets you design a transmitting array and immediately know how it will perform as a receiving array.

For example, the simplest practical antenna is the short dipole (length λ). Its pattern function is exactly the sin ⁡θ factor of electric-dipole radiation, and its radiated fields are those you already derived for an oscillating dipole moment l7_145.png.

Radiation Resistance

From the circuit point of view an antenna is a two-terminal device. Part of the power delivered by the source is dissipated as heat in the metal (providing an ohmic loss); the remainder is carried away by the radiation field. It is convenient to represent the radiated power by an equivalent resistance—the radiation resistance—so that ordinary circuit analysis can be used.

If the peak current at the feed point is l7_146.png and the time-averaged radiated power is P , the radiation resistance is defined by

l7_147.png

(7.36)

Definition 7.14 Radiation Resistance: The effective resistance that accounts for the power carried away from an antenna by electromagnetic radiation.

For a short dipole the result is

l7_148.png

(7.37)

a value that is only a few ohms when λ. That is why short antennas are inefficient unless careful matching networks are used. A half-wave dipole, by contrast, has l7_149.png , a convenient match to standard transmission lines.

Principle 7.9: Radiation resistance converts the field-theoretic calculation of radiated power into a simple circuit parameter. Maximizing l7_150.png relative to the ohmic resistance is one of the central goals of antenna design.

Directivity

Most antennas do not radiate uniformly in all directions. The directivity quantifies how strongly the radiation is concentrated in the preferred direction. If U(θ,φ) is the radiation intensity (power per steradian) and l7_151.png is the intensity of an isotropic radiator emitting the same total power, the directivity is

l7_152.png

(7.38)

The maximum value of D is simply called the directivity of the antenna. An isotropic radiator has D=1 (or 0 dBi); a half-wave dipole has D1.64 (or 2.15 dBi); large parabolic reflectors can reach tens of thousands.

Definition 7.15 Directivity: The ratio of the radiation intensity in a given direction to the intensity that would be produced by an isotropic radiator emitting the same total power.

For example, a Yagi–Uda array or a phased array increases directivity by constructive interference in one direction and destructive interference elsewhere—precisely the multipole interference you studied earlier, now engineered for a practical beam.

Directivity describes only the shape of the pattern; it does not include ohmic losses. The gain of an antenna is the directivity multiplied by the radiation efficiency, and is therefore the quantity that appears in real link-budget calculations.

You have now linked the fundamental radiation fields of accelerating charges to the practical language of antennas. Basic antenna theory supplies the pattern function, radiation resistance converts radiated power into a circuit element, and directivity measures how skillfully that power has been concentrated. These three concepts underlie every radio transmitter, every satellite link, and every radar system.

Electromagnetic Momentum and Angular Momentum

In the preceding sections you watched electromagnetic waves carry energy from accelerating charges, through multipole fields, past scatterers, and out of antennas. Energy, however, is only one of the conserved quantities stored in the field. The same waves also carry linear momentum and angular momentum. When light reflects from a mirror it exerts radiation pressure; when a circularly polarized beam is absorbed it transfers torque; when a helical wavefront winds around its axis it can set microscopic particles spinning. These mechanical effects are not optional add-ons—they are required by the conservation laws that follow directly from Maxwell’s equations.

You will first see how the conservation of energy, momentum, and angular momentum emerge in local form from the macroscopic field equations. Then you will examine the two distinct contributions that electromagnetic angular momentum can make in both an intrinsic spin part associated with polarization and an orbital part associated with the spatial structure of the wavefront.

Conservation Laws

The macroscopic Maxwell equations imply local continuity equations for every conserved quantity. You already know the Poynting theorem with the energy density

l7_153.png

(7.39)

and the energy flux (Poynting vector)

l7_154.png

(7.40)

together satisfy

l7_155.png

(7.41)

The same logic applied to momentum yields an electromagnetic momentum density

l7_156.png

(7.42)

(in vacuum) and a corresponding stress tensor—the Maxwell stress tensor—whose divergence accounts for the rate of change of field momentum. When the fields interact with matter, the mechanical momentum of the charges changes so that the total momentum (field plus mechanical) is conserved.

Definition 7.16 The Electromagnetic Momentum Density: The field contribution to the total linear momentum per unit volume; its flux is described by the Maxwell stress tensor.

Where the Maxwell stress tensor is the natural object that encodes the flux of electromagnetic momentum, just as the Poynting vector encodes the flux of energy. Once you accept that the fields themselves carry a momentum density l7_157.png, the requirement of local momentum conservation forces the existence of a rank-2 tensor T whose divergence gives the rate at which that momentum flows out of a volume. Explicitly, the Cartesian components in vacuum read

l7_158.png

(7.43)

the force on the charges and currents inside any volume is then equal to the integral of l7_159.png’s time derivative plus the surface integral of T. In this way the tensor unifies radiation pressure, the attraction of parallel currents, and the mechanical stresses that appear at the surface of a dielectric—all as manifestations of the same momentum flow already latent in Maxwell’s equations.

Angular momentum follows by taking the moment of the momentum density. The electromagnetic angular-momentum density is

l7_160.png

(7.44)

and a corresponding continuity equation guarantees that the total angular momentum (field plus mechanical) is conserved.

For example, a plane wave incident on a perfect absorber delivers both energy and momentum. The radiation pressure is simply l7_161.png, a direct consequence of the momentum density l7_162.png. The same pressure accelerates dust grains in interstellar space and must be included in the dynamics of precision spacecraft.

Principle 7.10: Energy, linear momentum, and angular momentum are not independent inventions; they arise together once the field is recognized as a dynamical physical system governed by Maxwell’s equations. The conservation laws are local and exact (in the absence of dissipation or when mechanical contributions are restored).

Spin and Orbital Angular Momentum of Light

The total electromagnetic angular momentum of a free wave can be partitioned, in the paraxial regime, into two contributions that transform differently under rotations and that can be measured independently.

Spin angular momentum is an intrinsic property linked to the polarization. A circularly polarized plane wave carries ±h of angular momentum along its propagation direction for each photon (or, classically, a torque density proportional to the degree of circular polarization). Linearly polarized light carries none.

Orbital angular momentum is associated with the spatial structure of the wavefront. A beam whose phase winds as l7_163.png around its axis (a Laguerre–Gaussian mode, for example) carries l h of orbital angular momentum per photon. Classically this appears as a circulating energy flow around the beam axis.

Definition 7.17 The Spin Angular Momentum of Light: The intrinsic, polarization-dependent contribution to the total angular momentum; the orbital angular momentum is the extrinsic contribution arising from the helical phase structure of the wavefront.

Both forms are transferred to matter when the light is absorbed or scattered. A circularly polarized beam can spin a microscopic birefringent particle; a beam with orbital angular momentum can set the same particle orbiting around the beam axis. Optical tweezers routinely exploit both effects.

For example, in atomic physics the selection rule Δ m=±1 for circularly polarized light is the quantum manifestation of spin-angular-momentum transfer. In modern optics, beams with l0 are used to rotate microparticles, to encode information, and to probe the rotational properties of materials.

The clean separation into spin and orbital parts is unambiguous for paraxial beams in free space. In strongly focused fields or inside media the separation becomes more subtle and gauge-dependent; the total angular momentum, however, remains well-defined and conserved.

You have now seen that electromagnetic waves are not mere carriers of energy. They transport linear momentum (giving rise to radiation pressure) and angular momentum (both spin and orbital). These mechanical properties follow directly from the same Maxwell equations that govern radiation, scattering, and antennas, and they open a rich interface between electromagnetism and mechanics that continues to be explored in laboratories and in the cosmos.

Applications to Optics

Throughout this lesson you have followed electromagnetic fields from their macroscopic foundations, through waveguides and cavities, out into radiation, multipoles, scattering, antennas, and even the momentum they carry. Optics is simply the same theory applied to wavelengths that are small compared with ordinary laboratory scales. When the wavelength is only a few hundred nanometers, the familiar phenomena of interference, diffraction, and image formation emerge as direct consequences of Maxwell’s equations. No new principles are required—only the recognition that light is a transverse electromagnetic wave and that every optical element is a boundary-value problem for that wave.

You will first see how the linear superposition inherent in Maxwell’s equations produces interference, then how the same equations, applied to an aperture or obstacle, generate diffraction, and finally how the entire process of propagation and imaging can be recast as a Fourier transform of the electromagnetic field—the domain of Fourier optics.

Interference

Because the macroscopic Maxwell equations are linear, any two solutions may be added. When the solutions are monochromatic waves that maintain a definite phase relation, their electric fields add coherently and the observed intensity—proportional to the time-averaged l7_164.png—contains cross terms. Those cross terms are the interference pattern.

For two plane waves of amplitudes l7_165.png and l7_166.png the intensity is

l7_167.png

(7.45)

where δ is the phase difference. Constructive and destructive interference are simply the extremes of this continuous expression.

Definition 7.18 Interference: The spatial or temporal modulation of intensity that results from the coherent superposition of two or more electromagnetic waves.

For example, Young’s double-slit experiment is the radiation from two coherent secondary sources (the slits) whose dipole-like fields overlap on a distant screen. The same mathematics describes the fringes in a Michelson interferometer, the colors of a soap bubble, and the standing-wave pattern inside a laser cavity.

Principle 7.11: Interference is not a new optical law; it is the linearity of Maxwell’s equations made visible once the wavelength is short enough for phase differences to vary rapidly across a laboratory apparatus.

Diffraction

When a wave encounters an obstacle or an aperture whose size is comparable to the wavelength, the simple ray picture fails. The field beyond the obstacle must still satisfy the macroscopic Maxwell equations and the boundary conditions on the screen. The rigorous solution of that boundary-value problem is diffraction.

In the scalar approximation (valid for many optical situations) the field in the observation plane is given by the Rayleigh–Sommerfeld or Kirchhoff integral—an expression that can be derived directly from Green’s theorem applied to the Helmholtz equation. Each point on the unobstructed portion of the wavefront acts as a secondary source (Huygens–Fresnel principle), and the observed field is the coherent superposition of all those secondary waves.

Definition 7.19 Diffraction: The deviation of an electromagnetic wave from rectilinear propagation caused by the presence of an obstacle or aperture, fully determined by Maxwell’s equations and the appropriate boundary conditions.

For example, the Airy pattern produced by a circular aperture is the Fourier transform of a uniform disk—exactly the far-field diffraction pattern of the aperture field. The same theory explains the resolving power of a telescope, the spread of a laser beam, and the colorful coronae seen around the Moon when thin clouds are present.

True vector diffraction must also enforce the divergence-free condition and the correct polarization boundary conditions; the scalar theory is an excellent approximation when angles are small and polarization effects are secondary.

Fourier Series

In the sections that follow you will see diffraction integrals turn into Fourier transforms and optical imaging become a matter of filtering spatial frequencies. Before those ideas can feel natural, you need a firm grasp of the simpler, periodic case that historically led to them, the Fourier series.  

The underlying physical idea is one you have already met many times. Any wave that repeats itself after a fixed interval—whether a periodic train of laser pulses, the field behind a regularly spaced diffraction grating, or the steady-state oscillation inside a cavity—can be built by adding together pure sinusoidal waves of discrete frequencies. The Fourier series is simply the precise mathematical statement of that construction. Once you can decompose a periodic electromagnetic waveform into its sinusoidal constituents, the step to non-periodic waves (the Fourier transform) and to the two-dimensional transforms of Fourier optics becomes almost immediate.

Consider a real-valued function f(t) that is periodic with period T, so f(t+T)=f(t). The fundamental angular frequency is l7_168.png. The claim of Fourier’s theorem is that, under very mild conditions, f(t) may be written as an infinite sum of harmonically related sines and cosines

l7_169.png

(7.46)

The coefficients are determined by orthogonality integrals over one period

l7_170.png

(7.47)

and

l7_171.png

(7.48)

Equivalently, and far more conveniently for later work, one may use the complex exponential form

l7_172.png

(7.49)

where

l7_173.png

(7.50)

Definition 7.19 A Fourier Series: The representation of a periodic function as a discrete superposition of complex exponentials (or sines and cosines) whose frequencies are integer multiples of the fundamental frequency.

Each term l7_174.png is a monochromatic wave at frequency l7_175.png. Because Maxwell’s equations are linear, the electromagnetic field produced by a periodic source is simply the sum of the fields produced by each Fourier component. The average power, the radiation pattern, and the response of any linear optical element can therefore be calculated one frequency at a time and then reassembled.

For example, a periodic train of short optical pulses has a broad spectrum of discrete lines spaced by the repetition frequency. The Fourier series makes that spectrum explicit: the shorter the pulse, the slower the decay of the coefficients l7_176.png, exactly as the uncertainty relation between time and frequency demands.

For another example, the transmission function of a diffraction grating is periodic in space. Its Fourier series coefficients are precisely the amplitudes of the various diffraction orders—an early glimpse of the Fourier-optics principle that a periodic structure selects discrete spatial frequencies.

Because the complex exponentials are eigenfunctions of differentiation and of time translation, the Fourier series behaves beautifully under the operations that appear in wave physics

Differentiation multiplies l7_177.png by l7_178.png.

Integration divides by l7_179.png (for n0).

A time shift l7_180.png multiplies every coefficient byl7_181.png.

These rules let you convert differential equations for linear circuits or driven oscillators into simple algebraic relations among the l7_182.png.

Principle 7.12: The Fourier series converts a periodic problem in the time (or space) domain into an algebraic problem in the discrete frequency domain. Linearity of Maxwell’s equations then guarantees that the electromagnetic response can be computed frequency by frequency.

When the period T is allowed to become infinitely long, the discrete spacing l7_183.png tends to zero and the sum over n becomes an integral over a continuous frequency variable. The Fourier series thereby turns into the Fourier transform. In optics the same limiting process is applied to spatial periods, where a single aperture is regarded as one period of an infinitely wide periodic array whose period has been stretched to infinity. The diffraction integral you will meet in the next section is exactly that continuous Fourier transform of the aperture field.

Thus, the modest one-dimensional series you have just written down is the conceptual seed of the entire edifice of Fourier optics. Once you are comfortable decomposing a periodic electromagnetic waveform into its harmonic constituents, the two-dimensional transforms that describe lenses, spatial filters, and coherent imaging will feel like natural generalizations rather than new inventions.

You now possess the essential language of Fourier series—both the real trigonometric form and the compact complex exponential form—together with the physical interpretation that links them directly to the linear wave equation. With this tool in hand you are ready to step from periodic trains of pulses and regular gratings into the continuous spatial-frequency domain of diffraction and imaging.

Fourier Transforms

In the preceding section you saw how any periodic electromagnetic waveform can be decomposed into a discrete sum of harmonic exponentials—the Fourier series. Most of the waves you encounter in optics and radiation, however, are not periodic. A single laser pulse, the field immediately behind an isolated aperture, or the transient radiation from an accelerating charge lasts for only a finite time or occupies only a finite region of space. To analyze these non-periodic signals you need the continuous counterpart of the Fourier series, the Fourier transform.

The physical idea remains exactly the same. You still express a complicated wave as a superposition of pure sinusoids; the only change is that the frequencies (or spatial frequencies) are no longer required to be discrete multiples of a fundamental. They form a continuum, and the discrete coefficients l7_184.png become a continuous amplitude density F(ω). Once that step is taken, every linear operation—propagation, diffraction, filtering by a lens—becomes a simple multiplication or phase shift in the frequency domain.

Let f(t) be a suitably well-behaved function of time (the generalization to a spatial coordinate x is immediate). Its Fourier transform is

l7_185.png

(7.51)

The original function is recovered by the inverse transform

l7_186.png

(7.52)

(The precise placement of the factor 2π and the sign convention vary among authors; the form above is common in physics.)

Definition 7.20 The Fourier Transform: Expresses an arbitrary (non-periodic) function as a continuous superposition of complex exponentials; the inverse transform reconstructs the function from its spectral amplitude F(ω).

When the variable is a transverse spatial coordinate x, one simply replaces tx and ωk; F(k x) is then the spatial-frequency spectrum of the aperture field—the central object of Fourier optics.

Because the complex exponentials are eigenfunctions of translation and differentiation, the Fourier transform converts the differential operations that appear in Maxwell’s equations into algebraic multiplications

Linearity: F(a f + b g)=a F +b G.

Time (or Space) Shift: l7_187.png

Frequency Shift: l7_188.png.

Differentiation: df/dt↔i ω F(ω).

Convolution theorem: the transform of a product is a convolution, and vice versa—precisely the relation that turns a multiplication by an aperture function into a convolution in the diffraction plane.

Parseval’s relation: l7_189.png, guaranteeing conservation of energy between the two domains.

Look at Parseval more closely

A short electromagnetic pulse (f(t)) has a broad spectrum F(ω). The shorter the pulse, the broader the spectrum—an immediate consequence of the scaling property of the transform. In the spatial domain the same statement reads: a narrow slit has a wide diffraction pattern. The Fourier transform therefore supplies the precise mathematical link between the size of a source (or aperture) and the angular spread of the radiated (or diffracted) wave.

When you expand an arbitrary monochromatic field in the plane z=0 as a superposition of plane waves,

l7_190.png

(7.53)

the amplitude l7_191.png is precisely the two-dimensional Fourier transform of the field. Propagation through free space then multiplies each plane-wave component by the simple phase factor l7_192.png with l7_193.png. This angular-spectrum representation is the foundation of the Fourier-optics treatment below.

Principle 7.13: The Fourier transform converts linear wave problems into algebraic problems in the frequency (or spatial-frequency) domain. Because Maxwell’s equations are linear and translation-invariant in free space, the complex exponentials are their natural eigenfunctions, and the transform is the tool that diagonalizes them.

With the Fourier transform in hand, the diffraction integral of the next section becomes transparent, with the far-field pattern of an aperture is the Fourier transform of the aperture field. A lens performs a Fourier transform between its front and back focal planes. Spatial filtering is multiplication of the transform by a mask. All of the concrete optical operations you will study are simply the physical realization of the mathematical properties listed above.

You have now completed the passage from discrete Fourier series for periodic waves to continuous Fourier transforms for pulses and apertures. The same electromagnetic field that can be written as a sum of discrete harmonics can equally well be written as an integral over a continuum of plane waves. That continuum representation is the language in which Fourier optics is written.

Fourier Optics Using Electromagnetic Theory

Once diffraction is expressed as an integral over the aperture field, a remarkable simplification appears. In the Fresnel or Fraunhofer regime the integral is precisely a Fourier transform. Propagation through free space, passage through a lens, and the formation of an image can all be rewritten as successive Fourier transformations of the transverse electromagnetic field.

A monochromatic field (l7_194.png) in an input plane evolves, after propagation a distance z, into

l7_195.png

(7.54)

where l7_196.png and F denotes the two-dimensional Fourier transform over the transverse coordinates. A thin lens simply multiplies the field by a quadratic phase factor, which is equivalent to a Fourier transform between its front and back focal planes.

Definition 7.21 Fourier Optics: The reformulation of diffraction and imaging as linear filtering operations on the spatial-frequency spectrum of the electromagnetic field.

For example, the Abbe theory of image formation follows at once, where  a microscope objective performs a Fourier transform of the object field; the aperture stop filters higher spatial frequencies; the tube lens performs the inverse transform. Resolution limits, spatial filtering, and the design of 4-f optical processors are all direct applications of this electromagnetic Fourier picture.

Principle 7.14: Because every solution of the Helmholtz equation can be decomposed into plane waves (the angular spectrum), free-space propagation is diagonal in the Fourier domain. Lenses, apertures, and gratings become simple multiplications or convolutions—turning the vector Maxwell problem into a practical language of spatial frequencies.

You have now closed the circle. The same macroscopic Maxwell equations that began this lesson, the same radiation fields that emerge from accelerated charges, and the same interference that follows from linearity, reappear as the foundational phenomena of optics. Interference is coherent superposition, diffraction is the boundary-value problem for those superpositions, and Fourier optics is the recognition that the Helmholtz equation is most naturally solved in the spatial-frequency domain. Light, radio waves, and X-rays are no longer separate subjects; they are simply electromagnetic waves viewed at different scales.

Doing This Stuff in Mathematica

Partial Differential Equations

Mathematica has many tools for handling PDEs. The documentation is quite extensive on this topic.

We will also adopt a convention of writing a general function of each of its arguments as an upper case letter, such as F. So a generic PDE can be written F(...)=G(...).

We can think of any differential equation as the application of a differential operator L to the function of our variables (or functions of our variables). Thus we can write a PDE as L(f)=F(...)=G(...).

In general, if we have a differential operator L such that for the functions u and v L(u+v)=L(u)+L(v) and for the constant or parameter c L(c u)=c L(u) then we say that the operator is linear. Any operator that is not linear is called nonlinear. It turns out that we can write any linear differential equation using a linear operator

Another criterion for classifying PDEs that we will consider here is the nature of G in the statement L(f)=F(...)=G(...). If G is identically 0 then the equation is homogeneous. Any PDE that has it defined by a linear operator that is homogenous is a linear PDE. If an equation is not homogeneous it is inhomogeneous. If a differential equation is linear and you have two solutions, then the sum of those solutions is another solution. This is called the superposition principle. It applies to ODEs and PDEs.

It is rare to find applications of PDEs that do not involve some kind of region where the solution needs to be found. If we symbolize the region with the Greek letter omega, Ω, we may then write its boundary as Ω. A wave propagating along a string is confined to the region of the string, for example. The endpoints of that string are the boundaries of the region. We will consider several types of boundary conditions.

The first kind of boundary condition we will consider is one where a specific value exists at the boundary. This is called a Dirichlet condition, named for work done by the German mathematician and physicist Peter Gustav Lejeune Dirichlet (1805-1859) published in 1828. Such Dirichlet conditions are usually in the form of some equation. There is also a Mathematica command that you can use with DSolve to establish Dirichlet conditions, it is (reasonably enough) called DirichletCondition[boundary equation,predicate], where the boundary equation is active when the predicate is met. We will see examples of this kind of boundary condition later, and we will examine the Mathematica command in later lessons. The act of finding a solution of this type is called a Dirichlet problem.

The second kind of boundary condition is one where there is a flux that is normal to the boundary defined as a partial derivative. What do we mean by normal? If we write l7_197.png as the unit normal vector on the boundary pointing outward from the region, then we define a normal derivative in terms of the unit normal vector and the directional derivative,

l7_198.png

(7.55)

If

l7_199.png

(7.56)

then (7.56) holds at the boundary then we have what we call a Neumann condition, named for the German mathematician Carl Neumann (1832-1925). In Mathematica there is a command called NeumannValue[value,predicate] that adopts the value whenever the predicate is met. This command is added to the existing PDE in DSolve. We will see examples of this later. The act of finding a solution of this type is called a Neumann problem.

The third kind of boundary condition arises when both the value of function u and its normal derivative are specified, then we call that a Cauchy boundary condition, and we call its solution a Cauchy problem. It is named for the French mathematician, physicist, and engineer  Augustin-Louis Cauchy (1789-1857).

The fourth kind of boundary occurs where both a linear combination of the values of function u and its normal derivative are specified,

l7_200.png

(7.57)

This is called a Robin condition, named for the French mathematician Victor Gustave Robin (1855-1897). It is interesting that we can use the NeumannValue command for Robin type boundary conditions, too.

The situation can exist where different boundaries of the same system have different, and disjoint, boundary conditions. Such a situation is called a mixed boundary condition.

It is also important to realize that these boundary conditions apply to ODEs and not just PDEs.

Examples

The simplest PDE is, given u(x,y)

l7_201.png

(7.58)

Solutions of this equation are of the form

l7_202.png

(7.59)

so whatever function we are applying to y is constant across x.

We can write a first-order equation with constant coefficients,

l7_203.png

(7.60)

By the method of characteristics, the characteristic curves for this equation satisfy the ODE

l7_204.png

(7.61)

This is a simple ODE to solve,

l7_205.png

(7.62)

By the nature of characteristic curves we know that u(x,y)=f(c).  Solving (7.62) for c we get,

l7_206.png

(7.63)

so the solution of (7.60) is

l7_207.png

(7.64)

We can also write a first-order equation with variable coefficients.

l7_208.png

(7.65)

If we again apply the method of characteristics, we get the ODE

l7_209.png

(7.66)

The solution of this ODE is

l7_210.png

(7.67)

and these are the characteristics. Solving for c gives us

l7_211.png

(7.68)

The solution to (7.65) is then,

l7_212.png

(7.69)

DSolve

Of course, we can use DSolve with PDEs in a manner similar to ODEs. Here we see the solution to our simplest PDE.

l7_213.png

l7_214.png

This is what we expected to see, the solution is some arbitrary function of y. What happens if we integrate this again?

l7_215.png

l7_216.png

Now we have the solution in terms of two arbitrary functions of y. We can examine some boundary value problems for this equation, too. Let’s say that whenever x=0 that u(x,y)=cos y/2.

l7_217.png

l7_218.png

Note that the boundary function replaces our arbitrary function.

Let’s try to solve the equation in (7.9). This is a slight modification of the example from the documentation.

l7_219.png

l7_220.png

We can use the same sort of solution for (7.14).

l7_221.png

l7_222.png

We can specify this arbitrary function using the rule-delayed command :→.

l7_223.png

l7_224.png

As an exercise read through the section on First-Order PDEs in the Scope section of the write up on DSolve.

As an exercise read through the sections Overview of PDEs and First-Order PDEs in the section PDEs in the tutorial Symbolic Differential Equation Solving.

As an exercise read through the Introduction to the tutorial Symbolic Solutions of PDEs.

DSolveValue

This is a command that does much the same thing as DSolve, but it does not return a rewrite rule like DSolve, it just returns a value. For example.

l7_225.png

l7_226.png

This negates the step of writing exp1 above.

l7_227.png

l7_228.png

As an exercise read through the write up for DSolveValue.

NDSolve

We can also apply the NDSolve command we used for ODEs. Again, the result is an interpolating function. For example, we can numerically solve our simplest PDE.

l7_229.png

l7_230.png

l7_231.png

l7_232.gif

Note the two warning messages. Convection-dominated PDEs are characterized by highly irregular solutions. What does that mean? It means there might be jumps, kinks, or other kinds of discontinuity. Let’s plot this and see what is happening.

l7_233.png

l7_234.gif

This is weird, there seems no value. Let’s look at the second message. It is suggesting some kind of boundary condition. Let’s say that we have the same condition we had above, whenever x=0 that u(x,y)=cos y/2.

l7_235.png

l7_236.gif

So the introduction of an initial condition seems to have solved the problem. Here is the new solution plot.

l7_237.png

l7_238.gif

What was the issue with the convection-dominated PDE message? The answer is based on an arbitrary function that is not defined.  Without an initial condition the solution surface is flat.

We need to do the same thing a second time for our second PDE from above.

l7_239.png

l7_240.png

l7_241.png

What do we use as a second constraint?

l7_242.png

l7_243.gif

Why not just write the derivative l7_244.png? Let’s see what happens.

l7_245.png

l7_246.png

l7_247.png

So, if you have an initial condition that does not seem to be working, try to change the way it is written. As an exercise look up the Derivative command and work through some examples.

For the solution that returned a result, here is the plot.

l7_248.png

l7_249.gif

NDSolveValue

NDSolveValue has the same relation to NDSolve as DSolveValue has to DSolve. It does not return a rewrite rule, it just returns an interpolating function.

l7_250.png

l7_251.gif

Note the difference in the plot.

l7_252.png

l7_253.gif

l7_254.png

l7_255.gif

l7_256.png

l7_257.gif

l7_258.png

l7_259.png

l7_260.png

l7_261.gif

Another Level of Classification

Many PDEs of interest are of higher than first-order. Let’s assume a function u(x,y) and a generic homogeneous second-order PDE is

l7_262.png

(7.70)

We choose a trial function f, that is twice differentiable, of the form

l7_263.png

(7.71)

We then find the relevant second-order partial derivatives for our PDE,

l7_264.png

(7.72)

Substituting this into (7.70) gives us

l7_265.png

(7.73)

There are three cases we can consider, f''=0, l7_266.png, or both.

If we look at the case f''=0, then we have the solution l7_267.png. This requires two arbitrary constants instead of two arbitrary functions and is thus not the most general solution. We may thus disregard this as a solution.

This leads us to consider the case l7_268.png, where we can choose a linear combination of the variables x and y and get a solution for any twice differentiable function f. This is a quadratic equation leading to the solutions

l7_269.png

(7.74)

The sign of the discriminant l7_270.png is extremely important in determining the number and type of solutions to the quadratic. There are three broad cases

l7_271.png

(7.75)

In the case l7_272.png, there are two distinct real solutions. PDEs of this type reduce to the form l7_273.png and are called hyperbolic.

In the case l7_274.png, there is a single degenerate root. PDEs of this type reduce to the form l7_275.png and are called parabolic.

In the case l7_276.png, there are two distinct complex roots. PDEs of this type are called elliptic.

Almost all of the PDEs used in theoretical physics fall into one of these categories. We shall cover each of these three kinds of equations in their own lessons later.

An Example of a Hyperbolic PDE

Here is an example of a second-order hyperbolic PDE in DSolve.

l7_277.png

l7_278.png

The solution surface can be  plotted, here we assume a=1.

l7_279.png

l7_280.png

l7_281.png

l7_282.gif

An Example of a Parabolic PDE

Here we have a parabolic type PDE.

l7_283.png

l7_284.png

This takes a while. The answer is in the form of a Fourier cosine series. We can get the first five terms.

l7_285.png

l7_286.png

l7_287.png

l7_288.png

We can plot this.

l7_289.png

l7_290.gif

An Example of an Elliptic PDE

Here we have an example of an elliptic PDE. Our PDE will be a Helmholtz-type equation on a rectangle. Here we introduce the Laplacian command and the UnitTriangle.

l7_291.png

l7_292.png

l7_293.png

l7_294.png

l7_295.png

l7_296.png

We again have an answer in the form of a Fourier series. We will take the first few terms and assume a=1.

l7_297.png

l7_298.png

Here is the plot of this function.

l7_299.png

l7_300.gif

What happens if we get more terms in the series?

l7_301.png

l7_302.png

l7_303.gif

Time-Dependent Problems

Up to now we have dealt primary with PDEs having only spatial derivatives. Many situations involve processes that evolve in time, what we now call a dynamical system. This allows us to consider problems of the form

l7_304.png

(7.76)

Such an equation is called a conservation law. Here is a good example of a time-dependent PDE. We can write the wave equation for some velocity v, and for some function u(x,t),

l7_305.png

(7.77)

So we have second derivatives instead of first derivatives. Is this still a conservation law? If we multiply by l7_306.png we get,

l7_307.png

(7.78)

We can factor this,

l7_308.png

(7.79)

Here we have two operators acting on the function. If we impose the condition that,

l7_309.png

(7.80)

this effectively destroys the second operator in (7.79), and (7.80) satisfies the wave equation in (7.77). This is a conservation law. Can we find a general solution to (7.77)?

l7_310.png

l7_311.png

This is a fairly well known result. This is very familiar as we have a solution of a constant function and is another example of the method of characteristics. We can impose a function form here and we assume the wave velocity to be 1.

l7_312.png

l7_313.png

Here we can plot the time evolution of this wave,.

l7_314.png

l7_315.gif

What's Going On Under the Hood?

The command NDSolve hides a lot of complications. In order to approximate the integration of a PDE we must first establish a grid, or mesh, covering the region we are concerned with. The vertices of this mesh (and it need not be rectilinear) form the cells of a matrix. This discretization of a spatial region into a matrix turns the problem of solving a PDE into a problem of solving a really large, and often complicated, matrix equation. The elements of this matrix equation are often created by some scheme such as finite differencing, like what we saw in Lesson 1.

Broadly speaking there are two kinds of  meshes. The first is a fixed mesh. Here the distance across one cell is the same throughout the solution of the PDE. The second is an adaptive mesh, and we will discuss that in a later lesson, when we talk about the finite element method. For now we can say that an adaptive mesh changes, or adapts, to the current situation.

It is possible to have a finer mesh for areas of particular interest to your study. That is called a nested grid. The danger of a nested grid is that you can propagate artifacts of instability outside of the finer grid.

The ability to make meshes is important to Mathematica, and we will discuss this in greater detail when we get to the Finite-Element method in a later lesson.

Created with the Wolfram Language