Part One: What Is Intensity Interferometry#
From ordinary interferometry to intensity interferometry: why can two telescopes measure the size of a star? Let us start from the most basic question: what is light, and what does a detector actually measure? Light is fundamentally an electromagnetic wave. For monochromatic light, the electric field can be written as
\(E(t)\) is the electric field oscillating in time; \(E_0\) is the amplitude; \(\omega\) is the angular frequency; \(t\) is time; \(\phi\) is the phase. This expression says that light is first of all a wave carrying amplitude and phase, not, from the outset, a count inside a detector.
where \(E_0\) is the amplitude, \(\omega=2\pi\nu\) is the angular frequency, \(\nu\) is the frequency, and \(\phi\) is the phase. What truly oscillates in time and propagates is the electric field \(E\).
More generally, the electric field is often written in complex form:
where \(\mathcal E\) is the complex amplitude, containing both amplitude and phase information:
Detectors, such as CCDs, PMTs, and APDs, cannot directly measure this rapidly oscillating electric field. The reason is simple: the frequency of visible light is roughly
corresponding to a period
Ordinary electronics fall far short of keeping up with such an oscillation speed. Therefore what the detector actually measures is not \(E(t)\) itself, but the time-averaged intensity. For an electromagnetic wave, the intensity is proportional to the time average of the square of the electric field:
If
then
So
In complex-amplitude notation, the intensity is often written as
\(I\) is the intensity recorded by the detector; \(\mathcal E\) is the complex amplitude; \(|\mathcal E|^2\) is the squared amplitude. This expression connects “the language of waves” with “the language of detectors”: the phase is hidden inside the complex amplitude, but an ordinary detector mainly responds to energy.
This is the single most important point in understanding intensity interferometry: the fundamental variable of light is the electric field, but what the detector measures is the intensity.
Ordinary interferometry starts precisely from the superposition of electric fields. Suppose there are two beams of the same frequency:
After the two beams superpose,
What the detector measures is
Expanding gives
where
The cross term is
Using the identity
we obtain
Therefore
that is,
If we define the intensity of a single beam as
then the total intensity is
\(I_0\) is the intensity of a single beam; \(\phi\) is the phase difference between the two beams; \(\cos\phi\) controls constructive or destructive interference. This result shows that ordinary interferometry measures the electric-field cross term: the phase difference is not seen directly, yet it manifests itself through the bright and dark fringes of the intensity.
Therefore, as the phase difference \(\phi\) varies, the detector sees bright fringes and dark fringes. This is ordinary interferometry. What truly does the work here is the cross term
In other words, ordinary interferometry is essentially comparing two electric fields, and so it is also called amplitude interferometry.
The problem is that starlight is not laser light. The phase of a laser can remain stable over a relatively long time, so it can form stable interference fringes. But the surface of a star has an enormous number of atoms and plasma elements, each of which radiates randomly. The total electric field can be written as
Each term can be approximately written as
Because the phase \(\phi_i(t)\) of each radiating unit varies randomly, the phase of the total electric field also changes constantly. After waiting a somewhat longer time,
This does not mean that the star does not shine, but rather that the electric field, as a quantity carrying phase, cancels itself out after its positive and negative oscillations and random phases are averaged. The detector can still measure a nonzero intensity, because
Although the phase of starlight is random, it does not vary randomly infinitely fast. Over a very short time, the electric field still maintains a certain correlation. This time scale over which correlation is maintained is called the coherence time, denoted
More precisely, the coherence time can be defined through the temporal coherence function. For the electric field at the same location, define
When \(|\tau|\ll \tau_c\),
When \(|\tau|\gg \tau_c\),
So \(\tau_c\) represents the time scale over which the electric field still retains its phase memory. If the frequency bandwidth of the light is \(\Delta \nu\), there is usually the order-of-magnitude relation
The corresponding coherence length is
Now consider two telescopes. One measures \(E_1(t)\), the other measures \(E_2(t)\). If the two telescopes are very close together, they receive almost the same wavefront, so
To describe how similar the electric fields at the two locations are, we define the first-order spatial coherence function
If we also account for a time delay \(\tau\), it is written as
\(g^{(1)}_{12}(\tau)\) compares the two electric fields at the two telescopes separated by \(\tau\); the numerator is the electric-field correlation, and the denominator normalizes by each intensity. Its modulus close to 1 means the two electric fields still share a common phase memory, while close to 0 means the phase relationship has already been lost.
The physical meaning of this quantity is very direct: it compares whether the electric fields at two locations and two instants are correlated. If
the two electric fields are fully correlated; if
the two electric fields have no correlation. What ordinary amplitude interferometry measures directly is exactly \(g^{(1)}_{12}\).
But ordinary optical interferometry is very hard to do, because it requires actually merging the two beams and requires the optical-path error to be smaller than the wavelength. For visible light,
This means the entire optical system must be stable to the level of tens of nanometers. For telescopes separated by a large distance, this is extremely difficult.
The idea of Hanbury Brown and Twiss begins precisely here. Since the detector cannot measure the electric field \(E\) in the first place and can only measure the intensity \(I\), might one avoid comparing \(E_1(t)\) and \(E_2(t)\), and instead directly compare the intensities measured by the two telescopes,
This is intensity interferometry.
On the surface this seems strange. The intensity is already \(|E|^2\), which appears to have averaged away the phase information, so why would there still be interference? The key is that although the mean intensity of thermal light is approximately constant, the instantaneous intensity is not strictly constant, but has small random fluctuations. It can be written as
\(I(t)\) is the instantaneous intensity; \(\langle I\rangle\) is the mean intensity; \(\Delta I(t)\) is the fluctuation about the mean. Intensity interferometry does not merge electric fields, but compares whether these small fluctuations at the two telescopes are synchronized.
where
Therefore
These fluctuations come from the random-phase superposition of thermal light. One can picture the electric field contributed by each atom or radiating unit as a small vector in the complex plane. At a given instant, if many of the small vectors point in nearly the same direction, the total electric field is large, and therefore
is also large. At another instant, if these small vectors point in disordered directions and cancel one another, the total electric field is small, and the intensity weakens. Thus thermal light naturally exhibits random flickering, that is, intensity fluctuations.
If the two telescopes receive light from the same coherence region, then these intensity fluctuations are not completely independent. At some instant, when the radiation from a certain part of the star happens to strengthen, both telescopes will see the increase; at the next instant, when it dims, both telescopes will also see the decrease. Therefore
What intensity interferometry measures is precisely this correlation between the intensity fluctuations, not the mean intensity itself.
If the distance between the two telescopes gradually increases, the wavefronts they receive are no longer completely identical, and the coherence regions they see also gradually differ. Then the bright-dark variations seen by one telescope are not necessarily seen synchronously by the other. The correlation gradually decreases, finally approaching zero:
Therefore, the degree of correlation of the intensity fluctuations tells us the spatial coherence scale of the light source.
This immediately connects to the size of the star. If the star is very small, approximately a point source, then the wavefront reaching the Earth is still coherent over a large range. Even if the two telescopes are separated by a relatively large distance, they can still see similar fluctuations, and the correlation decreases relatively slowly.
If the angular diameter of the star is relatively large, different regions radiate independently of one another. When the two telescopes are separated by a large distance, their phase weighting of different regions of the star differs, so the intensity-fluctuation correlation decreases more quickly. Thus the larger the star, the faster the coherence decreases with baseline. By measuring the correlation at different baselines \(B\), one can infer the angular diameter of the star.
Now this idea can be put into mathematical form. Define
where
The product of the intensities measured by the two telescopes is
Expanding gives
After averaging, since
the two middle terms vanish, and therefore
We then define the normalized second-order coherence function
\(g^{(2)}_{12}(0)\) is the normalized intensity product at zero delay; \(I_1,I_2\) are the intensities at the two telescopes; the denominator removes the effect of the mean brightness. If it is greater than 1, it means the intensity fluctuations of the two paths have excess synchrony.
Substituting into the above gives
If we consider a time delay \(\tau\), then
Correspondingly,
So the second-order coherence function is just the normalized intensity-fluctuation correlation.
The key question that follows is: why is this quantity related to the square of the first-order coherence function? That is, why does thermal light satisfy
\(g^{(2)}\) is the intensity correlation; \(g^{(1)}\) is the electric-field correlation; the squared modulus means only the magnitude of the first-order coherence is retained. This question is at the core of HBT: why, without measuring the electric-field phase, can the squared modulus of the first-order coherence still be read out from the statistics of thermal light.
This is not a coincidence, but is because thermal light can be well approximated as a complex Gaussian random field. For a Gaussian random variable, the fourth-order moment can be expressed in terms of the second-order moments; this is Isserlis’ theorem, which is the Wick theorem for classical random variables.
Let the electric fields at the two locations be \(E_1\) and \(E_2\), with intensities
Then
For a zero-mean complex Gaussian random field,
Since
the second term can be written as
Therefore
Dividing both sides by
gives
By the definition of the first-order coherence function,
so
Finally we obtain the Siegert relation:
On the left, \(g^{(2)}_{12}\) is the quantity measured directly by intensity interferometry; on the right, the 1 is the baseline in the absence of correlation, and \(|g^{(1)}_{12}|^2\) is the squared modulus of the electric-field coherence. It shows that the intensity correlation of thermal light retains the spatial-coherence information.
If one accounts for the finite number of polarizations, the detection time resolution, and the finite bandwidth, the correlation amplitude actually observed is diluted, and is often written as
where \(0<\beta\leq 1\). In the ideal case of a single polarization and sufficiently high time resolution, \(\beta=1\). In real intensity-interferometry experiments, \(\beta\) is affected by the detector time resolution, the spectral bandwidth, and the polarization selection.
This shows that although intensity interferometry does not directly measure the electric field, because of the Gaussian statistical nature of thermal light, the information of the electric-field correlation is still preserved in the correlation function of the intensity fluctuations. This is the theoretical core of intensity interferometry.
Now there is one last astronomical question to answer: why does measuring the correlation between different telescopes tell us how large a star is?
Suppose a star is very far away, for example
When the spherical wave emitted by the star propagates to near the Earth, since
it can be locally approximated as a plane wave. The wavefront is the surface composed of positions with the same phase. For a point source, the entire wavefront is almost uniform, so the spatial coherence is very high.
A real star is not a point source, but a disk with an angular diameter. One can divide the surface of the star into many small area elements, each of which is approximately treated as a point source. Different area elements lie in different directions in the sky, so when reaching the two telescopes they produce different path differences.
Let the baseline between the two telescopes be \(\mathbf B\), and let the direction of a certain area element be the unit vector \(\mathbf s\). The path difference is
The corresponding phase difference is
Here \(\lambda\) is the wavelength, \(\mathbf B\cdot\mathbf s\) is the path difference, and \(2\pi/\lambda\) converts the path difference into a phase difference.
If the brightness of the star in direction \(\mathbf s\) is \(I(\mathbf s)\), then adding up the contributions of all area elements gives
\(\mathbf B\) is the telescope baseline; \(I(\mathbf s)\) is the brightness in sky direction \(\mathbf s\); the exponential factor gives the phase difference of that direction between the two telescopes; the denominator normalizes by the total brightness. It turns the sky brightness distribution into the first-order coherence on the baseline.
The physical meaning of this formula is not complicated. Each area element contributes an electric-field amplitude, but because the directions differ, it carries a different phase factor between the two telescopes,
Adding up the contributions of all directions gives the first-order coherence function between the two telescopes.
In the small-field approximation, the sky direction can be expressed by angular coordinates \((\alpha,\delta)\). If the projection of the baseline on the sky plane is \((B_x,B_y)\), define
Then the Van Cittert–Zernike theorem can be written as
This shows that the first-order spatial coherence function of a far-field source equals the normalized Fourier transform of the source’s brightness distribution:
This normalized Fourier transform is usually called the complex visibility:
Changing the baseline length and direction is equivalent to sampling different positions in Fourier space. As the Earth rotates, the projection of the baseline relative to the sky also changes, so the sampling points move across the \(u,v\) plane; this is the basic idea of Earth-rotation synthesis.
For a star that is a uniform-brightness disk, the brightness distribution can be written as
where \(\theta\) is the angular diameter of the star and \(\rho\) is the angular distance from the center of the disk.
Because the disk is axisymmetric, the complex visibility depends only on the baseline length
The Fourier transform can be computed analytically, giving
\(V(B)\) is the visibility at baseline length \(B\); \(J_1\) is the first-order Bessel function; \(x\) combines the angular diameter, baseline, and wavelength into a dimensionless parameter. The larger the disk, the faster \(V(B)\) decreases with baseline.
where
Here \(J_1\) is the first-order Bessel function of the first kind. Therefore
This function decreases as the baseline increases, and becomes zero at certain baselines. The first zero satisfies
where the first zero is
Therefore
Solving gives
Since
we have
Conversely, if the first null baseline \(B_{\rm null}\) is observed, the angular diameter can be estimated:
This result has the same mathematical origin as the position of the first dark ring of the Airy pattern in circular-aperture diffraction, because both come from the Fourier transform of a disk.
Ordinary amplitude interferometry directly measures
so it can obtain both the amplitude and the phase of the Fourier transform. HBT intensity interferometry measures
Since
we have
In real observations this is often written as
\(g^{(2)}(B)-1\) is the HBT signal after subtracting the uncorrelated baseline; \(\beta\) is the dilution factor from finite time resolution, bandwidth, and polarization; \(|V(B)|^2\) is the squared modulus of the Fourier transform of the source brightness. This is the direct observational expression for measuring stellar size with intensity correlation.
That is to say, what intensity interferometry directly obtains is
It retains the squared modulus of the Fourier transform, but does not directly give the phase. Therefore intensity interferometry is very well suited to measuring model parameters such as stellar angular diameters and binary separations; but if one wants to directly recover a two-dimensional image, it is more difficult than amplitude interferometry, and requires additional information such as model fitting, closure phase, or higher-order correlations.
At this point, the complete logical chain is as follows.
The fundamental variable of light is the electric field \(E\), but what the detector measures is the intensity
Ordinary interferometry measures the first-order coherence function by comparing the electric fields at two locations,
and therefore requires maintaining optical-path stability and phase information. The electric-field phase of thermal light is random, but the intensity has small fluctuations. When the two telescopes receive light from the same coherence region, these fluctuations are correlated. What intensity interferometry measures is the normalized intensity-fluctuation correlation, that is, the second-order coherence function
For thermal light, since the electric field can be approximated as a complex Gaussian random field, the Siegert relation gives
And by the Van Cittert–Zernike theorem,
which is the normalized Fourier transform of the source brightness distribution. Therefore HBT intensity interferometry measures
that is, the squared modulus of the Fourier transform of the source brightness distribution. By changing the baseline length and direction, one can sample different spatial frequencies, and thereby infer the angular diameter and structure of the star.