How it works
What you see above is not a video. Your graphics card is computing, right now, how hundreds of thousands of particles — the stars, gas and dark matter of two galaxies — pull on each other by Newton's law, and how the gas compresses and forms stars. Below: which formulas are at work, what is simplified, and how we checked that it all adds up.
Scales and units
Galaxies are huge and slow. It is convenient to compute in units where the gravitational constant $G = 1$, length is the kiloparsec (3262 light years) and mass is $10^{10}$ solar masses. The unit of speed is then 207 km/s and the unit of time 4.7 million years. In these units the Milky Way is a disk of radius about 15 and mass 5 inside a halo of mass 100; Andromeda is somewhat larger.
A particle is not a star. On the page a star particle weighs a few million solar masses, a dark-matter particle eight times more, a gas particle four times less: "clouds" of many stars moving together. The film has 80 times more particles, each that much lighter.
Gravity: everything pulls on everything
The acceleration of particle $i$ is the sum of the pulls of all the others:
$$\mathbf a_i = G \sum_{j \ne i} m_j\, g(r_{ij})\, (\mathbf x_j - \mathbf x_i),\qquad g(r) = \frac{1}{r^3}\ \text{far away}.$$Up close the force is softened: a particle is a smeared cloud, not a point, so two particles cannot fling each other to infinity. We use the cubic spline of GADGET-2 (Springel 2005): beyond $h = 2.8\,\varepsilon$ the force is exactly Newtonian, and at the centre the potential is that of a Plummer sphere of radius $\varepsilon$. On the page $\varepsilon$ is 80 pc for stars and 250 pc for dark matter; with more particles the softening shrinks as $N^{-1/3}$.
The Barnes–Hut tree
A direct sum over all pairs costs $N^2$: forty billion terms per step for 200 thousand particles. Barnes and Hut (1986) noticed that a distant group of particles can be replaced by one: its total mass at its centre of mass plus a correction for its shape, the quadrupole moment:
$$\Phi(\mathbf y) \approx -\frac{M}{|\mathbf y|} - \frac{\mathbf y^{\mathsf T} Q\, \mathbf y}{2|\mathbf y|^5},\qquad Q_{ab} = \sum_k m_k \big(3 d_a d_b - |\mathbf d|^2 \delta_{ab}\big).$$Space is split by an octree, each cube into eight, until a cell holds at most 16 particles. A node is taken whole when a group of particles is farther from its centre of mass than $l/\theta + \delta$ ($l$ the node's size, $\delta$ the offset of its centre of mass from the cell's centre, $\theta = 1$ on the page, $0.9$ in the film); otherwise it is opened. A particle then needs about a thousand interactions instead of two hundred thousand.
The tree is rebuilt on the GPU every step. Each coordinate becomes a 20-bit integer, and the bits of the three axes are interleaved into a 60-bit Morton key. A radix sort orders the particles by key. After that, the particles of any cell sit next to each other, and its children are found by binary search. The walk needs no stack: each node knows "where next", and the 64 particles of a group walk the tree together, taking the same decisions.
Time: leapfrog
Positions and velocities advance by kick–drift–kick:
$$\mathbf v \mathrel{+}= \mathbf a\,\tfrac{\Delta t}{2},\qquad \mathbf x \mathrel{+}= \mathbf v\,\Delta t,\qquad \mathbf a = \mathbf a(\mathbf x),\qquad \mathbf v \mathrel{+}= \mathbf a\,\tfrac{\Delta t}{2}.$$The scheme is symplectic and time-reversible: the energy error does not accumulate but wobbles about zero, so orbits do not creep apart over thousands of steps. The step on the page is half a million years.
Building a galaxy
A galaxy has to start in equilibrium, or it would pulsate before it ever meets the other. Each is made of five parts:
- a dark-matter halo and a bulge with Hernquist's profile $\rho = \dfrac{M a}{2\pi r (r + a)^3}$; the halo's scale follows from its mass $M_{200}$ and a concentration $c = 10$, as for halos in cosmological simulations;
- an exponential stellar disk: $\Sigma(R) = \Sigma_0 e^{-R/R_d}$, $\operatorname{sech}^2(z/z_0)$ in height;
- a gas disk, twice as wide and thinner;
- a supermassive black hole in the middle: 4.3 million solar masses for the Milky Way, 140 million for Andromeda.
The halo and bulge particles get their speeds from Eddington's formula. For a spherical system with isotropic velocities it recovers the distribution over energy $\mathcal E = \Psi - v^2/2$ from the density:
$$f(\mathcal E) = \frac{1}{\sqrt 8\,\pi^2} \frac{d}{d\mathcal E} \int_0^{\mathcal E} \frac{d\rho}{d\Psi} \frac{d\Psi}{\sqrt{\mathcal E - \Psi}}.$$The potential $\Psi$ is the shared one of halo, bulge, disks and hole together. As a check, for a lone Hernquist sphere the numerical $f(\mathcal E)$ matches the known exact formula (Hernquist 1990) to a per cent.
The disk rotates as the Jeans equations say (Hernquist 1993). The vertical dispersion comes from the balance of a layer, $\sigma_z^2 = \pi G \Sigma z_0$. The radial one is set by Toomre's stability parameter
$$Q = \frac{\sigma_R\, \kappa}{3.36\, G \Sigma} = 1.3,$$with $\kappa$ the epicyclic frequency. The stars rotate a little slower than the circular speed $v_c$: this is the asymmetric drift,
$$\bar v_\phi^2 = v_c^2 + \sigma_R^2 \Big(1 - \frac{\kappa^2}{4\Omega^2} - \frac{2R}{R_d}\Big).$$The circular speed of a thin exponential disk is a combination of Bessel functions (Freeman 1970): $v_c^2 = 4\pi G \Sigma_0 R_d\, y^2 [I_0(y)K_0(y) - I_1(y)K_1(y)]$, $y = R/2R_d$. Our Milky Way turns at 230 km/s at the Sun's distance, like the real one. Left alone, such a disk grows spiral arms and a bar within a couple of hundred million years — and the real Milky Way is indeed a barred galaxy.
Where Andromeda is
Andromeda is 780 kpc (2.5 million light years) away today. Its speed towards us along the line of sight is known precisely: −301 km/s relative to the Sun, −109 km/s relative to the Galactic centre. Its sideways speed is harder: it is measured from the drift of Andromeda's stars across the sky over years. Hubble gave 17 km/s (van der Marel et al. 2012), Gaia 57 km/s (2019) and 80 km/s (2021), with errors of tens of km/s. The default is 17 km/s; the settings let you choose another.
The disks are oriented as the real ones. The Milky Way turns clockwise seen from the north Galactic pole. Andromeda's axis follows from its inclination ($i = 77.5^\circ$) and position angle ($37.7^\circ$), given that the north-west side of its disk is the near one and its north-east half recedes from us. That gives $(l, b) \approx (241^\circ, -30^\circ)$, as found by van der Marel.
The first three billion years, while the galaxies are far apart, are computed as a two-body problem with extended halos and dynamical friction. From 300 kpc on, the full simulation takes over.
Whether they will really collide is not known. Sawala and colleagues (2025) took the measurement errors and the pull of M33 and the Large Magellanic Cloud into account and found about a 50% chance of a merger within 10 billion years. This page shows a case in which the merger happens.
Dynamical friction
Why do the galaxies come back after flying apart from the first pass? A massive body moving through a cloud of light particles gathers a wake behind it — a crowd that pulls it back. Chandrasekhar (1943) derived the drag for a uniform medium:
$$\frac{d\mathbf v}{dt} = -\frac{4\pi G^2 M \rho \ln\Lambda}{v^3}\Big[\operatorname{erf}(X) - \frac{2X}{\sqrt\pi} e^{-X^2}\Big]\mathbf v,\qquad X = \frac{v}{\sqrt2\,\sigma}.$$In the full simulation nobody puts this formula in: the drag appears by itself, because the halo particles pull on each other and on the galaxies. We only use the formula on the far approach, where we compute two points.
Tidal tails
Stars almost never collide: the distances between them are tens of millions of times their sizes. The tails come from the tide: the near edge of a disk is pulled towards the other galaxy harder than its centre, the far edge more weakly. The stars stretched furthest are those whose rotation matches the direction of the pass: they ride along with the other galaxy the longest. Alar and Juri Toomre explained this in 1972 with test particles around two point masses on one of the early computers. Their models of the Antennae and the Mice are among the presets.
Gas: smoothed particle hydrodynamics
Unlike stars, gas has pressure and collides. We compute it with SPH (Lucy 1977; Gingold and Monaghan 1977): a gas particle is a smeared lump, and the density is the sum of its neighbours' contributions with the Wendland C2 kernel:
$$\rho_i = \sum_j m_j W(r_{ij}, h_i),\qquad W(r,h) = \frac{21}{2\pi h^3}\Big(1 - \frac rh\Big)^4\Big(1 + \frac{4r}{h}\Big).$$The smoothing length $h$ adjusts to give about 48 neighbours. The gas is isothermal: a temperature of about $10^4$ K (kept by starlight), pressure $P = c_s^2 \rho$, $c_s = 10$ km/s. The pressure force is symmetric in pairs, so momentum is conserved exactly:
$$\frac{d\mathbf v_i}{dt} = -\sum_j m_j \Big(\frac{P_i}{\rho_i^2} \nabla_i W(h_i) + \frac{P_j}{\rho_j^2} \nabla_i W(h_j) + \Pi_{ij}\, \nabla_i \bar W\Big).$$$\Pi_{ij}$ is Monaghan's artificial viscosity: it acts only on approaching pairs and turns a head-on clash of streams into a shock. Balsara's switch, which compares $|\nabla\cdot\mathbf v|$ with $|\nabla\times\mathbf v|$, keeps it from braking the disk's ordinary rotation.
Gas needs shorter steps than stars — the Courant condition $\Delta t < 0.25\, h / v_{\text{sig}}$ — so within one gravity step the gas takes up to eight substeps. Two more precautions: a Jeans pressure floor $c_{\text{eff}}^2 = \max(c_s^2,\ 3G\rho h^2)$ keeps the gas from collapsing into clumps smaller than the resolution, and where the Courant condition still pinches, the push of a substep is capped at half the signal speed.
Stars from gas
Where the gas is denser than about 0.1 hydrogen atoms per cubic centimetre and compressing, stars are born at the Schmidt–Kennicutt rate: one and a half per cent of the gas per free-fall time $t_{ff} = \sqrt{3\pi/32G\rho}$:
$$\dot\rho_\star = \varepsilon\, \frac{\rho}{t_{ff}},\qquad p = 1 - e^{-\varepsilon \Delta t / t_{ff}}.$$Each gas particle turns into a star particle with probability $p$ and remembers its birth time. The counter on screen shows how much mass a year becomes stars: 1.5–2 solar masses a year in the quiet Milky Way, like the real one.
How the picture is made
Starlight. A star particle is a population of one age. Its colour and light per unit mass follow the models of Bruzual and Charlot (2003): a young population is blue and hundreds of times brighter than an old, yellow-orange one. Between 10 million and 10 billion years the light fades roughly as $t^{-0.8}$. Stars younger than 10 million years ionise the gas around them, and a pink H II region glows (the hydrogen lines Hα and Hβ). Its brightness falls off smoothly from the middle, along a Moffat profile $(1 + r^2/a^2)^{-2}$: a bright knot a few tens of parsecs across and a halo that fades without an edge.
Dust. Dust travels with the gas, about a per cent of its mass. It absorbs the light behind it, blue more than red, by the law of Cardelli, Clayton and Mathis (1989), normalised to $N_H/E(B{-}V) = 5.8 \cdot 10^{21}$ cm⁻² (Bohlin et al. 1978). That is why dust lanes are dark brown and the light behind them reddens.
Putting a frame together. Each particle is drawn as a soft disc, the projection of its smoothing kernel. The particles are sorted by distance from the camera and laid down from far to near: stars add light, gas dims everything behind it, colour by colour, $C \leftarrow E + C\,e^{-\tau}$. Part of the starlight is drawn as a sharp point: real galaxies are not smooth, their light comes from bright giants and clusters.
Exposure and colour. The exposure adjusts so that the brightest 0.1% of pixels just touch white. Brightness is stretched by the $\operatorname{asinh}$ function, as in the colour images of the SDSS survey (Lupton et al. 2004): linear for faint light, logarithmic for bright, with no shift of colour. The brightest light rolls smoothly into white, with no flat clipping. Light brighter than white spreads into a halo, as in a telescope's optics, fading with distance roughly as $1/r^2$. So bright knots and cores glow while the disk around them stays sharp. Local contrast is lifted slightly, as in processing astrophotos. The colours are "camera" colours — light split into B, V and I filters, as in Hubble's famous images. To the eye the galaxy would look paler and less colourful.
The backdrop. Behind everything are distant galaxies. The twenty-odd brightest neighbours (M81 and M82, NGC 253, Centaurus A, M101, the Sombrero, the Virgo cluster) are in their true places on the sky at their true sizes. The rest is a deep field invented from the real statistics of galaxy counts, gathered into groups and clusters. It is scenery, not a catalogue.
What is simplified
- A particle is not a star but a cloud of millions of suns on the page and tens of thousands in the film. Small details — single clusters, thin dust filaments — are blurred to the softening scale.
- The gas is isothermal. No cooling below $10^4$ K, no shock heating to millions of degrees, no hot gas in the halo. In a real merger part of the gas would heat up and shut star formation down for a long time.
- No feedback: supernovae and the winds of young stars do not stir the gas.
- The black holes are just heavy particles. They sink to the middle and merge on their own, but we do not compute gas accretion or a quasar's light.
- No surroundings. No M33, no Magellanic Clouds or other satellites, no cosmological background: the Local Group is alone in the void.
- One step for all (apart from the gas substeps). In the densest cores orbits are computed more coarsely; real codes give every particle its own step.
- The population colours are for solar metallicity, and the dust is the same everywhere.
How it was checked
- Eddington's distribution function matches Hernquist's exact formula to 1% (5% at the very centre); halo and bulge hold the virial balance $2K = |W|$ in their potential to 3%.
- The circular speed of Freeman's disk agrees with a direct sum over the disk (0.2%) and with the derivative of its potential ($10^{-5}$).
- Andromeda's position and orientation agree with van der Marel's: $(l, b) = (121.17^\circ, -21.57^\circ)$, disk axis $(241^\circ, -30^\circ)$.
- The GPU tree's forces are compared with a double-precision direct sum: median error $2 \cdot 10^{-4}$, 99% of particles better than $10^{-3}$ at $\theta = 0.6$; $8 \cdot 10^{-4}$ and $4 \cdot 10^{-3}$ at $\theta = 1$.
- The SPH density on the GPU matches a direct count to $10^{-7}$; the total momentum of the hydrodynamic forces is zero to $10^{-9}$.
- An isolated Milky Way keeps its energy to $6 \cdot 10^{-5}$ over a billion years, the whole collision to $2 \cdot 10^{-4}$. The first close pass comes 3.9 billion years from now, as in van der Marel's calculations (3.87).
Sources
Barnes J., Hut P. (1986) A hierarchical O(N log N) force-calculation algorithm. Nature 324.
Bédorf J., Gaburov E., Portegies Zwart S. (2012) A sparse octree gravitational N-body code that runs entirely on the GPU. J. Comput. Phys. 231.
Springel V. (2005) The cosmological simulation code GADGET-2. MNRAS 364.
Hernquist L. (1990) An analytical model for spherical galaxies and bulges. ApJ 356; (1993) N-body realizations of compound galaxies. ApJS 86.
Freeman K. C. (1970) On the disks of spiral and S0 galaxies. ApJ 160.
Toomre A., Toomre J. (1972) Galactic bridges and tails. ApJ 178.
Chandrasekhar S. (1943) Dynamical friction. ApJ 97.
van der Marel R. P. et al. (2012) The M31 velocity vector I–III. ApJ 753; (2019) ApJ 872.
Sawala T. et al. (2025) No certainty of a Milky Way–Andromeda collision. Nature Astronomy.
Monaghan J. J. (1992) Smoothed particle hydrodynamics. ARA&A 30; Balsara D. (1995) J. Comput. Phys. 121.
Springel V., Hernquist L. (2003) Cosmological SPH simulations: a hybrid multiphase model for star formation. MNRAS 339.
Bruzual G., Charlot S. (2003) Stellar population synthesis at the resolution of 2003. MNRAS 344.
Cardelli J., Clayton G., Mathis J. (1989) The relationship between infrared, optical, and ultraviolet extinction. ApJ 345.
Lupton R. et al. (2004) Preparing red-green-blue images from CCD data. PASP 116.