Holding a Star: Fusion as a Control Problem
Tokamak plasmas are engineered to be unstable, and their instabilities unfold a thousand times faster than a human reflex. This essay argues that the real bottleneck of fusion energy is not plasma physics but the control loop, and looks at what happens when that loop is handed to a neural network.
There is a room in every tokamak facility where the operators sit, and during a plasma shot the striking thing about that room is what nobody does. Nobody steers. A hundred million degrees sits a meter or two from steel that would fail at a few thousand, held in place by magnetic fields, and for the seconds or minutes the plasma exists, no human hand adjusts anything that matters. This is not a policy choice or a union rule. It is arithmetic. When a tokamak plasma loses confinement, the thermal energy can be gone in under a millisecond. A trained human needs around two hundred milliseconds just to react to a light turning green. By the time a person perceives that something has gone wrong inside the vessel, the event has been over for two decades of time, on the plasma’s clock.
Fusion is usually told as a physics story: temperatures hotter than the sun’s core, exotic isotopes, the long wait for net energy. But the physics of why deuterium and tritium fuse has been settled for the better part of a century. What has kept fusion out of the grid is closer to an engineering confession: we are trying to operate a machine whose failure modes are faster than our nervous system, and until recently we did not have controllers worthy of it.
A bottle made of nothing
The core trick of a tokamak is that a charged particle cannot easily cross a magnetic field line; it spirals along it. Bend the field lines into closed rings, nest those rings into a doughnut, and the particles are trapped on invisible rails. The plasma touches nothing. The bottle is made of field.
The field has to fight the plasma’s own pressure, and the currency of that fight is magnetic pressure, B squared over two mu zero. The formula just says that the stiffness of the bottle grows with the square of the field strength, which is why fusion companies obsess over magnets: doubling the field buys four times the push. At the 20 tesla that the strongest new superconducting magnets reach, magnetic pressure is on the order of 1,600 atmospheres, roughly the crush at the bottom of the Mariana Trench, exerted by empty space on a wisp of gas thinner than a laboratory vacuum.
So far this sounds static, a bottle that simply holds. It is not. A tokamak equilibrium is a balance actively maintained, more standing on a ball than sitting in a chair. And here is the part that surprises most people: modern tokamaks are deliberately built to be unstable.
A plasma with a circular cross-section is well behaved but mediocre. Stretch it vertically into a D shape and everything improves: more current fits, pressure rises, confinement gets better. The price is that an elongated plasma is vertically unstable. Nudge it up and the fields pull it further up; nudge it down and it accelerates down. Left alone, it would slam into the vessel in fractions of a millisecond. The conducting metal wall slows this runaway to a few milliseconds through induced currents, and active feedback coils do the rest, sensing the drift and pushing back, thousands of times per second, for as long as the plasma lives.
Aircraft engineers know this bargain as relaxed stability. A modern fighter jet is designed unstable because instability buys agility, and a flight computer keeps it in the air, since no pilot’s hands are fast enough. A tokamak makes the same wager with higher stakes. When vertical control is lost, the result is a vertical displacement event: the plasma column drifts, touches the wall, and disrupts. The stored heat leaves in the thermal quench, typically under a millisecond, and the megaamperes of plasma current collapse over one to ten milliseconds, dumping enormous electromagnetic forces into the surrounding structure. On a research machine this scorches tiles. On a reactor-class machine it bends things that took years to build. Studies of ITER, the international flagship, are blunt: a limited number of full-scale uncontrolled disruptions could damage the device beyond practical repair. Control, in fusion, is not a performance feature. It is the condition of the machine’s survival.
The classical answer, and its ceiling
For forty years the answer has been a pipeline. Hundreds of magnetic sensors around the vessel measure fluxes and fields. A reconstruction code solves an inverse problem to estimate, in real time, where the plasma boundary is. A shape controller compares that boundary to the requested one, and a stack of feedback loops translates the error into voltages on a dozen or more coils. Every piece is hand-designed, hand-tuned, and provably stable within its assumptions. It works; it is why tokamaks run at all.
But the pipeline has a ceiling, and it is the cost of change. Each new plasma shape means months of control engineering: new reconstruction settings, new loop tunings, new commissioning shots on a machine whose experimental time is priced like telescope nights. The controller is not a small part of the research program. In practice it gates the research program.
In 2022, a team from EPFL’s Swiss Plasma Center and DeepMind published a different answer in Nature. They trained a single neural network, by deep reinforcement learning in simulation, to command all nineteen control coils of the TCV tokamak directly from sensor readings, replacing the reconstruction-plus-loops cascade with one learned mapping. Moved onto the physical machine, it held and shaped real plasmas: conventional shapes, negative triangularity, an exotic snowflake divertor, and, in one experiment, two separate plasma droplets sustained in the vessel at once. Asking for a new shape stopped being a months-long engineering campaign and became, to a first approximation, a new line in the objective the agent was trained on.
The idea has since crossed the Atlantic and moved up in machine class. Work reported in 2026 on DIII-D, the largest tokamak in the United States, demonstrated a reinforcement-learning magnetic controller that maps raw magnetic diagnostics straight to actuator commands, standing in for the isoflux algorithm that has anchored shape control for decades. In parallel, a DIII-D collaboration supported by the US Department of Energy has used deep reinforcement learning against a different enemy, the tearing instability, teaching a controller to forecast the approach of the instability from real-time monitoring and steer the plasma around it before it locks and triggers a disruption. That is a qualitative shift: not reacting faster than a human, but acting before the event, in a regime where reaction, at any speed, is already too late.
Where this stands in 2026
The machine that will test all of this at reactor scale is taking shape outside Boston. SPARC, built by Commonwealth Fusion Systems, is the compact high-field tokamak designed around those 20 tesla magnets, and by company accounts it was roughly 80 percent assembled as of July 2026, with all eighteen magnets, each around 24 tons, expected in place by the end of the summer. First plasma and the attempt at scientific breakeven, more fusion power out of the plasma than heating power in, are now targeted for 2027, a slip from the 2026 date the company long advertised. The money keeps arriving regardless: a round announced on 30 July 2026 brought total funding to about 4 billion dollars, reportedly close to a third of all private fusion investment worldwide. Those figures come from the company and its investors and should be read as such, but the direction is unambiguous. Serious capital is betting that the bottle will hold.
The slip is worth pausing on, because it is the pattern of the whole field, and it is rarely a physics slip. Magnets, assembly tolerances, control commissioning: fusion timelines die in engineering, in the thousand-step choreography needed before anyone gets to do plasma physics at all. The harder the machine leans on performance, the more it leans on the controller, and SPARC leans harder than anything built before it.
What breaks, and what is not yet proven
Honesty requires drawing the boundaries of the result. The learned controllers demonstrated so far are shape and position controllers on research machines, not autopilots for burning plasmas. TCV is a small, flexible, forgiving device; the Nature result, striking as it is, tracked its targets with accuracy that dedicated classical controllers can still match or beat after tuning. Reinforcement learning’s real advantage to date is flexibility, not raw precision.
The deeper problem is that these agents learn in simulators, and plasma simulators are incomplete by necessity. Turbulence, edge physics, and the plasma’s interaction with the wall are only partially captured, so every learned controller carries a sim-to-real gap it has never been fully tested across. Worse, each tokamak is its own dialect: a controller trained for one machine transfers to no other, and there is no shared corpus of plasma experience remotely comparable to what made modern machine learning work elsewhere. Data in this field is priced in machine-days.
And then there is trust. A PID loop can be certified: gain margins, phase margins, proofs within a model. A neural network holding megajoules a meter from the first wall cannot be certified that way, only tested, bounded, and wrapped in supervisory layers that must themselves be fast enough to matter. Prediction is not avoidance, avoidance is not a guarantee, and no learned controller has yet faced the environment that matters most: a self-heated, burning plasma, which no machine on Earth has sustained. Anyone claiming the control problem is solved is selling something.
The reflexes we write
Still, the direction of travel seems settled, and it says something larger than fusion. A fusion power plant will be the first major energy technology that cannot, even in principle, be operated by hand. Not supervised loosely, not overridden in emergencies by a quick-thinking operator: the physics fixes the timescales, and the timescales exclude us. Human control does not disappear in such a machine; it migrates upward, from the hand to the intention. Engineers will choose objectives, define the envelope of allowed states, decide what the controller should value and fear, and then stand outside a loop that closes a thousand times a second without them.
We have made this bargain before, in aircraft, in power grids, in anti-lock brakes, but always with the comforting fiction of a manual mode somewhere below the automation. Fusion removes the fiction. If we get our star, we will not fly it. We will write its reflexes, hand them to the machine, and watch, two hundred milliseconds behind, as it holds.
Further reading
- Magnetic control of tokamak plasmas through deep reinforcement learning, the Nature paper by Degrave et al. on the TCV learned controller.
- AI tackles disruptive tearing instability in fusion plasma, the US Department of Energy summary of the DIII-D instability-avoidance work.
- Accelerating fusion science through learned plasma control, DeepMind’s account of the EPFL collaboration.
- SPARC: proving commercial fusion energy is possible, Commonwealth Fusion Systems’ technical page on the machine now being assembled.