The Therac-25 Disaster Explained: How A Software Bug Gave Patients Massive Radiation Overdoses
The Machine That Could Not Fail
The Radiation Machine That Changed Software Safety Forever
A cancer patient lies alone inside a radiation treatment room. Outside, an operator enters a prescription into a computer, checks the screen and starts the machine. The console says that almost no radiation has been delivered. The patient, however, has just received an enormous overdose.
Between June 1985 and January 1987, six known patients were massively overdosed while being treated by the Therac-25 medical linear accelerator. Several suffered devastating injuries and some died. The accidents became one of the defining warnings in the history of software engineering because the computer was not merely keeping records or displaying information: it was controlling a machine capable of delivering potentially lethal radiation.
The story is often reduced to a notorious software bug. That is true, but incomplete. The deeper failure was that software defects were allowed to become physically catastrophic because independent hardware protections had been removed, operators received misleading information, safety analysis largely discounted software failure, earlier warning signs were not connected quickly enough and the manufacturer initially struggled to accept that its machine could be responsible.
That is also why Therac-25 still matters. Modern medical devices, cars, aircraft, industrial equipment and AI-enabled systems increasingly place software between a human decision and a physical consequence. Therac-25 showed what can happen when the software becomes part of the safety system but is still treated as if it were merely another convenience.
A Cancer Treatment Machine With Enormous Power
The Therac-25 was a medical linear accelerator developed by Atomic Energy of Canada Limited, or AECL. Linear accelerators produce high-energy beams that can destroy cancerous tissue. Electrons can be used for relatively shallow treatment, while the machine can also convert the electron beam into X-rays capable of reaching deeper tissue.
The technology itself was not inherently reckless. Radiation therapy depends on precisely delivering powerful radiation to a carefully selected part of the body while limiting damage to surrounding healthy tissue. The danger lies in how enormous the difference can be between a therapeutic dose and an uncontrolled beam.
AECL had previously worked on the Therac-6 and Therac-20. Those machines incorporated computer control, but the computer largely supplemented machinery that retained conventional hardware protection. The Therac-20, for example, had independent protective circuitry and mechanical interlocks capable of preventing unsafe operation even if the software behaved incorrectly.
The Therac-25 represented a significant step forward. It was compact, versatile and designed around computer control from the beginning.
That innovation also introduced the central vulnerability.
The Safety System That Had Quietly Disappeared
The Therac-25 gave software far more responsibility for safety than its predecessors. AECL did not duplicate all of the hardware interlocks that had existed in earlier machines because the computer could monitor and control those functions itself.
This produced an important distinction.
On the Therac-20, a software error might command something dangerous and an independent physical system could still stop it. On the Therac-25, software increasingly had to detect the dangerous condition, understand it and prevent itself from carrying it out.
A later discovery made the difference painfully clear. A related software problem existed in the Therac-20, but its independent hardware protective circuits prevented the beam from firing dangerously. The software could fail without the failure reaching the patient.
The Therac-25 had eliminated much of that final defensive layer.
Its 1983 safety analysis also contained a revealing assumption. Residual programming errors were effectively excluded after AECL reasoned that extensive testing had reduced them sufficiently. Computer failures considered in the fault analysis were largely treated as hardware faults or random execution errors rather than systematic software defects.
The result was not simply buggy software.
It was a machine whose architecture placed extraordinary trust in software while its safety reasoning did not treat software failure with equivalent seriousness.
The First Patient Said The Machine Had Burned Her
The first known accident occurred at Kennestone Regional Oncology Center in Marietta, Georgia, on June 3, 1985.
A 61-year-old woman who had undergone surgery for breast cancer was receiving follow-up radiation treatment near her collarbone. When the Therac-25 activated, she experienced an extraordinary sensation of heat and immediately believed that she had been burned.
At first there were no obvious external marks. Soon, however, the treatment area reddened and swelled. Pain increased. Her shoulder became immobile and the injury eventually developed into an unmistakably severe radiation burn.
The facility physicist contacted AECL and asked whether the machine could operate improperly in electron mode. AECL engineers responded that the suspected failure was not possible.
Later estimates suggested the patient may have received somewhere in the region of 15,000 to 20,000 rads, although the exact dose could never be established. For comparison, the investigation noted that a typical individual therapeutic dose was around 200 rads. These figures should not be compared simplistically with whole-body radiation exposure because the effects depend heavily on the area and tissue exposed, but the difference in scale was extraordinary.
Her breast eventually had to be removed because of radiation damage. She lost the use of her shoulder and arm and experienced continuing pain.
Yet the machine was not immediately recognised as the cause.
That mattered enormously.
A Second Overdose And The Wrong Explanation
Seven weeks later, another patient was injured on a Therac-25 in Hamilton, Ontario.
On July 26, 1985, a 40-year-old woman receiving treatment for cervical cancer was being treated when the machine stopped after approximately five seconds and displayed an “H-tilt” error. Crucially, its dosimetry display indicated that no dose had been delivered.
The machine had entered what it called a treatment pause rather than a full suspension.
Operators were used to Therac-25 interruptions. Because the computer said no radiation had been delivered, the operator followed normal procedure and pressed the command to proceed.
The machine stopped again.
The process was repeated several times.
Only after the fifth pause did the system enter a treatment suspend. Afterward, the patient reported a burning sensation and what felt like an electrical tingling shock in her hip. She later developed pain and swelling.
AECL investigated but could not reproduce the malfunction. Suspicion fell on a microswitch associated with the rotating mechanism that positioned components required for different treatment modes. Hardware and software changes followed.
AECL subsequently described the modifications as producing an improvement in hazard rate of at least five orders of magnitude, even though the investigation itself had not firmly established the precise cause of the accident.
The patient died in November 1985 from her cancer. An autopsy concluded that cancer caused her death, but the radiation injury was so severe that a total hip replacement would otherwise have been required. An AECL technician later estimated that she had received between 13,000 and 17,000 rads.
Canadian authorities had already identified a deeper problem. Recommendations included changing the system so certain malfunctions caused treatment to terminate rather than allowing an operator simply to press a key and continue. An independent means of verifying the position of critical machinery was also proposed.
Not all of those protections were immediately implemented.
Another Warning Appeared In Washington
In December 1985, a patient at Yakima Valley Memorial Hospital in Washington developed an unusual striped reddening of her skin after Therac-25 treatment.
Doctors and physicists tried to determine what had caused it. Chemotherapy was considered and rejected as an explanation. A heating pad used by the patient was investigated. Staff examined whether its internal wires could explain the pattern, but they did not match.
The hospital contacted AECL.
In February 1986, AECL responded that, after consideration, it did not believe the damage could have been produced by a Therac-25 malfunction or operator error. Among the reasons offered was the apparent absence of similar accidents.
That assumption was dangerously misleading.
There had already been serious incidents in Georgia and Ontario.
The Yakima patient continued suffering. When the case was reassessed after another Therac-25 accident the following year, she was found to have developed a chronic ulcer, tissue necrosis and persistent pain. Surgery and skin grafting eventually relieved much of the damage.
Three warning events had now appeared.
The decisive breakthrough came in Texas.
The Tyler Accident That Finally Exposed The Bug
On March 21, 1986, a patient at the East Texas Cancer Center in Tyler arrived for the ninth treatment in a course of radiation therapy after the removal of a tumour from his back.
His prescribed treatment was a 22-MeV electron beam delivering 180 rads.
The operator was experienced and fast at using the Therac-25 terminal. While entering the prescription, she accidentally selected X-ray mode rather than electron mode. She noticed the mistake, moved the cursor back up the screen and corrected the entry.
Everything on the screen appeared to behave normally.
The prescription was shown as verified.
The terminal indicated that the beam was ready.
She activated treatment.
The machine abruptly stopped and displayed an obscure warning:
Malfunction 54.
The terminal simultaneously suggested that only a tiny fraction of the requested radiation had been delivered. The error was classified as a treatment pause rather than a more serious suspension. The documentation available to the centre provided little useful explanation of what Malfunction 54 actually meant.
The operator did what experience had taught her to do.
She pressed proceed.
Why Malfunction 54 Was So Dangerous
The patient knew immediately that something was wrong.
He had already undergone eight treatments and understood what normal radiation therapy felt like. This time he experienced a sensation he later compared to an electrical shock or hot coffee being poured onto his back. He heard buzzing and tried to rise from the table.
At precisely the wrong moment, the operator initiated the next attempt.
He felt another shock, this time affecting his arm.
Normally, the operator could communicate with the patient using audio and video equipment. On that day, however, the video system was unplugged and the audio monitor was broken. The patient eventually reached the door and began pounding on it.
Doctors initially suspected an electrical shock.
What had actually happened was far more serious.
Later simulations suggested the patient may have received between approximately 16,500 and 25,000 rads concentrated in a small area in less than one second. Exact doses in the Therac accidents remain uncertain because reproducing the malfunction produced substantial variations between machines.
The patient subsequently developed devastating neurological injuries. He lost the use of his left arm and both legs and experienced further complications before dying approximately five months later from consequences of the overdose.
AECL engineers tested the machine but initially could not reproduce Malfunction 54.
Once again, an enormous radiation overdose appeared impossible.
The machine eventually returned to service.
Three Weeks Later, It Happened Again
On April 11, 1986, another patient entered the same East Texas Cancer Center for radiation treatment, this time for skin cancer on the side of his face.
The same operator entered the prescription.
Again, she needed to correct the treatment mode from X-ray to electron.
Again, she rapidly edited the information.
Again, the screen eventually indicated that the beam was ready.
And again, the Therac-25 stopped with Malfunction 54.
This time the intercom was working.
The operator heard the disturbance and rushed into the room. The patient described a feeling of fire on the side of his face. He had seen a flash and heard a sizzling sound.
He died on May 1.
The autopsy identified acute high-dose radiation injury affecting parts of his brain.
After the second Tyler accident, the centre did something crucial.
It stopped assuming that the machine was right.
Physicist Fritz Hager worked with the operator to reconstruct exactly what she had done. Reproducing the fault was difficult at first, because the dangerous sequence depended on something apparently harmless: speed.
The operator had become too good at using the machine.
The Eight-Second Race Condition
The Therac-25 software performed multiple tasks concurrently and passed information between them through shared variables.
When an operator entered a treatment prescription, different parts of the program processed the selected mode, energy and physical configuration. Setting the machine's bending magnets took roughly eight seconds. The software was supposed to notice prescription changes and adjust the machine accordingly.
But under a particular sequence, it did not.
If an experienced operator completed an initial prescription, moved back up to change the mode or energy and returned to the command line within the critical timing window, information displayed on the terminal could be updated without all of the machine's internal settings being correctly updated.
Different software tasks could therefore effectively disagree about what treatment had been requested.
This is a classic example of a race condition: the result depends upon the timing and ordering of concurrent operations.
The Therac code used shared variables and flags to coordinate those processes. Under the Tyler sequence, an edit could occur after part of the software had already passed the point where it would recognise the change. The screen could reflect the new prescription while other machine parameters remained based on the earlier one.
That created an extraordinarily dangerous state.
In X-ray mode, the accelerator generated a powerful electron beam intended to strike a target inside the machine. The target converted that energy into therapeutic X-rays and equipment spread and moderated the resulting radiation appropriately.
If the system delivered the high-current beam while configured incorrectly for electron treatment, the protective hardware that should have intercepted and transformed that energy was absent from the beam path.
The patient could therefore receive an enormously concentrated electron beam.
Hager eventually reproduced Malfunction 54 by entering and editing the information quickly enough. Once AECL understood that speed was essential, its engineers reproduced the failure too. Measurements by AECL reached around 25,000 rads at the centre of the field under the simulated condition.
The supposedly impossible failure had become repeatable.
Why The Software Bug Became Lethal
The race condition explains the Tyler accidents, but stopping the story there misses its most important lesson.
Software has bugs.
Safety-critical engineering exists partly because nobody should assume that sufficiently complex software will never behave incorrectly.
The older Therac-20 contained related software behaviour. Yet independent hardware protective circuits prevented dangerous irradiation. New students using unusual editing patterns could trigger failures, but physical safeguards interrupted the machine before those failures reached a patient.
Therac-25 relied far more heavily upon software to enforce the same safety requirements.
That transformed a software fault into a radiation accident.
There were other layers of failure. Error messages such as Malfunction 54 were cryptic. Operators had become accustomed to frequent non-catastrophic interruptions and therefore learned that pressing proceed was often normal. The screen could report an underdose while the patient had actually been massively overdosed. Early accidents were not connected efficiently enough. Safety analysis had discounted residual software bugs. Independent interlocks that might have blocked unsafe states were absent.
Even the operator interface contributed.
A good safety system does not merely detect an error. It communicates the seriousness of that error clearly enough that a human knows what to do.
Therac-25 sometimes did the opposite.
It could make a catastrophic situation look routine.
The Final Yakima Accident
The story was still not finished.
On January 17, 1987, another patient was overdosed using a Therac-25 at Yakima.
This accident involved a different software flaw. The precise coding sequence differed from the Tyler race condition, demonstrating something even more worrying: correcting one known defect did not make the underlying design safe.
Investigators determined that another software race condition could allow treatment to begin while the machine was in an inappropriate configuration.
Preliminary measurements suggested a dose of roughly 4,000 to 5,000 rads could be delivered under the fault condition. Because treatment had been attempted twice, the patient may have received approximately 8,000 to 10,000 rads instead of the prescribed 86.
The patient died in April. He already had terminal cancer, complicating attribution of his death, but legal action alleged that the overdose shortened his life and caused additional suffering.
The extraordinary detail was what followed the investigation.
A quality assurance manager reportedly concluded that hardware and software changes already planned after the Tyler accidents would have prevented the Yakima event had they been installed.
By then the regulatory response was escalating rapidly.
The Regulator Steps In
The first Tyler reports reached the US Food and Drug Administration through Texas health authorities and triggered a deeper investigation.
The FDA subsequently declared the Therac-25 defective and required AECL to notify purchasers, investigate the problem and produce a corrective action plan. The final programme contained more than 20 hardware, software and documentation changes.
One early temporary measure now looks almost surreal.
Operators were instructed not to use the cursor-up key for editing. Users were told to remove its key cap and prevent the switch from operating, forcing incorrectly entered prescriptions to be reset and entered again from the beginning.
The FDA considered the initial notification inadequate because it failed to describe properly the defect and its hazards.
More substantial changes followed.
Malfunctions that previously caused treatment pauses were changed so they produced treatment suspensions. Independent hardware circuitry was introduced to shut down the beam when inappropriate radiation levels were detected. Software was changed, documentation revised and the machine underwent extensive safety analysis.
By February 1987, US and Canadian authorities were moving to stop routine use of Therac-25 machines until corrective work was completed. The corrective action process continued through multiple revisions and extensive testing.
The machine eventually became safer.
The damage that forced those changes could not be reversed.
It Was Never Just One Programmer’s Mistake
It is tempting to imagine the Therac-25 disaster as a morality tale about one incompetent programmer.
The evidence supports a much more uncomfortable conclusion.
Investigators Nancy Leveson and Clark Turner argued that accidents of this kind are usually system accidents. Therac-25 involved interactions between software design, hardware architecture, testing, safety analysis, organisational communication, regulatory oversight, interface design, operator behaviour and assumptions about what could or could not fail.
The programmer matters.
But so does the engineer who decides that a software check makes a physical interlock unnecessary.
So does the organisation that assumes extensive testing means remaining software defects can be omitted from hazard analysis.
So does an interface that communicates a catastrophic condition through an obscure numerical error.
So does a reporting system that prevents one hospital from immediately learning what happened at another.
And so does a culture in which evidence contradicting the assumed safety of the machine is initially easier to dismiss than the assumption itself.
The most dangerous bug in Therac-25 was therefore not simply buried in its source code.
It was the belief that the rest of the system could safely depend on the code always being right.
What Therac-25 Changed
Medical-device regulation in the United States had already expanded significantly before the Therac accidents. The 1976 Medical Device Amendments created the modern risk-based classification system, while medical device reporting rules required manufacturers to notify the FDA of certain serious incidents from December 1984.
The regulatory environment continued evolving afterward. The Safe Medical Devices Act of 1990 expanded post-market surveillance and introduced reporting requirements affecting healthcare facilities, while later quality-system requirements made software validation an explicit part of medical-device development.
The FDA's later software-validation guidance describes software validation, design controls, verification and disciplined development as essential components of medical-device safety. The agency's analysis of recalls from 1992 to 1998 found hundreds attributable to software failures, illustrating that Therac-25 was not the end of the problem.
Nor could regulation ever guarantee bug-free software.
The stronger lesson is architectural.
When software controls something capable of killing a person, designers should ask what happens when the software is wrong rather than merely trying to prove that it is right.
Independent protection matters.
Fail-safe states matter.
Clear warnings matter.
Hazard analysis must include software.
Unusual input sequences matter.
Operators who use equipment differently from developers' expectations matter.
And near misses must be treated as evidence rather than noise.
Why Therac-25 Matters Even More In The AI Era
More than four decades after the Therac-25 was developed, software has moved much further into the physical world.
Modern software can influence drug delivery, diagnostic decisions, surgical equipment, transport, industrial systems and patient monitoring. FDA guidance now explicitly addresses device software functions, including software whose failure could pose a risk to patient safety, while the agency continues researching methods for assessing software-controlled medical systems.
AI adds another layer. Some medical systems increasingly use machine-learning models rather than purely deterministic software, and regulators are now developing lifecycle approaches for AI-enabled medical-device functions. The specific technology is different from Therac-25, but the foundational question is strikingly similar: how much authority should software receive when an incorrect output can reach a human body?
Therac-25 therefore survives in engineering education because its lesson is larger than radiation therapy.
The machine did not become dangerous simply because somebody wrote defective code. It became dangerous because designers trusted software enough to remove other defences, the system sometimes concealed its own failure, human operators were given misleading information and separate warning events were not transformed quickly enough into a shared understanding of risk.
The six known accidents occurred between 1985 and 1987. The engineering problem they exposed has not disappeared.
We have simply given software control over far more of the world.

