Site icon UK Aviation News

NATS report reveals one-millisecond software defect behind September air traffic chaos

NATS Swanwick ATC Centre

NATS Swanwick ATC Centre

A previously unknown software defect with an exposure window of just around one millisecond was responsible for the major National Airspace System failure that caused widespread disruption across UK aviation on 8 September, according to a preliminary investigation by NATS.

The report, published on 16th of September, reveals how a highly specific sequence of events corrupted flight data within the NATS National Airspace System (NAS), eventually forcing London Area Control (LAC) to operate using fallback procedures and triggering severe restrictions on flights across the UK.

More than 2,000 flights were delayed, cancelled or diverted as a result of the incident, with NATS handling around 1,800 fewer flights than it had expected to handle that day.

NATS says the incident was not caused by incorrect actions by either military or civil operators, was unrelated to the 2023 NATS flight planning system failure and was not connected to the radar issue experienced in July 2025.

There is also no evidence at this stage that the incident was caused by malicious activity or a cyber attack.

The preliminary report instead identifies a legacy software defect within a module responsible for processing requests to allocate or reassign aircraft squawk codes.

Squawk codes are used by aircraft to identify themselves to air traffic control systems, allowing controllers to associate radar information such as altitude, speed and direction with the relevant flight plan.

In this case, a perfectly valid manual request for a squawk code was being processed when a higher-priority request arrived.

The NAS temporarily paused processing of the original request while dealing with the higher-priority activity. When processing resumed, the software defect caused the original request to be handled incorrectly, resulting in corrupted data.

The circumstances required an exceptionally precise combination of events.

According to NATS, the vulnerable section of software had an exposure window of approximately one millisecond. The higher-priority request had to arrive during that exact period while the original request was part-way through updating a value.

Had the second request arrived even one millisecond earlier or later, the update would have completed normally.

The corrupted data did not immediately bring the system down.

At 10:02 on 8th of September, the London Area Control system detected that its connection with the NAS had been lost. The connection automatically recovered after approximately 45 seconds, and at the time there was no apparent operational impact.

Engineers subsequently investigated the incident and found no evidence of a hardware failure.

However, at 12:32, the connection began dropping repeatedly as the LAC system encountered increasing difficulty processing the corrupted flight data.

As the problem worsened, some automated functions became unavailable and controllers had to undertake tasks manually.

At 12:45, the first direct impact on airlines, airports and passengers began as NATS imposed restrictions on the number of aircraft entering certain sectors and temporarily stopped departures from UK airports.

By 13:32, the connection between the LAC system and the NAS had failed completely and controllers switched to established fallback procedures.

Further restrictions were imposed, with some sectors limited to as few as 30 aircraft per hour. Restrictions were also placed on arrivals at some UK airports, meaning aircraft due to arrive in Britain could not necessarily depart their overseas airports.

The result was a rapidly developing imbalance between arriving and departing traffic.

With more aircraft arriving than departing, congestion began building on the ground at airports. Established diversion procedures were subsequently used to divert some inbound aircraft that were already airborne.

NATS says the response was driven by safety considerations throughout.

Controllers retained the ability to communicate with aircraft and monitor them on radar, while standard fallback procedures were used to maintain separation.

The report states that all aircraft remained safely separated throughout the incident and that safety margins were maintained at all times.

A major decision was then taken to restart the NAS.

At 13:45, engineers and the Major Incident Manager concluded that the problem was most likely within the NAS and that the optimum recovery strategy was a controlled restart followed by reloading flight data from the London Area Control system.

The process could not simply be undertaken immediately, however, because the NAS supplies flight data to multiple national and international air traffic control units and several airports.

Restarting it therefore required extensive planning and coordination.

The NAS restart began at 15:17 and was completed at 16:09.

But restoring the system did not immediately restore normal operations.

Because the LAC system and NAS had been disconnected, flight data had become unsynchronised. Engineers had to reconcile duplicate flight plans, correct aircraft code and callsign associations and resolve inconsistencies caused by new and amended flight plans submitted while restrictions were in place.

This reconciliation process continued until 18:50, when NATS says all systems had returned to stable operation and normal ATC operations could resume.

All remaining airspace restrictions were subsequently lifted at 19:30.

The UK departure stoppage had been applied intermittently during the incident, amounting to around four and a half hours of stopped departures over a six-hour period.

Although the technical incident was resolved on the evening of 8th of September, the consequences continued for several days as airlines and airports worked to recover their schedules.

NATS held six coordination calls with industry stakeholders on 9th of September, followed by further calls on 10th of September. The final coordination call was held on 11th of September.

The organisation says a software fix for the defect has already been developed and delivered by its supplier. It is currently undergoing testing and safety assurance before being deployed permanently.

In the meantime, NATS has introduced additional engineering measures to reduce the risk of a repeat incident.

The new arrangements include additional monitoring and reporting around link failures between the LAC system and NAS, with clearer escalation procedures for suspected failures.

NATS says these measures should allow a faster recovery if a similar problem occurs while the permanent software fix is being tested and deployed.

However, the preliminary report is not the end of the investigation.

NATS has committed to completing a full Major Incident Investigation within 60 days of the 8th of September incident.

That investigation will examine the underlying causes, the scale and duration of the disruption, the resilience and recovery arrangements, the effectiveness of NATS’ command and control structure and the clarity and timeliness of communications with airlines, airports, EUROCONTROL and other stakeholders.

It will also examine system health monitoring, software defect records, reliability trends, asset lifecycle management and the implementation of recommendations made following previous major incidents.

The report is particularly significant because the 8th of September disruption came after previous major NATS incidents had already prompted changes to its resilience and incident-management arrangements.

NATS says recommendations arising from previous internal investigations and the CAA’s independent review following earlier incidents had been implemented and independently verified by the CAA.

These included the introduction of a dedicated Major Incident Manager, a role that was used during the September incident.

The Civil Aviation Authority is separately conducting its own independent review of the disruption, following a request from the Government. The CAA has said its review will consider what happened and how well NATS is set up to provide a resilient service in the future.

For passengers and airlines, the most striking finding from the NATS report is therefore not simply that software failed, but how an apparently routine and valid flight-data request was able to trigger a chain of events across a critical national system.

A defect that could remain hidden unless another request arrived within a window measured in milliseconds ultimately led to the restriction of thousands of flights and disruption that continued long after the original technical problem had been resolved.

NATS’ preliminary findings are subject to change as the full investigation continues, with the final report expected to provide a more detailed assessment of why the defect existed, why it was not previously identified and whether other characteristics of the wider system increased the likelihood of the failure occurring.

Exit mobile version