The Last System Between an Error and an Explosion
Key Intel / TL;DR
  • › A safety instrumented system exists to take a process to a safe state when everything else has failed, and your risk assessments credit it as independent.
  • › Triton showed in 2017 that attackers will go after the safety system directly, and it came to light because the plant tripped and shut down.
  • › Independence is an engineering claim that stops being true the moment the safety system shares a network path, a workstation, or a remote access route with the control system.
  • › In most plants the safety system belongs to process safety engineering and is invisible to the security program, which is the ownership gap attackers use.
  • › Start with an assessment that physically traces every connection to the safety system, because that tells you whether the credit you are taking is real.

Every process plant runs two systems that look similar from a distance. The basic process control system keeps the plant producing, adjusting valves, pumps, and temperatures thousands of times a day. The safety instrumented system does almost nothing, almost all of the time. It watches a small number of conditions, and when pressure, temperature, or level moves outside the range the process was designed for, it takes the process to a safe state regardless of what the control system is trying to do.

That second system is the one your business is counting on when everything else goes wrong. It is also the one that your risk assessment, your insurer, and your process safety program all assume is independent of the first. For most facilities that assumption has never been tested from a cybersecurity standpoint, and that is the gap this article is about.

Why the Safety System Carries So Much Weight

In a formal process hazard analysis, the safety instrumented system is credited as an independent protection layer. That credit is what allows a facility to accept a level of risk that would otherwise require a physical redesign, a larger buffer zone, or a different process entirely. Functional safety standards such as IEC 61511 set out how these systems are designed, tested, and maintained so that the credit holds up.

In business terms, the safety system is load-bearing. It shows up in the risk numbers that justify your operating model, in the assumptions behind your property and liability coverage, and in the compliance posture you present under process safety regulations such as OSHA’s Process Safety Management standard. If it cannot be relied on, several documents you have signed are describing a plant that does not exist.

The engineering discipline behind these systems is serious and mature. The trouble is that it was built to defend against random hardware failure and human error, and an adversary who studies the system and acts deliberately is a different kind of problem.

What Triton Showed the Industry

In June 2017, a petrochemical plant in Saudi Arabia shut down unexpectedly. The trip was initially put down to a mechanical fault. A second unplanned shutdown in August prompted a closer investigation, which found malware, later named Triton and also known as Trisis, targeting the plant’s Schneider Electric Triconex safety controllers.

The attackers had reached the safety system itself, the layer that exists to prevent an explosion or a toxic release. The work was later attributed to a Russian government research institute, the Central Scientific Research Institute of Chemistry and Mechanics, and in 2022 the FBI warned that the same group was still targeting the global energy sector after a US indictment against one of its employees was unsealed.

The detail that should stay with every plant manager is how it came to light. The safety system tripped and shut the plant down, which is the safety system doing its job, and that trip is why anybody went looking. An intrusion into a safety system that did not cause a trip could, in principle, sit undetected until the day the system was needed and did not respond.

Independence Is a Claim You Can Test

Independence in a safety case means the safety system does not share a failure mode with the control system it backs up. Engineers achieve that with separate sensors, separate logic solvers, and separate final elements such as shutdown valves.

The cybersecurity question is whether the digital paths are separate too, and in many facilities they are not. Common patterns we see include a shared network segment between the safety and control systems for convenience, an engineering workstation that programs both, and a remote access path that an integrator uses to support the whole site. We have written about how the integrator still has a way in at many facilities, and that same access frequently reaches the safety layer as well.

Each of those connections was added for a legitimate operational reason, and each one is a shared failure mode from an attacker’s perspective. If a single compromised laptop can reach both systems, the two layers are no longer independent in the way your risk assessment assumes. The Purdue model gives you the vocabulary for where the safety system should sit, and an honest diagram of where it actually sits is usually more instructive.

Many safety controllers also have a physical mode switch that determines whether they accept program changes. A controller left in a mode that allows remote changes will accept new logic from anything that can reach it on the network, and in a busy plant the switch position tends to reflect whatever the last maintenance task needed.

The Ownership Gap

In most facilities, the safety system belongs to process safety engineering. It is specified, tested, and maintained under a functional safety management program with its own procedures and its own auditors. The security team often has limited visibility into it, and corporate IT typically has none.

That arrangement made sense when the main risks were hardware failure and procedural error. It leaves a gap when the risk is an adversary, because the people who understand the safety system are not usually the people tracking remote access, patch levels, and network paths, and the people doing that tracking often have no idea the safety system is on the network at all.

The result is a system that everybody assumes somebody else is watching. Process safety assumes the network is secured, and security assumes the safety system is isolated because the engineering documentation says it is.

What This Costs When It Goes Wrong

A compromised business system is expensive. A compromised safety system belongs to a different category of loss entirely, because the consequence is physical. The realistic outcomes run from an extended unplanned shutdown, with the lost production and restart costs that come with it, to a loss-of-containment event with injuries, an environmental release, regulatory investigation, and litigation that follows a company for years.

There is also a quieter cost that arrives without any incident at all. If an assessment shows your safety system is not independent, the risk assessments built on that independence need to be revisited, and that can change your insurance position and your compliance position before anything has happened. Finding that out on your own schedule is far cheaper than finding it out from an insurer’s engineer or a regulator.

Where to Start

The first step is an assessment, and it needs to be a physical one. Documentation describes how the safety system was designed, which is often different from how it is connected today.

Trace every connection to the safety system. Walk the cabling, check switch configurations, and identify every device that can communicate with a safety controller, including engineering workstations, historians, and remote access gateways. Anything that can reach both the safety system and the control system is a finding.

Check the mode switch and the procedure around it. Confirm where each safety controller’s mode switch sits today and who is allowed to change it, and make returning it to the locked position part of every maintenance close-out.

Give the safety system a dedicated, controlled engineering workstation. The machine that programs the safety logic should not be the machine used for everything else, and it should be disconnected when nobody is actively using it.

Bring the safety system into both programs. Put it in the security program’s asset inventory and monitoring scope, and add a cybersecurity review to the functional safety management process. Neither team can cover this alone, and the gap between them is exactly where the risk sits.

Revisit the risk assessment with the findings in hand. If the assessment shows shared paths, fix them or update the protection layer credits to reflect reality. Either outcome is better than carrying a number you know is wrong.

A safety instrumented system is the one control in your facility that is meant to work when every other control has failed. That is a strong reason to verify it is still standing apart from everything else, and the verification is a short, contained project compared with what it protects against.

Dusten Trounce is Director of Physical Security at Grab The Axe.

Distribute Intel
Dusten Trounce
Director of Physical Security
Dusten Trounce
The Growth Architect.

A leader defined by a 'bias for action,' Dusten specializes in physical security assessments that impact profitability. He leverages high-logic strategies to pinpoint high-ROI vulnerabilities, ensuring defense measures actually scale with the business.

View Author Page →