Why Air-Gapped OT Environments Are Still Insecure

Noman Nasir MinhasPublished ot-securityicsair-gapmalware-analysis

A technical analysis of OT isolation: Stuxnet, TRITON, offline malware channels, industrial protocols, WirelessHART, measurement integrity, and maintenance—with reproducible simulation labs.

Share

An air gap can remove an attacker's continuous network path to a controller. It cannot establish that the engineering workstation programming that controller is trustworthy, that the running application matches the approved project, or that an authenticated sensor value represents the physical process. Those are different claims, requiring different evidence.

The difficult question is not whether a USB drive can carry malware. It is how an isolated system decides that an imported artifact, authorized command, replacement instrument, or maintenance action deserves the authority it receives. A useful assessment must follow that authority all the way from the boundary to the physical consequence. Counting disconnected Ethernet cables stops too early.

This article assumes familiarity with PLCs, engineering workstations, industrial networks, and safety instrumented systems. It focuses on mechanisms rather than a checklist of basic controls. The historical cases are deliberately bounded: observed incident means publicly documented deployment or intrusion; analyzed capability means behavior established from tooling; laboratory demonstration means a result under an experiment's prerequisites; engineering inference means the analysis or illustrative model developed here. Capability is not proof of successful use, and an OT attack is not automatically evidence of an air-gap crossing.

#1. Define the Isolation Claim Before Evaluating It

#A disconnected network is only one boundary

Consider four architectures that are often described using the same word. A physically isolated network has no permitted live communication path to the external network in the stated operating configuration. A segmented network has paths mediated by firewalls, routing, or gateways. A temporarily connected network admits a maintenance path for a bounded interval. A network exporting through a unidirectional gateway permits a particular direction of communication rather than no communication at all. NIST discusses unidirectional gateways as a distinct architectural control, not a synonym for total disconnection. NIST SP 800-82r3, final, architecture guidance.

For an assessment, the claim needs a subject, direction, operating state, and time interval. “The safety engineering station has no inbound routed path from corporate IT during production” is testable. “The plant is air-gapped” usually leaves unanswered whether the historian, wireless gateway, vendor laptop, remote support appliance, or commissioning network belongs to the claimed boundary. Missing scope does not prove compromise; it makes the assurance claim impossible to falsify.

A boundary can be effective against one transfer and irrelevant to another. A firewall may block a TCP session but say nothing about a project archive carried on approved media. A data diode may constrain network direction but leave an inbound maintenance workflow outside its scope. A locked controller cabinet may prevent unauthorized physical handling while a legitimate engineering station retains the authority to replace application logic. These are not contradictory observations. Each control constrains a different operation.

The first technical task is therefore to enumerate state-transfer mechanisms, not just interfaces. They include application downloads, firmware updates, recipes, configuration databases, calibration values, certificates, device descriptions, and backups. A mechanism counts even when it moves a file rather than a network packet. The receiving application may grant the artifact more authority than the transport review anticipated.

#Time-ordered paths defeat static topology arguments

Engineering inference. Let a graph G(t) represent permitted transfers at time t. A conventional network review asks whether an external source and a protected controller are connected in G(t). An operational review must also ask whether a sequence of transfers exists at increasing times, with intermediate equipment retaining state between them.

A laptop can receive an artifact from a vendor network on Monday, disconnect, and deploy it to an OT workstation on Tuesday. There need never be an instant at which both networks are reachable. Persistent storage supplies the missing edge. The laptop's memory of Monday is available to Tuesday's trust domain. A packet-capture-based proof of “no simultaneous bridge” can be completely correct and still fail to answer this separate question.

Time-ordered trust path from a vendor workstation through a maintenance laptop to an OT workstation and controller, without a simultaneous network bridge

We can express the distinction without pretending to calculate breach probability:

TEXT
Instantaneous isolation:
    no permitted external -> controller path in G(t)
 
Lifecycle isolation assurance:
    every time-ordered external -> artifact -> tool -> controller path
    has controlled acceptance boundaries and independently checked outcomes

The second statement is harder because the acceptance boundary may be a software parser, a person approving a change, or a controller accepting a programming request. A network diagram does not show whether a project import invokes a vulnerable component, whether its dependencies are resolved from local mutable directories, or whether the approved file is the one eventually downloaded.

A practical inventory should represent each transfer with five attributes: producer, transport, consumer, granted authority, and verification method. “USB to engineering station” is incomplete. “Signed vendor package, transferred on dedicated media, interpreted by firmware updater, allowed to alter executable code, checked against an independently acquired release manifest” is substantially more useful. It still does not prove the updater or signing authority is uncompromised, but it exposes those remaining dependencies.

#Integrity boundaries differ from reachability boundaries

Suppose a controller accepts commands only from one workstation. That reduces the set of origins, but every program executing with sufficient authority on that workstation may inherit its permitted path. If the workstation's identity is the authorization principal, malicious software does not need to create a second identity. It needs to act through the first.

An assessment should ask where authority changes form. A file becomes parsed project data. Parsed project data becomes compiled controller logic. A programming operation becomes a running task. A running task becomes an actuator command. At each conversion, the previous check may no longer cover the property that matters. A hash proves equality of bytes, not equivalence of process behavior under all operating modes.

The distinction also explains why stronger isolation remains useful. Removing an online path can eliminate classes of remote exploitation, reduce interactive attacker control, and constrain exfiltration. Those are real benefits. The mistake is to promote a reachability result into a comprehensive integrity claim. The remainder of the article examines the transformations that such a promotion overlooks.

#2. Stuxnet Beyond the USB Story

#Host exploitation and controller manipulation are separate stages

Analyzed capability. ESET's original analysis describes removable-media execution through the Windows LNK vulnerability, additional internal propagation paths, and a replacement of the Step7 communication library s7otbxdx.dll. Its Export 17 analysis identifies a wrapper that forwards ordinary operations while intercepting selected variable and block operations, including s7blk_read and s7blk_write. These mechanisms concern particular samples and vulnerable software; they do not reconstruct the first physical insertion into Natanz. ESET, Stuxnet Under the Microscope, revision 1.31, §§3.1–3.4 and Export 17, pp.53–54.

That separation changes the security question. Windows compromise supplies execution in an engineering context. The engineering context supplies a path with controller-programming authority. The physical process is affected only after the malicious application logic reaches an appropriate target. Treating all three as one “USB attack” hides the different prerequisites and the checks that might interrupt each stage.

Engineering inference. Think of the chain as conversions of authority rather than a list of vulnerabilities. A Windows parser flaw turns data encountered by a host into host execution. Control of the engineering communication layer turns host execution into the ability to influence controller operations. Knowledge of the target process turns that influence into meaningful sabotage. Exploiting the first boundary is neither necessary nor sufficient to explain the final physical behavior in every possible intrusion.

A project-based workflow deserves separate attention. If a project or its associated environment is a carrier, its nominal business purpose can move it into a higher-trust context. The acceptance question becomes: which files and dependencies are interpreted when the project opens, which are included when it builds, and which outputs are actually downloaded? This is an engineering inference about artifact handling, not a claim that any arbitrary project format automatically executes malware.

The following simplified call flow highlights the shared trust boundary without reproducing the malware:

TEXT
Engineering application
    -> loaded communication library
        -> selected block / variable operations
            -> controller engineering interface
                -> running process application
 
A compromised library can influence both the write path
and the operator's subsequent view through that same path.

The operator may have a legitimate account, an approved work order, and the correct target address while the implementation underneath supplies a different operation. Access logs identifying the workstation are therefore insufficient to establish application integrity. They answer where an operation originated, not which code constructed it.

#Two PLC attacks must not be collapsed into one

Analyzed capability. Langner distinguishes the S7-417 overpressure attack from the S7-315 rotor-speed attack. The former includes recording and playback of process inputs; the latter does not use that same replay mechanism. He explicitly calls out this common conflation. His analysis also separates controller manipulation through product functionality from the Windows exploit chain. Langner, To Kill a Centrifuge, overpressure/rotor-speed chapters, especially printed p.13.

Stuxnet engineering compromise branches into a distinct S7-417 overpressure payload with replay and S7-315 rotor-speed payload without that same replay mechanism

Engineering inference. These differences matter to monitoring design. A defender looking only for repeated sensor sequences is asking a useful question about replay, but not necessarily about a payload that suspends or replaces legitimate execution. Conversely, detecting an abnormal actuator trajectory does not establish that the operator's measurements are faithful. The monitored property has to correspond to the mechanism under consideration.

A process-aware review should therefore maintain separate evidence for the write path, the execution path, and the observation path. The write path establishes what was deployed. The execution path establishes which tasks and parameters currently govern behavior. The observation path establishes how the reported process state was obtained. If all three are mediated by the same compromised engineering stack, three matching observations may represent one failure rather than three independent confirmations.

Consider an illustrative controller with an approved function block and a separate retained parameter bank. A project comparison can report identical logic while a retained setpoint differs. A communications capture can show an authorized download while a later mode change activates a different branch. A readback can show expected bytes while its interpretation depends on an untrusted engineering library. None of these possibilities is a new claim about Stuxnet; they are reasons to specify what each verification actually establishes.

#Offline execution changes response assumptions

A physically isolated target may deny an attacker continuous interactive control. It does not prevent a payload from carrying its own schedule, target conditions, and action sequence. Once the required state has been transferred, execution can depend on local inputs rather than further commands. The relevant question is whether the device possesses malicious executable authority now, not whether it can reach a command server now.

That distinction shapes containment. Disconnecting an already compromised workstation may stop additional communication but leave malicious controller logic running. Removing a workstation may leave altered controller parameters intact. Reinstalling Windows may remove the compromised engineering component while restoring its contaminated project archive. An isolation action must be paired with a hypothesis about where persistent state resides.

A defensible investigation does not assume a clean project is the same thing as a clean controller, or that a clean controller is the same thing as a correctly calibrated process. It enumerates the possible persistence domains and checks them with evidence appropriate to each domain. The cost of that investigation is one reason incident preparation must include controller-specific recovery knowledge before an emergency.

Stuxnet's strongest lesson for this discussion is the combination of access and interpretation. The same trusted engineering relationship can convey instructions and shape the defender's view of those instructions. The historical payload details are important, but the broader engineering question is whether the evidence of correctness is independent of the component whose correctness is in doubt.

#3. Offline Channels Are Communication Protocols

#Fanny: transported storage as a message bus

Analyzed capability. Kaspersky describes Fanny's agentcpd.dll as an offline reconnaissance backdoor that allocates a hidden container using its own FAT16/FAT32 handling when removable media arrives. The media carries information between disconnected systems. The report leaves the actual target uncertain and frames possible preparation for Stuxnet as a hypothesis, not an established operational history. Kaspersky, A Fanny Equation, USB Backdoor and Conclusion, February 2015.

Engineering inference. A transported drive has several properties of a network channel: addressing conventions, message storage, serialization, collection timing, and an opportunity to acknowledge or replace state during the next visit. Its bandwidth and latency differ from Ethernet, but those differences do not eliminate its communication role. For an operation whose objective is periodic collection or delivery of a precomputed task, a slow channel may suffice.

Imagine a benign simulator in which a drive carries a queue of approved configuration changes. Each insertion reads pending messages and writes a receipt. That is already a bidirectional protocol over disconnected transport. Its security depends on who may create messages, how recipients are identified, whether a receipt can be trusted, and whether old messages can be accepted again. The absence of a socket does not answer any of those questions.

This model also exposes an ambiguity in “read-only media.” A device may be physically write-protected during one stage but previously writable in a different domain. A receiving workstation may write its own local execution state even when the transported medium cannot be modified. Conversely, a drive that collects output from the isolated side may become an outbound information channel without any return command capability. Directionality belongs to the whole workflow, not the label on the device.

#Ramsay and the limits of visibility

Analyzed capability. ESET's Ramsay report describes malicious document and installer delivery, along with a version whose spreader infects PE files on removable and shared drives. It reports few visible victims and evidence suggesting development and testing. It does not establish OT controller sabotage. ESET, Ramsay, Attack vectors and Conclusion, May 2020.

Engineering inference. For a plant, an installer can be a more attractive carrier than an obviously irrelevant attachment because support workflows already expect installers. The technical review should follow what the installer changes, not merely whether its name resembles an approved utility. A legitimate executable that has been modified also challenges a filename-only allowlist. The object of verification needs to be the actual bytes and their authority, with the verification performed outside the potentially compromised execution path.

Visibility is a separate concern. An isolated environment can have strong boundary controls and weak external telemetry at the same time. A low public victim count can reflect genuine rarity, inaccessible telemetry, or both. It cannot be converted into a probability of local safety without additional evidence. This is why the article avoids deriving prevalence from a handful of malware discoveries.

#Survey findings should remain date-scoped

ESET's 2021 survey examined 17 air-gap-targeting frameworks. Its findings emphasize USB transport and distinguish connected-side assistance from fully offline workflows; it did not identify analyzed covert physical transmission in that surveyed set. That is a finding about a defined historical corpus, not proof that other channels are impossible or unused today. ESET, Jumping the air gap, Key findings and Communication and exfiltration channel.

The engineering priority follows the evidence: enumerate real transfers before spending the assessment on spectacular covert channels. A routine vendor-media workflow is a concrete path that can be audited. A speculative electromagnetic pathway requires a different threat model and evidence. Both can belong in a sophisticated assessment, but they should not be described with the same confidence or given priority solely because one sounds more exotic.

#4. TRITON: Programming Authority Reaches the Safety System

#What the incident establishes

Observed incident. Mandiant reported an intrusion involving a Windows SIS engineering workstation and TRITON's interaction with Triconex safety controllers; a failed manipulation produced a shutdown that helped expose the activity. The report does not establish that the initial intrusion crossed a physically isolated boundary. Mandiant, original TRITON analysis, December 14, 2017.

Analyzed capability. CISA's updated HatMan report describes a PC-side component and controller payload, in-memory firmware modification, and the analyzed sample's MP3008/firmware 10.0–10.4 scope. Its discussion of operating modes and mitigations is specific to that system; it is not a universal rule for every safety controller. CISA, HatMan Update B, Vulnerable Systems, Technical Details, and Mitigations, February 27, 2019.

TRITON authority chain from a compromised SIS workstation through programming prerequisites to controller memory changes and loss of independent safety assurance

Engineering inference. The case moves the debate from confidentiality to the authority to change the safety function. A safety controller can be physically distant from the corporate network while its engineering station remains the practical gatekeeper of executable behavior. Protecting that relationship requires more than proving that corporate packets cannot reach the controller directly.

The keyswitch is particularly useful as an assessment object because it makes an otherwise abstract authority transition observable. The relevant question is not merely whether the switch has a secure position. It is when programming is permitted, who can change the state, how the state is observed, and how the organization proves the expected position was restored. A well-designed control can still be absent during the exact interval when its protection matters.

#Redundancy does not imply security independence

Consider three redundant channels that all receive application logic from the same engineering project. They may tolerate a hardware fault in one channel. They do not automatically provide three independent judgments about whether the project is malicious. If the same untrusted authority programs every channel, the shared input is a possible common cause.

This is an engineering distinction, not a claim that particular certified hardware is ineffective. Fault tolerance and resistance to malicious authorized changes answer different questions. An assessment should state what kinds of failures a redundancy scheme detects and which configuration relationships remain shared. “Three channels” is not a substitute for identifying common engineering tools, credentials, libraries, and change authorities.

A similar dependency exists between control and safety. Suppose the basic process control system and SIS are administered from separate workstations but both consume the same change-management archive and vendor account. Physical separation improves one boundary, yet the shared upstream authority may defeat the intended independence of changes. The finding depends on the actual workflow; it cannot be inferred from workstation count alone.

#Separate the action from the permission window

An engineering transaction is more than a packet that downloads a block. It has a preparation state, a permission window, an action, verification, and closure. The closure step is where temporary privileges, bypasses, and altered modes are removed. If the workflow treats “download succeeded” as completion, it can return a plant to service with engineering access still enabled.

The expected evidence should pair every opening with a closing event. An approved maintenance window establishes why authority was temporarily needed. Controller readback and independent checks establish what changed. Operating-mode evidence establishes that normal restrictions returned. Missing closure evidence is a specific assurance gap even when no malicious activity is detected.

Finally, response must distinguish workstation recovery from safety-system recovery. Removing the host component does not itself demonstrate that controller memory, application logic, or persistent parameters are in an accepted state. A recovery plan needs a trusted source of engineering tools and a controller-specific method of re-establishing the safety function. Restoring a backup through the same suspect tool is not an independent confirmation of restoration.

#5. Valid Protocol Messages Can Produce Invalid Process Behavior

#Use a layered acceptance model

Engineering inference. A command can pass several checks and still be unsafe for the current process state. It is useful to separate these checks explicitly:

TEXT
Accept(command, context) =
    well_formed(command)
    AND expected_origin(command)
    AND permitted_operation(command, role)
    AND fresh_for_this_transaction(command)
    AND allowed_in_this_mode(command, context)
    AND within_process_constraints(command, context)

This is an analysis model, not a statement that any one protocol implements all six predicates. A network allowlist might approximate expected origin. A secure session can authenticate an application. Application permissions can restrict an operation. Controller logic can enforce an interlock. None should be credited with predicates it does not actually evaluate.

A protocol parser normally decides whether a message conforms to its encoding and supported operations. Process authorization asks whether the operation belongs to a permitted purpose. Process safety asks whether the operation remains acceptable under current physical conditions. These may be evaluated by different components, or some may be absent. A diagram that places all three under “secure protocol” conceals the missing enforcement point.

LayerExample evidenceWhat it does not establish by itself
Transport reachabilityAllowed source/destination pathIntegrity of software at the source
Message syntaxValid lengths and supported functionAuthority to request the function
Cryptographic identityAccepted certificate or sessionSafety of the requested value
Application permissionsWrite allowed for a principalSuitability during the present operating mode
Transaction contextSequence and freshness acceptedTruth of the sensor state used for the decision
Process constraintInterlock or bounded commandIntegrity and independence of that interlock

The table is not an argument against cryptography. It identifies where cryptography ends and the process-specific policy begins. Security improves when those boundaries are explicit enough that a reviewer can ask who supplies each predicate and what evidence would falsify it.

#Classic Modbus/TCP: decoding is not authorization

The official Modbus application specification defines function 0x06 as Write Single Register, with a two-byte register address and two-byte value; protocol addressing is zero-based. The TCP implementation guide defines a seven-byte MBAP header containing transaction ID, protocol ID, length, and unit ID. The length counts bytes following the length field, including the unit identifier. These fields provide framing and transaction association, not cryptographic client identity. Modbus Application Protocol V1.1b3, §§4.2, 4.4, 6.6, TCP Implementation Guide V1.0b, §3.1.3.

The first lab constructs this request in memory:

TEXT
00 01 | 00 00 | 00 06 | 01 | 06 | 00 64 | 01 f4
 TXID | proto | length|unit|func| address| value
    1 |     0 |      6|   1|   6|     100|   500

The packet is twelve bytes long. Its six bytes after the length field consist of the unit identifier and a five-byte PDU. Address 100 is a protocol address, not a claim about a universal manufacturer's register map. Value 500 acquires meaning only through the application's documented mapping. It could represent a setpoint, a scaled measurement, or something else in a real product.

Engineering inference. In our synthetic application, address 100 is a configurable limit. A perfectly encoded request to write 900 is rejected because the permitted maintenance envelope is 400–600. The protocol did not supply that envelope; the application did. An equally valid request from a viewer is rejected because the role lacks authority. A request from an engineer during production is rejected because the maintenance operation is inappropriate for that mode.

The important result is that all five candidate messages can be syntactically valid. A malformed-packet detector would not identify the policy violations. A source-IP filter would not distinguish an approved operation from an unapproved operation issued by compromised software on an allowed workstation. A useful detector needs the application context against which the operation is evaluated.

Modbus Security is a separate specification that combines the protocol with TLS, X.509v3 authentication, and port 802. It must not be conflated with classic Modbus/TCP on port 502. Enabling it can improve identity and integrity guarantees, but deployment and application authorization still require verification. Modbus Organization, Modbus Security announcement.

#Lab 1: syntactic validity versus command authority

Download modbus_authorization.py and run:

BASH
python3 modbus_authorization.py

Python 3.10 or later is sufficient for all three labs. They use only the standard library, synthetic in-memory objects, and deterministic calculations. This script opens no socket, discovers no device, and sends no command to a controller. Its authorization principal and operating mode are explicit inputs supplied by the simulator, not fields secretly present in Modbus.

The core policy is deliberately visible:

PYTHON
def authorize(unit, address, value, role, mode):
    return (unit == 1 and address == 100 and role == 'engineer'
            and mode == 'maintenance' and 400 <= value <= 600)

Expected output:

TEXT
approved: syntax=valid tx=1 address=100 value=500 authorized=True
wrong role: syntax=valid tx=1 address=100 value=500 authorized=False
wrong mode: syntax=valid tx=1 address=100 value=500 authorized=False
out of envelope: syntax=valid tx=1 address=100 value=900 authorized=False
wrong register: syntax=valid tx=1 address=101 value=500 authorized=False
frame: 00 01 00 00 00 06 01 06 00 64 01 f4
malformed frames: rejected (3/3)

The negative framing cases test truncation, an inconsistent MBAP length, and an unsupported function for this deliberately narrow decoder. This is not a complete Modbus implementation. It accepts only the function-06 request shape needed for the experiment. In particular, it does not model TCP stream reassembly, exception responses, gateways, multiple outstanding transactions, or device-specific register behavior.

For a useful extension, change the permitted role or envelope and predict which case should pass before running it. Do not infer that adding this Python predicate to a network appliance automatically provides plant safety. In a real application, the predicate's inputs, timeliness, fail behavior, and independence would also need evidence.

#IEC-104: attacker intent inside supported operations

Observed deployment and analyzed capability. ESET reported Industroyer2's deployment in an April 2022 attempt against a Ukrainian energy provider. The payload used IEC-104 and embedded configuration; the reported attempt was thwarted. It should not be retold as an accomplished April blackout or as a proven air-gap crossing. ESET, Industroyer2, timeline and payload analysis.

Engineering inference. The lesson is about a command channel after access. A protocol-speaking payload can select operations a device already understands. Detection cannot rely exclusively on unusual ports or malformed syntax. It needs the relationship among the commanding origin, addressed process point, approved operation, and process state.

Suppose a deployment uses a command sequence with an explicit selection and execution step. Observing a correctly ordered sequence only establishes that the sender followed the expected transaction pattern. It does not establish that the selected operation was approved for the plant at that moment. A compromised permitted master can satisfy a sequence check while violating operational intent. The details of such mechanisms are implementation-specific and must be checked against the relevant device profile rather than inferred from the protocol name.

There is also a distinction between a protocol acknowledgment and a physical result. A command's acceptance may not demonstrate that an actuator moved, that it moved in the expected time, or that a separately reported state is trustworthy. An investigation needs the device response, independent feedback where available, and a time-consistent process record. Treating the first acknowledgment as proof of the final physical outcome is an avoidable analytical shortcut.

A later Mandiant report describes a distinct October 2022 Ukrainian disruption involving native OT tooling. It should remain a separate case with its own chronology, not be merged with April's Industroyer2 attempt. Mandiant, November 2023 incident analysis.

#OPC UA: security mode, identity, and permissions are different controls

OPC UA provides security mechanisms. Part 2 distinguishes None, Sign, and SignAndEncrypt, and describes SecureChannel confidentiality, integrity, and application authentication. It does not follow that every deployed endpoint uses the strongest configuration or grants appropriately narrow user permissions. OPC UA Part 2, §4.

Engineering inference. Review three relationships separately: which application certificate is accepted; which user or service identity operates inside the session; and what that identity may do to particular resources. An authenticated client application is not automatically an authorized engineering user. A trusted certificate is not automatically evidence that the software using its key remains uncompromised.

The OPC Foundation's practical guidance restricts anonymous access to non-critical resources. That is a permissions boundary, distinct from transport encryption. Historical cipher recommendations should not be copied into a current policy without checking current support and requirements. OPC Foundation practical security guidelines.

Consider an illustrative server that accepts a valid application certificate and lets a service account write a control tag. An attacker possessing the service's execution context may submit a cryptographically valid write. The SecureChannel is functioning as designed: it protects the message from other parties. The failure is that the principal's authority is broader than the intended operation, or that the trusted principal is compromised.

A semantic policy might limit writable resources, values, operating modes, and change windows. It should also identify where the state needed to evaluate those limits originates. If the mode indicator can be changed by the same account writing the tag, “maintenance only” may be a circular restriction. Independent approval and independently observed mode make a stronger claim than two variables under one mutable authority.

#INCONTROLLER: capability is not a breach narrative

Analyzed capability. Mandiant's INCONTROLLER report describes TAGRUN, CODECALL, and OMSHELL tooling aimed at OPC UA and particular controller ecosystems. TAGRUN can enumerate data and alter tag values. This is evidence about a toolset's capabilities after access, not a confirmed successful intrusion into an air-gapped plant. Mandiant, INCONTROLLER tooling overview.

The OPC Foundation subsequently emphasized that enabling and correctly configuring OPC UA security constrains the described access. A blanket statement that “OPC UA has no security” would misrepresent the evidence. OPC Foundation response to AA22-103A.

Engineering inference. Multi-vendor environments create a practical assessment problem: a single trust label such as “engineering network” can conceal very different acceptance mechanisms. One controller may require a particular credential; another may depend on operating mode; a gateway may terminate one secure protocol and forward another protocol with different guarantees. Assurance must follow each translation point. Protection of an upstream channel should not be credited to the downstream device unless the gateway actually preserves the relevant identity and policy.

#6. Wireless Sensors Create a Different Boundary

#Identify what crosses the radio boundary

An industrial wireless deployment is not automatically an Internet bridge. It can nevertheless expose a radio medium and administrative components outside the scope of an Ethernet-only isolation argument. The assessment should distinguish a field sensor, mesh participants, network manager, security manager, gateway management plane, and the application consuming sensor values.

WirelessHART trust boundaries from join-key provisioning through mesh and session keys to gateway authority and application interpretation

Emerson's technical note describes WirelessHART's AES-128-based protection, session keys, and device admission controls. That establishes architectural mechanisms, not a proof that a given plant's provisioning practices, firmware, or gateway are correct. Emerson Wireless Security technical note.

Engineering inference. Four questions should remain separate. Can an outsider receive emissions? Can that outsider interpret protected application data? Can it produce a packet the receiver accepts? Can it disrupt the medium or timing without producing an accepted application packet? Answering yes to reception does not answer yes to decryption or authenticated injection. Answering no to injection does not prove resistance to interference.

The physical RF boundary also differs from the organizational boundary. A transmitter can be inside a fenced site while parts of its signal are observable outside. Conversely, a radio range estimate does not establish exploitability: antenna position, interference, scheduling, cryptographic prerequisites, and device behavior still matter. A useful finding identifies the receiver conditions rather than presenting “wireless” as a vulnerability by itself.

#Keys have different scopes of authority

For analysis, distinguish admission authority from ongoing communication authority. A join secret participates in establishing a device's membership. Network-level protection and session-level protection have different scopes. Compromise of a secret used broadly across a deployment may have a different consequence from compromise of a secret provisioned uniquely to one device.

The 2026 SSTIC paper examines those distinctions explicitly. Its threat model assumes a previously compromised join key and the ability to receive and inject radio traffic; the experiments use Dust devices. Captured association exchanges support subsequent recovery of relevant communication keys. The reported suspension/reassociation and time-desynchronization experiments are conditional results, not a break of AES or evidence of keyless compromise. SSTIC 2026 paper, §3.1 and experiments, pp.22–27.

Engineering inference. This is the kind of result that makes a practical boundary review more precise. The finding is not simply “encryption failed.” The question is why an admission secret was available, what other authority could be derived under observed conditions, and which devices or operations fell inside the resulting scope. A control focused only on cipher strength would address a different failure mechanism.

A provisioning review should therefore record which entity generates keys, which devices receive them, where copies remain, whether replacement equipment inherits old credentials, and what constitutes evidence of removal. A key that survives in a support archive or retired gateway can remain relevant after the original device leaves service. A rotation event is meaningful only when the assessor can identify which remaining paths still accept the old authority.

#Timing is part of an industrial message's meaning

A valid value delivered too late can be unusable for control even if its confidentiality and integrity remain intact. A wireless schedule and a controller's sampling assumptions form an operational contract. If that contract is disturbed, the consumer needs to recognize missing or old observations rather than silently treating the last available value as current.

This concern is broader than deliberate RF interference. A maintenance change can alter reporting intervals, retries, or routing. A gateway restart can change the availability pattern seen upstream. The analytical task is to identify the consumer's deadline and the system's behavior when the deadline is missed. “Eventually delivered” is not equivalent to “delivered in time to support this decision.”

For a hypothetical controller, define an allowed observation age of five seconds. That number is chosen for the simulator below, not asserted as an industrial norm. If the gateway reports a value sampled ten seconds earlier but stamps it with the forwarding time, an upstream age check can pass the wrong property. The data model should distinguish source sampling time, intermediate receipt time, and consumer receipt time.

NIST's wireless field-network guidance includes encryption and device allowlisting. Those are useful controls, but an assessment still needs to check freshness handling and gateway administration in the specific architecture. NIST SP 800-82r3, Wireless Field Networks.

#Do not conflate monitoring traffic with safety authority

The consequence of compromised telemetry depends on what consumes it. A vibration sensor feeding a maintenance dashboard creates a different immediate control path from a measurement used to regulate a vessel or trigger a protective function. A path to a historian is not automatically a path to an actuator. Conversely, data initially described as “monitoring only” may influence recipes, maintenance decisions, or operator interventions over longer periods.

A consequence model should name the decision. If the value merely changes a chart, the direct result is incorrect visibility. If an operator acts on the chart, the result depends on the operator workflow and independent checks. If a controller uses the value, the result depends on the controller's logic and deadlines. If a protection function uses it, the required independence and failure behavior deserve a separate safety analysis.

This is why the article keeps sensor spoofing, RF denial of service, gateway compromise, and control manipulation separate. They can be linked in a plausible chain, but each link requires evidence. Treating the whole chain as established because the first link is technically possible would overstate the result and mislead the defender about which control must change.

#7. Authenticated Measurements Can Still Mislead Control

#Separate truth, freshness, and transport integrity

Engineering inference and illustrative model. Let h(t) be the true liquid level in a hypothetical tank, and let y(t) be the level reported to a controller. A message can be authentic in the communications sense while y differs from h. The sender may be a legitimate sensor with compromised configuration, an endpoint running malicious software, or an instrument with an incorrect calibration. Authentication identifies a source under the key assumptions; it does not measure the liquid independently.

An equally important distinction is between an old truthful measurement and a fresh false measurement. An age check can detect the former when the source timestamp is trustworthy. It cannot detect the latter from age alone. If an adversary controls both the reported value and its metadata, a “good” quality flag and current timestamp may simply accompany the false value.

Illustrative feedback loop separating the tank's true state, reported measurement, controller command, and independently observed actuator behavior

Use the following deliberately simple model:

TEXT
Tank cross-sectional area A = 1.0 m²
Outflow q_out = 0.01 m³/s
Maximum inflow q_max = 0.03 m³/s
Target level h_target = 1.0 m
Controller gain K = 0.04 m²/s
Sample interval dt = 1.0 s
 
q_in = clamp(q_out + K * (h_target - y), 0, q_max)
h_next = h + (q_in - q_out) * dt / A

Every constant is selected for this illustration. None is a value measured at an attacked facility. The units matter: multiplying a level error in metres by K in square metres per second yields a flow correction in cubic metres per second. Dividing net flow by tank area yields a level rate. The controller is proportional with outflow feed-forward; it is not an industrial PID implementation.

For an honest measurement, y equals h. At the chosen initial level, inflow balances outflow and the system remains at one metre. With a constant measurement bias b of minus 0.25 metres, y equals h minus 0.25. The controller sees a low level and adds inflow until the reported value approaches the target. The actual equilibrium approaches 1.25 metres, while the reported equilibrium approaches one metre.

This simple result isolates a specific failure: a trusted measurement path can translate a false observation into a physically wrong operating point without any malformed command. The control law itself can be implemented correctly. Transport authentication can work correctly. The problem is that the input does not represent the physical state that the law assumes.

#Stale does not always mean immediately dangerous

The stale scenario holds the initial measurement and its original sample time. At the selected operating point, the controller's calculated inflow still balances the fixed outflow. The physical level remains unchanged even though the observation is too old after the configured deadline. This is intentional: it prevents the example from suggesting that every stale value necessarily produces an immediate physical excursion.

The deficiency is loss of current knowledge. A disturbance would expose it. If outflow changed while the sample remained frozen, the controller would continue acting on the original observation. To interpret the result properly, separate “no damage in this trajectory” from “valid basis for future control decisions.” A test that never perturbs the system can miss the latter failure.

The same distinction applies to stable-looking plant trends. A displayed line that remains constant might indicate steady operation, a stale source, or a filtered signal. The interpretation depends on source time, data quality, expected process variation, and independent context. Constancy is a property of the reported sequence, not proof of an unchanged physical system.

#Lab 2: fresh bias versus stale truth

Download control_loop.py and run:

BASH
python3 control_loop.py

Expected output:

TEXT
honest: final=1.000m peak=1.000m stale=0 trip=False
stale: final=1.000m peak=1.000m stale=74 trip=False
bias: final=1.240m peak=1.240m stale=0 trip=False
bias+independent-trip: final=1.201m peak=1.201m stale=0 trip=True

The simulation executes 80 one-second steps. It marks a sample stale when age is strictly greater than five seconds; steps 6 through 79 therefore produce 74 stale events. The biased case uses current source times and produces none. Its final level is below the theoretical 1.25-metre equilibrium because the simulation stops after a finite interval.

The independent-trip case uses the true level as an idealized separate observation and stops the simulation at the next sampled threshold crossing. The chosen threshold is 1.20 metres. The reported peak slightly exceeds it because the model observes discretely. This example stops both flows and ends the simulation; that behavior is not a recommended generic safety action. Real trip actions depend on the process, actuator dynamics, and a validated safety design.

The comparison demonstrates what an independent observation could contribute, not that an ideal sensor can be purchased or that this script calculates a safety integrity level. It omits sensor noise, delays, valve dynamics, controller scheduling jitter, tank geometry changes, failure rates, and uncertainty bounds. Those omissions are explicit because adding realistic complexity without validated parameters would create an appearance of precision rather than stronger evidence.

To investigate the stale case further, modify the synthetic outflow after a chosen step and compare true level with reported level. To investigate detector sensitivity, introduce a smaller bias and compare it with a hypothetical measurement uncertainty band. Keep the question concrete: does the proposed check distinguish the condition of interest under the stated assumptions?

#Independent observations must be independent in the relevant way

Two displayed values may share one transmitter, one gateway, or one mutable scaling table. Agreement between them then provides little protection against failure of that shared component. Independence has to be evaluated along the path of the particular threat: acquisition, configuration, transmission, interpretation, and authority to alter each stage.

A physically separate sensor can still share an engineering account or configuration package with the first. Two historians can consume the same gateway output. Two dashboard tiles can read the same cached value. None of these facts automatically invalidates the architecture, but they limit what agreement can establish. The assessment must identify which failures or malicious changes remain common to both observations.

A model-based residual is another observation, but its evidentiary value depends on its inputs. If a detector predicts level from the same compromised flow values that an attacker can modify, the residual may remain small by construction. A useful residual uses constraints or observations the attacker cannot simultaneously control under the stated threat model. That might be independent actuator feedback, a separate balance measurement, or a bounded physical relationship.

False positives also have physical consequences. If a detector's chosen response interrupts a process, its uncertainty and failure behavior need engineering review. The correct security outcome is not “trip on every suspicious packet.” It is a documented decision rule whose sensing assumptions, deadlines, operating modes, and consequences are understood and tested.

#8. The Trusted System Changes Throughout Its Lifecycle

#Stable operation is not immutable state

Engineering inference. “No static system is possible” is too absolute. A device can operate unchanged for a long period, and an industrial process can remain near a steady operating point. The narrower and more useful claim is that long-term assurance must account for legitimate transitions. Maintenance, calibration, replacement, certificate renewal, and restoration may alter the state or the authority available to alter it.

The IAEA explicitly discusses increased exposure during maintenance, temporary interfaces, and computer-based testing or calibration tools. Its recommendations concern nuclear facilities; applying the lifecycle reasoning to other OT environments is an engineering analogy, not an assertion that every requirement has identical scope. IAEA NSS 17-T Rev.1, §§6.33–6.35.

A baseline should identify what is being held stable. Consider this illustrative state vector:

TEXT
S = {
    engineering_project_bytes,
    running_logic_identity,
    retained_parameters,
    field_calibration,
    operating_mode,
    temporary_bypasses,
    active_trust_and_credentials
}

Each element needs a collection method and an authority model. A filesystem hash can directly establish equality for a file, assuming the collector is trustworthy. It cannot directly observe a transmitter's physical calibration or prove that the controller executes the corresponding compiled application. A file manifest is useful precisely when its scope is explicit.

#Configuration changes can preserve the file hash

Suppose the approved engineering archive is unchanged, but a retained limit is altered online. A subsequent comparison of the archive with its approved hash can pass. That result is accurate: the archive is unchanged. It does not establish that the operational limit equals the archive's intended value.

The same distinction applies to field scaling. An unchanged controller application can interpret a measurement differently after a transmitter configuration or gateway scaling change. A certificate renewal can preserve application logic while changing accepted identities. A replacement device can restore nominal function while introducing different firmware or default credentials. These are separate state transitions, not necessarily malicious ones, but they belong inside the security argument.

A useful change record states both the intended difference and the invariants that must remain true. For a calibration operation, the intended difference may be an offset, while controller logic and permissions must remain unchanged. For a firmware update, firmware changes while accepted configuration, process behavior, and trust relationships require explicit verification. A generic “maintenance completed” entry does not expose those distinctions.

#Maintenance is a transaction with an exit condition

Model maintenance as a controlled transaction: establish preconditions, open necessary authority, perform the approved change, verify the resulting state, close temporary authority, and accept return to service. If verification fails, the transaction needs a defined alternative such as rollback, further diagnosis, or continued restricted operation under the site's engineering procedures.

Lifecycle transition from approved operation through maintenance and changed state to independently verified restoration

The exit condition is not merely a successful tool message. It is evidence that the accepted operational state has been reached. That can include controller readback, parameter verification, device identity, restored mode, removed bypasses, and a functional check. The actual combination depends on the changed component and its consequences.

Authority should also be restored, not just functionality. A temporary programming permission, support account, or maintenance route can survive the technical repair. If the workflow closes the work order without revoking that authority, the system's future exposure differs from the pre-maintenance state. A successful repair and a failed security closure can occur in the same transaction.

#Lab 3: a file check can pass while restoration fails

Download maintenance_integrity.py and run:

BASH
python3 maintenance_integrity.py

Expected output:

TEXT
approved: file_match=True trusted=True
maintenance: file_match=True trusted=False
naive return: file_match=True trusted=False
verified restoration: file_match=True trusted=True
independent state checks: passed (3/3)

The simulator changes a retained limit from 600 to 900 while keeping the approved project hash constant. It also enables a temporary bypass during maintenance. A naive return clears the bypass and switches the mode back to production but fails to restore the limit. File comparison still passes; the state comparison correctly fails.

Restoring the limit makes the illustrative state satisfy its acceptance predicate. Additional assertions show that changing running logic, calibration, or bypass state independently invalidates the predicate. These checks are deterministic comparisons over synthetic objects, not evidence from actual hardware. In particular, the script assumes trustworthy observations of every field. It does not solve the collection or attestation problem that a real deployment must address.

The experiment's value is the counterexample: equal project hashes do not logically imply equal operational state. It does not require malware to demonstrate that gap. A legitimate but incompletely closed maintenance change is enough. A malicious change could exploit the same missing check, but that is an inference about the model rather than a new historical incident.

#Recovery can reintroduce the state being removed

A backup has a capture time, a source, and an acceptance basis. “It restores successfully” is not equivalent to “it predates compromise” or “it contains the intended parameters.” The evidence should distinguish software authenticity, configuration correctness, and operational suitability for the present equipment.

Restoration can also involve translations. An old project opened in a newer tool may undergo conversion. A replacement controller may require a different firmware package. A replacement instrument may have a different device profile. A certificate chain may have expired since the archive was created. These transitions need review because the recovered artifact is no longer necessarily the exact object previously accepted.

For an isolated environment, the tools and verification data needed for recovery must be available through a controlled workflow. Otherwise an emergency can create pressure to import unverified utilities or enable ad hoc connectivity. Preparation is an integrity control: it reduces the number of unknown authority transitions required during a time-critical restoration.

#9. Build an Assurance Case

#Start from a claim that evidence can falsify

Engineering inference. A useful assurance statement might be: “During approved production operation, only the accepted controller application and parameter set can govern this process; changes require a bounded engineering transaction whose result is verified through independent evidence.” That claim is much more demanding than proving the absence of a network route, but it names the property that matters.

It can be decomposed into supporting claims: execution originates from an accepted toolchain; change authority is restricted and observable; imported artifacts have accountable provenance; running state is checked; measurement interpretation is bounded; temporary exceptions are closed; and recovery reconstructs an accepted state. Each supporting claim should have concrete evidence and declared limitations.

A single control rarely proves a whole supporting claim. Signing provides provenance under a key-trust assumption. Malware scanning provides a result under a detection model. An engineering comparison provides a result under a tool-integrity assumption. A functional test provides observations for selected conditions. Combining them can strengthen assurance, provided their dependencies are understood rather than silently assumed independent.

Failure mechanismEvidence to seekResidual question
Transported artifact crosses trust domainsProducer identity, manifest, transfer record, consumer behaviorCan the receiving parser or trusted producer be compromised?
Engineering authority changes logicWork order, bounded privilege, download record, readbackIs verification independent of the altered toolchain?
Valid operation violates process intentPrincipal, resource permissions, mode, bounded valuesAre the policy's context inputs trustworthy and timely?
Wireless provisioning exposes authorityPer-device provisioning record, key scope, removal/rotation evidenceWhich old devices or archives retain usable secrets?
Fresh authenticated value is falseIndependent observations and physical constraintsWhat shared configuration can corrupt all observations?
Maintenance leaves altered stateExit checks for parameters, modes, bypasses, accessDoes restored functionality mask an unclosed exception?
Recovery restores contaminated artifactsTrusted tools, accepted backup, controller-specific checksDoes the backup's acceptance basis still apply?

#Controlled transfer needs an acceptance point

The receiving side should decide what authority an artifact receives before it is interpreted by a high-privilege consumer. The acceptance record should connect the approved artifact to the exact deployed output. If a project is rebuilt, dependencies and build outputs belong in that connection; the hash of the original archive alone does not cover newly generated objects.

This does not mean a universal content-sanitization step can make every industrial project safe. A project format may contain necessary logic that cannot be stripped without changing functionality. A firmware image may be opaque to the receiving organization. The appropriate control depends on the artifact's meaning and the verification available. The assurance case should state those limits rather than promise an inspection capability the organization does not possess.

A dedicated transfer station is similarly a component to assess, not a magical boundary. Its update channel, detection engine, removable-media handling, and logs affect the evidence it produces. If the transfer station is compromised, its “passed” result may be unreliable. Independence from the engineering workstation can improve the design, but the station still requires its own trustworthy lifecycle.

#Independent readback is a property, not a product label

Readback is valuable when it observes the relevant running state through a path whose failure is not identical to the path under investigation. A second software window on the same compromised station is not automatically independent. A second tool using the same altered library may share the same failure. Even an independent byte read requires correct interpretation and an accepted reference.

The independence requirement should match the suspected failure. If the concern is a replaced communication library, a collection path that avoids that library is relevant. If the concern is controller firmware that lies about its state, a normal engineering readback may not suffice. If the concern is field calibration, application readback alone cannot resolve it. The point is to choose evidence based on the hypothesis rather than adding redundant screenshots.

#Data diodes constrain a direction, not every source of change

A unidirectional gateway can substantially constrain communication in its enforced direction. Its assurance claim belongs to a defined placement, interface set, and data flow. It does not automatically govern inbound removable media, local engineering actions, or a separate management path. The physical enforcement mechanism and any protocol mediation should be reviewed as implemented, not inferred from a marketing diagram.

Outbound-only data can also carry an integrity problem to its consumer. A diode can faithfully transmit false telemetry from a compromised source. That does not make the diode defective; truth of the source measurement is outside the directionality property. The consumer still needs to understand the authenticity, freshness, and consequence of the data it receives.

An OT architecture may therefore combine strong unidirectional export with a tightly controlled inbound maintenance process. Those controls address different needs. The coherent assurance case explains both rather than claiming the export control eliminated all inbound change paths.

#Test transitions as well as steady operation

An assessment should include production, maintenance, degraded operation, recovery, and return to service. The critical observation is often at a transition: when access opens, when a bypass activates, when a device rejoins, when a certificate changes, or when the system accepts a restored state. A test performed only in stable production can miss the interval in which the trusted configuration is actually mutable.

A useful exercise begins with a specific invariant and attempts a benign violation in an isolated replica. For example: change a retained parameter while preserving a project file, or delay a synthetic observation while leaving its value plausible. The expected result is not merely an alert. It is an evidence trail showing which component rejected the condition, which authority approved any exception, and which checks supported restoration.

The three labs in this article demonstrate the reasoning pattern at deliberately small scale. They do not model a whole plant, establish universal thresholds, or justify production changes. They make the missing predicates visible: syntax is not authorization, freshness is not truth, and file equality is not operational-state equality. Those distinctions remain useful when the implementation becomes more complicated.

Air-gapping is strongest when its claim is narrow enough to verify and embedded in a lifecycle that controls every permitted source of authority. The engineering task is to preserve that narrow benefit while proving the additional integrity properties the process actually needs. A plant is not secure because an attacker lacks a continuous route. It becomes more defensible when every path from an external artifact or compromised principal to physical behavior has an explicit boundary, an acceptance rule, and evidence that the resulting state is the intended one.

Found it useful? Share it
← Back to all posts