Predictive Maintenance

Why Factory Cybersecurity Must Include A Recovery Plan

Photo by Sasun Bughdaryan (@sasun1990) on Unsplash

A cyberattack on a factory does not end when the malicious software has been removed from the network. Production may still be suspended, control systems may no longer be trusted, and engineers may be unable to confirm whether machine configurations, safety parameters or production recipes have been altered. Even when corporate IT has restored email, files and business applications, the plant can remain offline because operational technology follows a different logic. This is why factory cybersecurity cannot be built around prevention alone. Firewalls, network segmentation, access controls and employee training remain essential, but they do not answer the question that matters once an incident reaches production: how does the company restore operations without creating a second failure through an unsafe or premature restart?

For manufacturers, recovery is not simply a matter of bringing servers back online. It requires confidence that physical processes, machine behaviour and safety functions remain intact. A production line restarted with corrupted control logic may appear operational while producing defective goods, damaging equipment or exposing employees to danger. The pressure to resume production quickly can therefore conflict with the need to verify that the plant is safe. A credible recovery plan resolves that conflict before an attack occurs.

Factory Recovery Is Different From IT Recovery

Traditional disaster recovery is usually organised around systems, applications and data. A company identifies which services are critical, maintains backups and establishes the order in which they should be restored. The same principles apply in manufacturing, but operational technology adds another layer because the systems control physical equipment.

A programmable logic controller does not merely store information; it governs the sequence in which a machine moves. A production recipe can determine temperatures, pressures, speeds and tolerances. A safety controller may prevent a worker from entering a hazardous area while equipment is active. Restoring these systems incorrectly can have consequences far beyond lost data.

The manufacturing environment is also highly interconnected. A single line may depend on controllers, drives, sensors, vision systems, industrial networks, manufacturing execution software and interfaces with enterprise planning. Restoring one component does not guarantee that the process will function correctly, particularly when system versions, configurations and communication settings no longer match.

This makes recovery an engineering challenge as much as a cybersecurity task. IT teams may understand the attack and the affected infrastructure, while production engineers understand which systems must be restored together and what needs to be checked before equipment can move again. Neither group can manage the process effectively without the other.

The Fastest Restart Is Not Always The Safest

When production stops, commercial pressure arrives immediately. Orders are delayed, materials accumulate and customers begin asking when deliveries will resume. Management naturally wants the shortest possible outage, yet this urgency can encourage teams to restore systems before they understand the full scope of the incident.

A controller that appears functional may still contain altered logic. A backup may be available but no longer represent the approved production configuration. An attacker may have changed user accounts, remote-access settings or alarm thresholds without disrupting the system visibly.

The danger is not limited to deliberate sabotage. Recovery work performed under pressure can introduce new errors. Engineers may load the wrong programme version, overlook a local modification or reconnect a machine before the surrounding systems are ready.

The recovery plan should therefore define the conditions under which each part of the factory can return to service. These conditions may include technical validation, safety checks, confirmation of production parameters and approval from named responsible functions. A restart decision should follow evidence rather than the assumption that an apparently quiet system is a clean one. The aim is not to delay production unnecessarily. It is to prevent a rushed restart from creating equipment damage, quality failures or a second shutdown that lasts even longer.

Known-Good Configurations Matter More Than Ordinary Backups

Factories often maintain backups of controllers and machine programmes, but the existence of a file does not prove that it is current, complete or safe to restore. A backup can be copied after a configuration has already been compromised, while older versions may omit legitimate changes introduced during maintenance or process improvement.

Recovery therefore depends on maintaining known-good configurations: verified versions of control logic, safety settings, machine parameters and network configurations that the company can trust.

These records need clear ownership and version history. Engineers should know when a configuration was approved, which machine it belongs to and which hardware and software versions it requires. Changes made directly on equipment must be captured rather than remaining undocumented local knowledge.

Offline or isolated copies are particularly important because backups connected continuously to the same environment can be encrypted or corrupted during an attack. The company should also test whether the files can actually be restored. A backup that has never been loaded onto replacement hardware provides less assurance than it appears to offer.

For legacy machines, the challenge may include obsolete software, unavailable licences or hardware no longer supplied by the manufacturer. Recovery planning should identify these dependencies before an emergency reveals them.

Recovery Priorities Must Follow Production Dependencies

Not every system can be restored simultaneously, and the sequence matters. A company may be tempted to begin with the most visible production line, even though that line depends on supporting systems that are not yet available.

A recovery plan should map the dependencies between utilities, safety systems, networks, machines and production software. Compressed air, power management, cooling or material-handling systems may need to return before the line itself. Identity services and time synchronisation may be necessary before controllers and applications can communicate reliably.

The order should also reflect business priorities. A plant may choose to restore one strategically important product line before other operations, but that decision must be compatible with technical and safety requirements. Management can define which products and customers matter most, while engineering determines what can be brought back safely and in what sequence.

This dependency map should be detailed enough to guide action without becoming so complicated that no one can use it during an incident. The most effective plans identify critical systems, prerequisite services, responsible teams and the evidence required before moving to the next stage.

Manual Operation Needs To Be Planned, Not Improvised

Some factories can continue limited production manually or through simplified operating modes when digital systems are unavailable. This can reduce the impact of an attack, but only when the alternative process has been designed and practised in advance.

Employees need to know which tasks can be performed safely without normal automation, which quality checks must be added and how production data will be recorded. Manual procedures may require more people, slower speeds and different supervision.

The process also needs clear limits. A machine designed for automated operation may not be suitable for prolonged manual control, while employees who no longer perform certain tasks regularly may lack the necessary confidence or experience.

Companies should resist treating manual operation as a universal fallback. In some environments, stopping production remains the safest decision. The recovery plan should distinguish between processes that can continue in degraded mode and those that require full restoration before work resumes.

IT, OT, Safety And Production Need One Incident Structure

Cyber incidents often expose organisational divisions that were manageable during normal operations. IT may focus on containing the attacker, OT teams on restoring equipment, production managers on output and safety specialists on risk. Without a shared command structure, these priorities can conflict.

A factory recovery plan should establish who leads the response, who has authority to isolate systems and who approves the restart of equipment. It should define how decisions are communicated across shifts and sites, particularly when the incident affects several facilities.

External partners may also be essential. Machine builders, automation integrators and software vendors often hold specialised knowledge that the manufacturer does not possess internally. Their contact details, access procedures and contractual responsibilities should be prepared before the incident rather than negotiated while production is already down.

The company should also plan how remote vendor access will be handled during recovery. Reopening external connections too early can undermine containment, while blocking them completely may prevent engineers from receiving the support they need.

Recovery Testing Must Go Beyond A Tabletop Exercise

Many organisations discuss cyber scenarios in conference rooms, which helps clarify roles and communication. Factories also need practical tests that examine whether systems can be restored under realistic conditions.

A recovery exercise may involve loading a known-good controller programme, rebuilding an industrial workstation or restoring a production application in an isolated environment. Teams can test whether documentation is accurate, replacement hardware is available and passwords or licences can be retrieved when needed.

The exercise should include operational verification rather than ending when the software starts. Engineers need to confirm that sensors, alarms, machine movements and safety interlocks behave as expected. Quality teams may need to inspect initial products before the line returns to normal output.

Testing often reveals dependencies that documentation missed. A restored application may rely on an old driver, a machine may require a particular engineering laptop, or a safety system may need a specialist who is not available outside working hours.

These findings are valuable because they convert an assumed capability into a real one. Recovery plans improve through rehearsal, not through confidence alone.

Quality Control Is Part Of Cyber Recovery

When a cyber incident affects production systems, the company must consider whether products manufactured before detection can still be trusted. An attacker may have altered process settings gradually, while corrupted data may have affected traceability or inspection results.

The recovery process therefore extends beyond the machines. Quality teams may need to quarantine inventory, review production records and determine which batches require additional inspection. If traceability data is incomplete, the company may have to apply broader containment than the technical evidence alone would suggest.

The first products manufactured after restart also deserve special attention. A machine can pass a basic functional test while still operating outside the tolerances required for consistent quality. Controlled ramp-up, additional sampling and engineering supervision help identify problems before normal production resumes.

This connection between cybersecurity and quality is easy to overlook when the incident is treated mainly as an IT outage. In manufacturing, system integrity and product integrity are closely linked.

Recovery Plans Need A Clean-Environment Strategy

Restoring compromised systems into the same network without sufficient isolation can allow an attacker to regain access or spread into newly rebuilt equipment. Manufacturers need a clean environment in which systems can be examined, restored and tested before they return to production.

This may involve separate recovery networks, isolated engineering stations and controlled pathways for transferring approved software and configurations. Devices should be scanned and verified before connection, while access credentials compromised during the incident must be replaced.

The company also needs to decide how much of the environment should be rebuilt rather than repaired. Reinstalling a workstation from a trusted image may provide greater assurance than attempting to clean it, but older industrial systems can be difficult to recreate without affecting compatibility.

The strategy should balance confidence, time and operational practicality. The decision becomes much easier when replacement images, configuration files and hardware inventories already exist.

Legacy Equipment Requires Special Attention

Factories often operate machinery for decades, long after the associated software and operating systems have left mainstream support. These assets can be difficult to patch and even harder to recover.

A legacy machine may depend on a specific computer, cable, licence key or engineering application that is no longer readily available. The organisation may have only one employee who remembers how the system was configured.

Recovery planning should identify these fragile dependencies and preserve the tools needed to restore them. This may include spare hardware, archived installation media, offline documentation and relationships with specialist service providers.

Where replacement is not immediately possible, the company can reduce exposure through segmentation, restricted access and close monitoring. None of these measures removes the need to understand how the asset would be recovered after failure or compromise.

The oldest machine in the factory can become the longest part of the recovery process.

A Recovery Plan Must Include Business Decisions

Technical teams can restore systems, but management must decide which level of disruption the company can tolerate and what trade-offs it is prepared to make.

Should the plant restart at reduced capacity while some systems remain unavailable? Can products be manufactured without full digital traceability? Which customers should receive priority when output resumes? At what point should the company notify regulators, insurers or business partners?

These decisions are difficult to make during a crisis because the information is incomplete and the financial pressure is high. A recovery plan cannot predetermine every answer, but it can establish principles, thresholds and decision authority.

The company should also define what information senior leadership needs. A useful status report distinguishes between systems that are technically restored, processes that are operationally validated and production that has returned to acceptable quality. These are different stages, even though they are sometimes presented as one.

Recovery Readiness Should Influence Technology Investment

A new industrial system is often evaluated according to productivity, integration and purchase price. Recoverability deserves equal attention.

Manufacturers should ask whether configurations can be exported, how quickly replacement hardware can be obtained and whether the vendor supports offline backups and controlled restoration. Proprietary systems that cannot be rebuilt without one external provider may create a hidden continuity risk.

The same principle applies to connected platforms and industrial cloud services. The company needs to understand what happens when the service is unavailable, how data can be retrieved and whether local operations can continue in a degraded mode.

Recovery requirements should be included during procurement rather than added after the technology becomes critical to production.

The Plan Must Remain Current

Factories change constantly. Machines are upgraded, lines are reconfigured and software versions evolve, which means that a recovery plan can become obsolete even when it was accurate at the time it was written.

Ownership is therefore essential. Named teams should review configurations, dependencies and contact details regularly. Significant production changes should trigger an update, while recovery exercises should confirm that the plan still reflects reality.

The documentation also needs to remain accessible during an incident. A plan stored only on the compromised network may be unavailable at the moment it is needed most. Secure offline or isolated copies ensure that teams can still reach procedures, inventories and contact information. Recovery readiness is not a document completed for an audit. It is an operating capability that has to follow the factory as it changes.

Resilience Begins Before The Attack

Manufacturers cannot eliminate every cyber risk, particularly as factories become more connected and dependent on software. They can, however, reduce the uncertainty that follows an incident. A strong recovery plan establishes which systems matter most, how trusted configurations will be restored and who decides when production can resume. It brings IT, OT, production, quality and safety into the same process, while recognising that a successful restart requires more than functioning computers. The central question is not how quickly the company can switch the machines back on. It is how quickly it can return to safe, stable and trustworthy production. Prevention reduces the likelihood of an attack reaching the factory floor. Recovery planning determines whether the company can survive when prevention is not enough.

  Why Factory Cybersecurity Must Include A Recovery Plan