The server went down at 7:43 on a Monday morning. The ERP won't start. The production line has stopped. The shift supervisor is waiting. How long can your company hold on like this — and do you know, to the minute, how long it takes to restore everything on a clean machine?

This article makes an uncomfortable case: most Portuguese industrial SMEs don't have a backup problem. They have a recovery problem. And the choice between cloud and on-premise is secondary as long as the restore process has never been rehearsed.

The distinction nobody makes — and that decides everything

Copies get made. Sometimes. But last month's backup file may be corrupted. The ERP may need a specific restore sequence that only the vendor knows. The database may have dependencies that the snapshot doesn't capture. None of these problems show up while there's no incident — and they all show up at once when there is.

The 2025 Verizon Data Breach Investigations Report documents that ransomware was present in 88% of data breaches at SMEs, against 39% at large organisations. SMEs are the disproportionate target precisely because they have less-tested backups and longer recovery times. Storage architecture matters less than knowing, to the minute, how long it takes to get back up and running.

That said: architecture does matter. Let's get to it — with concrete criteria, a decision matrix and an audit checklist ready to use.

What you need to have defined before choosing any architecture

Without this inventory, any technical decision is arbitrary. For each critical system — ERP, MES, invoicing files, customer database, work-in-progress production data — define two numbers:

RTO (Recovery Time Objective): how much downtime is acceptable. For the ERP of a garment factory with 120 employees, four hours of downtime on a Monday at month-end has a very different operational cost from four hours on a Friday afternoon. The RTO has to be approved by management, not merely estimated by IT.

RPO (Recovery Point Objective): how many hours of data you can lose without serious damage. If the last confirmed order before the incident is not recoverable, what is the impact? Answering this question forces the CEO and CFO to sit down at the table with the IT lead — and that's where the real numbers appear.

To these two, add: the system dependency map (the ERP depends on the database server, which depends on the storage, which depends on the UPS), the physical and digital location of the licences and activation keys — kept away from the main server —, the available bandwidth measured at peak production hour, and a three-year TCO budget, not just the initial acquisition cost.

Step 1 — Classify data by real criticality, not by intuition

A garment factory in the Vale do Ave with 120 employees has three distinct data categories. Treating them with the same backup policy is the most common mistake we see in industrial infrastructure projects.

Critical data — ERP database, invoicing files, work-in-progress production data — requires an RTO of under four hours and continuous replication with an RPO of minutes. Important data — order history, technical sheets, HR data — tolerates an RTO of up to 24 hours. Archive data — documentation of closed projects, historical tax-compliance backups — can live on tape or cloud cold storage at minimal cost, with an RTO over 72 hours.

The classification has to be reviewed annually and validated by management. We regularly see situations where IT classifies a system as "important" and the CEO, when confronted with the real cost of 24 hours of downtime, immediately reclassifies it as "critical". That conversation has to happen before the incident.

Step 2 — Assess the network with arithmetic honesty

Cloud backup presupposes connectivity. A company with a 100 Mbps connection shared between the office and the shop floor will feel continuous replication. The calculation is straightforward: 500 GB of new data per day requires, over eight hours of continuous transfer, about 139 Mbps dedicated. If the line can't handle that without degrading the ERP, pure cloud is not viable without network upgrades.

Hybrid cloud solves this problem. The primary backup stays on-premise — fast, with no bandwidth consumption — and replication to the cloud happens outside peak hours. It's the architecture we most often see work in Portuguese industrial facilities with 50 to 300 employees. Also check whether you have an uptime SLA guaranteed by the ISP (minimum 99.5%) and whether there is a failover connection via 4G or 5G for when the fibre goes down.

Step 3 — Compare the architectures with real criteria

Criterion On-Premise Cloud Hybrid
Typical RTO 2–8 hours (local restore) 1–4 hours (depends on bandwidth) 30 min–2 hours
Typical RPO 1–24 hours (nightly backup) Minutes (continuous replication) Minutes (local replication) + cloud
Initial cost High (hardware, licences) Low (monthly OPEX) Medium
3-year cost Predictable, amortisable Grows with data volume Controllable if well sized
Network dependency None for local restore Total Partial
GDPR compliance Full control of location Requires DPA contract + EU region Requires DPA contract + EU region
Ransomware protection Weak if backup on the same network Good (natural air gap) Very good (air gap + local speed)
Internal management required High Low to medium Medium
Suitability for PT industrial SMEs Companies with dedicated in-house IT Companies without dedicated IT Most industrial SMEs

A note on the egress cost that almost no contract highlights in the initial proposal: transferring large data volumes back from the cloud for a restore has a per-GB cost with most providers. Calculate this cost in the worst-case scenario — a full restore of the ERP database — before signing. In infrastructure projects we've supported, this surprise cost has reached triple the projected monthly cost in the month of a serious incident.

Step 4 — Implement the 3-2-1-1 rule

The classic rule was 3-2-1: three copies, on two different media, one off-site. With ransomware encrypting network-connected backups, the current version is 3-2-1-1 — it adds one immutable copy that cannot be altered or deleted during a defined period, typically 30 to 90 days. Most cloud providers offer object lock or WORM storage for this purpose.

In practice, for a Portuguese industrial company with MULTI ERP or another vertical ERP, the four copies are distributed as follows: a local snapshot on a dedicated NAS, outside the production domain; replication to a second site or backup appliance with deduplication; cloud replication with 30-day retention and object lock active; and the immutable copy in a separate account, not accessible via the day-to-day administrator credentials.

The retention period has to cover the average ransomware detection time, which is around 20 to 30 days. A 15-day backup isn't enough if the malware got in three weeks ago and has only now encrypted the files.

Step 5 — Test the restore. Now. Not tomorrow.

A backup that has never been tested is not a backup. It's a hope stored on a disk.

The test is not verifying that the file exists. The test is restoring on an isolated machine and confirming that the ERP starts, that the database is consistent, and that the actual restore time fits within the defined RTO. There's an operational detail the manuals rarely mention: many vertical ERPs for the textile or footwear industry have licence dependencies tied to the server's hostname or MAC address. A restore on a different machine can fail at activation even with all files intact. Confirm this point with the ERP vendor before the incident — not during.

Set a schedule: monthly for critical systems, quarterly for the rest. Document each test with date, system, volume restored, elapsed time and anomalies found. This record is also evidence of compliance for GDPR audits and, if your company falls under the NIS2 Directive transposed in Portugal by Decree-Law No. 65/2025, for the risk management and service continuity obligations.

Step 6 — Document the plan in operational language

The continuity plan can't be 40 pages nobody reads. It needs a laminated A4 sheet in the server room and another on the IT lead's phone. With three sections: what happened (a list of scenarios — ransomware, hardware failure, fire, ISP failure — with the first step for each); who does what (names, mobile numbers, responsibilities with no ambiguity); where the access is (the physical and digital location of recovery credentials, licence keys, ERP vendor support contacts).

We regularly see in infrastructure projects that the plan exists in a PowerPoint on the IT manager's computer — who is on holiday when the server goes down. The documentation has to be accessible to whoever is present, not to whoever wrote it. This seems obvious. It isn't: in more than half of the recovery projects we've supported after an incident, access to the restore credentials was blocked by a dependency on a single person.

Mistakes that keep repeating — and how to break the pattern

Backup on the same server as the production data. A ransomware attack encrypts everything on the same network. Separate physically or use a cloud air gap with object lock. There is no middle ground here.

RTO defined by IT without management validation. IT says "we restore in four hours". The CEO doesn't know that four hours of downtime on a Monday at month-end has a concrete cost in stopped production, delayed orders and contractual penalties. Align the numbers and the cost before the incident — not after.

Cloud contracts without a data location clause. GDPR requires that European citizens' data be processed on servers with adequate guarantees. Require a Data Processing Agreement and confirm that the storage region is the EU. Some cloud providers place backups in regions outside the EU by default, to reduce latency or cost. Check.

Ignoring prevention and focusing only on recovery. Backup is recovery. Prevention reduces the frequency with which you need to recover. See perimeter protection and SOC to understand how the two layers complement each other — and why a poorly configured firewall invalidates the best backup policy in the world.

The regulatory context you can't ignore

The NIS2 Directive (Directive (EU) 2022/2555), transposed in Portugal by Decree-Law No. 65/2025, requires medium and large companies in critical sectors — including manufacturing — to implement risk management measures that explicitly include backup and business continuity policies. Non-compliance can result in fines of up to 10 million euros or 2% of global turnover.

The CNCS recorded 2,758 cybersecurity incidents in Portugal in 2024, a 36% increase on 2023, with about 78% occurring in private entities (source: CNCS, Cybersecurity in Portugal Report, 2024). For an industrial SME in the North of the country, this is not an abstract statistic — it is the context in which the next incident will happen. The question is not if, it's when.

To understand what NIS2 means in factory practice, see NIS2 in practice: technical controls for the factory in the first 90 days. And if your company operates with warehouse management or real-time production capture, the data generated by these systems has to be covered by the same backup policy — not just the central ERP.

Quick decision matrix

Company profile Recommended architecture Immediate priority
Industrial SME, no dedicated IT, <50 employees Cloud managed by a partner Define RTO/RPO and test the restore
Industrial SME, in-house IT, 50–200 employees Hybrid (local NAS + cloud with object lock) Implement the 3-2-1-1 rule and schedule tests
Industrial company, structured IT, >200 employees Hybrid with a second site + cloud NIS2 audit and formal continuity plan
Retail with multiple stores and POS Centralised cloud + local backup per store Ensure the POS works offline in the event of a network failure

The remaining question: if the server went down right now, who in your company would know exactly what to do in the first 15 minutes — and would have access to everything they need to do it?

Frequently asked questions

What is the difference between RTO and RPO and why do both matter?

RTO (Recovery Time Objective) is the maximum acceptable downtime. RPO (Recovery Point Objective) is the maximum amount of data you can lose. An ERP may have an RTO of 4 hours (acceptable downtime) but an RPO of 15 minutes (it can't lose more than 15 min of orders). Defining them precisely avoids wrong technical decisions and unnecessary costs.

Why does on-premise backup fail against ransomware?

If the backup is on the same network as the main server, ransomware accesses both simultaneously. By the time you discover the incident, the backup file is already encrypted. The solution is an air gap (physical disconnection) or cloud with immutable retention. Without this, you have copies you can't use when you need them.

Is pure cloud viable in a factory with a shared 100 Mbps connection?

No. Continuous replication of 500 GB daily requires ~139 Mbps dedicated. If the line is shared with the ERP, it degrades. The solution is hybrid cloud: fast local backup (on-premise) and replication to the cloud off-peak. It's the architecture that works best in Portuguese industrial SMEs.

What is the real cost of 24 hours of ERP downtime?

It depends on the activity. A garment factory with 120 employees loses about 2,400 hours of direct work, plus the impact of delayed orders and contractual fines. The CEO and CFO should quantify this before choosing an architecture. That conversation defines whether the system is "important" or "critical" — and changes everything.

How do you know if the backup is actually functional?

Only by testing a full restore on a clean machine, without access to the original data. Most SMEs never do this. An annual test of each critical system is recommended, with a stopwatch, to validate the real RTO. If the test fails, the backup doesn't exist — you only have files you can't use.

Tape or cloud for tax-compliance archiving?

Tape is cheaper in the long run if access is rare (tax compliance requires 10 years). Cloud cold storage is more practical if you need occasional access. The criterion is: if you're never going to restore in under 72 hours, tape pays for itself. If you might need it within days, cloud is safer and more traceable.

Do I need a DPA contract for cloud backup within the EU?

Yes. Even if the provider is European, the data processing agreement (DPA) is mandatory under GDPR. Check that it is explicit that the data resides in EU data centres and that the provider does not transfer it outside. Without this, you are at risk of a regulatory fine regardless of the technical security.

What is the most common mistake in choosing between cloud and on-premise?

Choosing the architecture before defining RTO, RPO and the dependency map. This leads to investment in unnecessary hardware or cloud contracts that don't solve the real problem. The architecture is secondary. What decides everything is knowing, to the minute, how long it takes to get back up and running.

Sources

  • Verizon Data Breach Investigations Report 2025 — Annual data breach investigations report, published by Verizon Communications Inc.
  • Regulation (EU) 2016/679 (GDPR) — General Data Protection Regulation, applicable to the processing of personal data in the context of backup and compliance
  • Directive (EU) 2022/2555 (NIS2) — Directive on measures for a high common level of cybersecurity, including operational continuity and backup requirements
  • ISO/IEC 27001:2022 — International information security management standard, including requirements for backup, data recovery and continuity planning
  • ENISA Guidelines on Backup and Recovery — Guidelines from the European Union Agency for Cybersecurity on backup, data recovery and RTO/RPO