Eighty per cent of the Portuguese industrial companies that ask us to implement a new system have the same problem: the infrastructure did not fail for lack of investment. It failed for lack of alignment. They bought servers without knowing the ERP's RTO. They migrated to the cloud without measuring the latency of the production line. They renewed switches without asking the warehouse manager what brings his operation to a halt. The result always appears at the same moment: month-end, invoicing peak, ERP slowing down, afternoon shift at a standstill.

This guide is not a neutral survey of best practices. It is an argument: an infrastructure audit only has value if it starts from the business processes and works backwards — never the other way round. Five dimensions, one decision template, and at least one mistake you have probably already made.

What you need before you start

Without these elements in hand, the exercise produces pretty lists that no one executes. Gather them before convening any audit meeting.

You need the map of critical processes — production, invoicing, dispatch, sales — even if it is a hand-drawn block diagram. You need the current systems inventory: ERP, MES, WMS, BI, POS, portals, with version and date of the last upgrade. You need the SLA contract of the infrastructure or datacentre provider — read the availability clauses and the contracted RTO/RPO values. And you need the report of the last IT incident with operational impact, even if it is just a support e-mail with a start time and a resolution time.

Add to this the approved IT budget for the current year, split into CAPEX and OPEX, the list of applicable regulatory obligations — NIS2, DL 28/2019, GDPR — and at least one representative from each operational area willing to answer ten direct questions. Without the voice of operations, the audit ends up describing the infrastructure that IT knows, not the infrastructure the business needs.

Dimension 1 — People: who decides and who suffers

The most common mistake we see in INFOS projects is not technical. It is organisational: the infrastructure is designed by IT without input from operations, and operations have never been asked what brings them to a halt.

In a typical garment factory in the Vale do Ave with 120 employees, the IT manager single-handedly manages servers, ERP, shop-floor Wi-Fi and helpdesk. When there is an incident at 7 a.m. at the start of the shift, the priority decision is theirs — with no formal criterion, no escalation matrix, no one telling them which process costs most per hour of downtime. The concrete result we have seen repeated: the electronic time-and-attendance system is down for 40 minutes because the production server was restarted first. The decision was technically reasonable. Operationally, it cost 40 minutes of manual paper recording that no one was later able to reconcile.

The solution is not to hire more IT. It is to define a minimum RACI matrix for infrastructure incidents — three columns are enough: affected system, technical owner, business owner to notify. Identify the two or three processes whose downtime costs most per hour and appoint a "process owner" for each: it is not IT, it is operations. Document who authorises a planned shutdown and with how much advance notice. Check whether IT has a substitute for absences — holidays, sick leave, departure. In a company with a single IT technician, an unplanned absence is itself a continuity incident.

Dimension 2 — Processes: mapping what the infrastructure has to support

Infrastructure does not exist to "support the business" in the abstract. It exists to ensure that specific processes run within defined parameters. If you do not have those parameters written down, any infrastructure decision is a gamble disguised as an investment.

For each critical process, answer four questions without ambiguity. What is the maximum tolerable downtime — the RTO? What is the maximum tolerable data loss — the RPO? What is the predictable load peak — month-end, collection close, Black Friday, shift start-up? Which external systems depend on this process — customer, carrier, the tax authority, bank?

The detail the manuals do not mention: in footwear companies in Felgueiras with collections of 800 to 1,200 SKUs, the load peak is not month-end — it is the week before the visit of international buyers, when the entire sales team accesses the ERP simultaneously to generate technical sheets, price lists and availabilities. If the infrastructure was sized for average load, that week is always a problem. Map the real peaks of your business, not the peaks that IT imagines.

Map dependencies between systems — the ERP that feeds the WMS that feeds the POS. Identify single points of failure: a server, a network connection, a supplier. Check whether backups are tested — not just executed. Restore a test file quarterly, with a stopwatch. Document the manual contingency procedures for each critical process: when the system goes down, the team cannot waste time working out what to do.

Dimension 3 — Systems: the inventory no one wants to do

Audit what exists before buying what is missing. In ERP MULTI implementation projects at industrial companies, it is common to discover systems in production that no one knew were still active — a file server with data from 2014, an order management software that was "replaced" three years ago but still receives data from an old customer because no one communicated the change. These invisible systems are security and compliance liabilities.

System Version / Last update Criticality (High/Medium/Low) Integrated with Owner Action needed
ERP __________ High WMS, invoicing, BI __________ __________
WMS / Warehouse __________ High ERP, carriers __________ __________
BI / Dashboards __________ Medium ERP, production __________ __________
Shop floor / MES __________ High ERP, terminals __________ __________
File server __________ __________ __________ __________ __________
E-mail / Communication __________ __________ __________ __________ __________
POS / Retail __________ __________ __________ __________ __________

Fill in one row per system. Any empty cell is an unmanaged risk. Identify systems without active vendor support — they are immediate security liabilities. Check licensing: unlicensed software exposes the company to audit and to unpatched vulnerabilities. Classify each system by criticality — they are not all equal and should not receive the same resilience investment.

Dimension 4 — Data: the asset the infrastructure has to protect

Infrastructure does not protect servers. It protects data. If you do not know where the critical data is, you do not know what you are protecting.

The global average cost of a data breach reached 4.88 million dollars in 2024 — a 10% increase over 2023 (IBM, 2024). For a Portuguese industrial SME, the impact is not that absolute figure: it is the production stoppage, the loss of trust from the international customer, and the regulatory exposure. The national figures confirm the trend: CERT.PT recorded 2,758 cybersecurity incidents in Portugal in 2024, 36% more than in 2023, with 78% occurring in private entities (CNCS, 2024). And according to the Verizon DBIR 2025, ransomware was present in 88% of data breaches at SMEs — against 39% at large organisations. Small and medium-sized enterprises are the disproportionate target, precisely because the infrastructure is simpler to compromise and the response capacity is smaller.

The NIS2 Directive, transposed in Portugal by Decree-Law No. 65/2025, requires medium and large companies in critical sectors — including industry — to implement documented security controls and to report incidents. Ignoring this is not a management option: it is a legal exposure with fines and director liability.

To align data and infrastructure, classify the data by sensitivity — customer data, financial data, production data, HR data — and map where each category is stored: local server, cloud, employee laptop, USB drive. Check whether access to sensitive data is restricted by profile. Confirm that backups of critical data are off the main site, following the 3-2-1 rule: three copies, two different media, one offsite copy. Audit the access logs — if they do not exist, that is the first thing to implement, before any other control.

To delve deeper into the technical security controls applicable to industry, see the article on NIS2 in practice: technical controls for Portuguese factories. To understand the risks that go unnoticed, read the silent cybersecurity risks in the company.

Dimension 5 — Compliance: what the law already requires

Compliance is not a project separate from infrastructure. It is a set of requirements the infrastructure has to meet — and which, if ignored, generate fines, audits and forced stoppages. The most frequent problem we see is not ignorance of the law: it is the company that knows about the obligation but assumes that "the supplier takes care of that". The supplier takes care of the software. The responsibility is the company's.

DL 28/2019 requires invoicing software to be certified by the tax authority: check whether the version in use has valid certification and whether it issues ATCUD correctly — an outdated version may invalidate issued invoices. GDPR and Law 58/2019 require a legal basis for processing personal data, a defined retention period, and the capacity to respond to access or deletion requests within 30 days. NIS2 and DL 65/2025 require medium and large companies in critical sectors to document the security policy, the incident response procedures, and the formal point of contact with the CNCS. Law 93/2021 makes the whistleblowing channel mandatory for companies with 50 or more employees — check whether the infrastructure supports it with guaranteed confidentiality, which rules out a simple e-mail address shared with the HR department. And if you use a digital signature on contracts or legal documents, confirm that it is a qualified signature under eIDAS 2 — not a scanned image of a handwritten signature, which has no equivalent legal value.

Five mistakes the audit will find — and how to fix them

Mistake 1 — Buying infrastructure without defined RTO/RPO. The result is a new server that no one knows how long it can be down for. Before any acquisition, define RTO and RPO per critical process. If the ERP cannot be down for more than two hours, that has direct implications for the backup architecture and for the decision between a local server and the cloud.

Mistake 2 — Treating the cloud as a destination, not a decision. Migrating to the cloud without assessing latency, egress cost and connectivity dependence creates new problems. A factory with a production line that depends on ERP in a public cloud and an internet connection with no redundancy is more fragile than a well-managed local server. Analyse the real ROI of the migration, including the cost of failure and the monthly cost of a redundant connection — which is mandatory, not optional, if the critical process depends on connectivity.

Mistake 3 — Ignoring the shop-floor network. Industrial Wi-Fi is not the same as office Wi-Fi. Production terminals, barcode scanners and real-time production capture equipment require stable coverage, network segmentation and redundancy. We regularly see factories where the OT — Operational Technology — network shares the same segment as the sales team's laptops. When a laptop is compromised, the production network is exposed. Segmentation is not an advanced cybersecurity project: it is switch configuration, and it should be done before any industrial terminal is connected to the network.

Mistake 4 — Confusing backup with recovery. Having a backup does not mean being able to recover. Test the full restoration of a critical system at least once a year, with a stopwatch and with the person who will do the real recovery — not the technician who configured the backup. The result will surprise you, almost always negatively. At a distribution company in the Lousada/Paços corridor, the test revealed that the WMS backup was being made to a folder that had not existed for six months. The backup ran without error. There was nothing stored.

Mistake 5 — Leaving compliance for the end. Regulatory compliance is not a layer applied on top of the existing infrastructure. It is a design requirement. When NIS2 requires auditable access logs and the company has no logging configured, the cost of retroactive implementation is always higher than the cost of having done it from scratch. Integrate the regulatory requirements into the audit from the very first dimension — not as a final checklist, but as a decision criterion in every infrastructure choice.

The five-dimension audit is not an academic exercise. It is two to three weeks of work that avoids six months of recovery after an incident. The difference between the companies that recover quickly and those that do not lies not in the technology they had — it lies in the fact that they knew exactly what they had, where it was, and what to do when it stopped working.

Frequently asked questions

What is RTO and why does it matter for infrastructure?

RTO (Recovery Time Objective) is the maximum time a critical process can be unavailable. It defines how much downtime the company tolerates before suffering serious operational losses. Without a defined RTO, the infrastructure is sized at random, causing unnecessary stoppages or wasted investment in redundancy.

How do I identify my company's critical processes?

Start with the processes that, if they stop, cost money immediately: production, invoicing, dispatch and sales. Bring together the managers of each area and ask what the impact is per hour of downtime. Also map the dependencies — which system feeds which. A hand-drawn block diagram is enough to get started.

What is the difference between RTO and RPO?

RTO is the maximum tolerable downtime. RPO (Recovery Point Objective) is the maximum amount of data you can lose — for example, "we lose data from the last 15 minutes". An ERP may have an RTO of 4 hours but an RPO of 5 minutes, requiring frequent backups even if recovery is slower.

Why do traditional infrastructure audits fail?

They fail because they start from the technology instead of the business. They describe what exists without questioning whether it serves what is needed. They produce pretty lists that no one executes because they have no connection to the real processes. The correct audit starts from the critical processes and works backwards to the necessary infrastructure.

What is a RACI matrix for infrastructure incidents?

It is a simple document with three columns: affected system, technical owner, business owner to notify. It defines who decides the priority of each incident and who communicates with users. It prevents IT from deciding alone which system to restart first, without knowing which one costs most per hour of downtime.

How should I test my backups?

It is not enough to run backups automatically. Quarterly, restore a test file with a stopwatch and document the time. Check whether the data is intact and whether you can access it without problems. A backup that has never been tested is just a file taking up space — it may be corrupted and no one knows.

What information do I need to gather before doing an infrastructure audit?

The map of critical processes, the systems inventory with versions, SLA contracts with RTO/RPO values, the report of the last IT incident, the IT budget split into CAPEX and OPEX, applicable regulatory obligations, and at least one representative from each operational area. Without the voice of operations, the audit describes the infrastructure that IT knows, not the one the business needs.

Sources

  • Directive (EU) 2022/2555 (NIS2) — Security of Network and Information Systems, operational continuity requirements and incident management
  • Decree-Law No. 28/2019 — Transposition of Directive (EU) 2016/1148 (NIS1), obligations of operators of essential services in Portugal
  • Regulation (EU) 2016/679 (GDPR) — Protection of personal data, including systems availability and integrity requirements
  • Standard ISO/IEC 27001:2022 — Information security management, business continuity and disaster recovery
  • ENISA — Guidelines on Incident Handling and Response (2016) — Best practices in infrastructure and IT incident management