An operations director at a weaving mill in the Vale do Ave told us recently: "We tested a generative AI to read supplier receipts. It worked well for the first 200 documents. Then it started inventing numbers — and we only realised when the Tax Authority queried our invoicing." The AI was not wrong. It was overconfident. And that confidence cost it an audit.

This is not an argument against generative AI in document management. It is an argument against the illusion that it works well at everything. In Portugal, only 11.5% of companies with 10 or more employees use artificial intelligence technologies — and adoption is dramatically uneven: 49.1% in large companies, but just 9.4% in small ones (10-49 employees), according to INE in 2025. When these small and medium-sized enterprises finally test AI on documents, they discover an uncomfortable truth: it works very well on some very specific things, and dangerously on others — and the difference is not obvious at first glance. Worse still: the illusion of AI confidence is more dangerous than human incompetence, because it passes visual validation and reaches regulatory records before being caught.

What generative AI does well with documents

Let us start with the obvious: generative AI is excellent at tasks where variation is high, the pattern is linguistic and the margin for error is forgiving. Extracting contract terms from a supplier email. Summarising a non-conformity report. Classifying a customer complaint into categories. Filling in metadata fields in a proposal file. Generating a return-refusal response.

In these situations, what you are doing is transforming natural language into structure — and generative AI is, by definition, a model trained for precisely that. The accuracy rate is high. The computational cost is low. And, importantly: the user can quickly validate whether the response makes sense.

We see this working well in companies that use generative AI in document management for automatic triage of incoming invoices (supplier, date, number), or for extracting terms from purchase contracts. Manual processing time drops by 60-70%. The errors that slip through are caught at review stage — because review is quick, structured and predictable.

Generative AI in documents is only as good as your ability to validate the result. If you cannot validate it in 30 seconds, do not let the AI do it alone.

Where generative AI fails silently

Now the problem. Generative AI is terrible at decisions that demand absolute truth or regulatory compliance. And industrial document management is full of them. The reason is simple: generative models do not understand the domain. They understand statistical patterns in language. When the pattern is ambiguous, or when the context is critical, their confidence is an illusion — and that illusion passes visual validation because the result "looks correct".

Example: a generative AI reading a raw-material traceability document and extracting the batch number. It seems simple. But if the batch number is invented — because the model has never seen that specific pattern and "hallucinated" it — and you record it in your DMR (Device Master Record, or its equivalent in textile/food compliance), you have a regulatory problem. It is not a wrong opinion. It is a false fact on record. An operator reviewing in 45 seconds using a checklist of 3 critical fields (batch, quantity, supplier) can catch this — but only if they know they have to look for it.

Or: a generative AI extracting quantities from a purchase order. If it says "100 units" when the document says "1000 units", the cost is a wrong order. If it says "1000 units" when the document says "100", it is waste. The AI does not know it made a mistake — because it was never trained to know what "wrong" means in a business context. Worse: if you implement a system that automatically flags with "confidence <85%" the records that need review, the AI may be 100% confident in a completely false answer — because the statistical pattern it saw during training was misleading.

The difference between "AI that works well" and "AI that is dangerous" is not technological. It is contextual. A generative AI summarising an email is safe because the error is obvious — you read the summary and the email and see whether they match. A generative AI extracting a serial number that goes into a compliance record is dangerous because the error is invisible — you have no "reference truth" to compare against, and the AI looks confident.

The practical rule: where to put the brakes on

Here is what we advocate: use generative AI in document management ONLY for tasks where an error does not cause immediate regulatory, financial or operational harm. And always with structured human validation.

Do not use generative AI alone (without validation) for:

  • Extracting serial, batch or traceability numbers that go into compliance or warranty records.
  • Document approval/rejection decisions (a "yes" or "no" that triggers a downstream process).
  • Filling in tax audit fields (values, dates, tax references).
  • Classifying confidential or access-restricted documents (the AI may expose sensitive data while trying to "understand" the content).
  • Extracting contractual obligations that bind the company (deadlines, penalties, automatic renewal).

Use generative AI with structured human validation for:

  • Initial document triage (supplier, type, urgency) — an operator validates in 20-30 seconds, with a checklist of 3-4 fields.
  • Content summaries for review (a manager reads the summary and then the document, if they feel they need to).
  • Filling in non-critical metadata fields (responsible department, associated project).
  • Generating standard responses to routine documents (receipt confirmation, review scheduling).

The difference is not technological. It is one of risk — and of who is responsible when the AI gets it wrong.

What we have seen change — and where we got it wrong

Three years ago, we believed generative AI would be the solution to the classic problem of Portuguese industrial SMEs: disorganised documentation, slow manual processing, transcription errors. And it is true that it helps — but not in the way we imagined.

What changed: reality imposed a distinction that software vendors did not want to make. You do not want "AI that reads documents". You want "AI that reads documents AND ensures compliance". These are two different things. The second requires an architecture that generative AI alone does not provide. It requires rigid business rules, workflow orchestration, and clear accountability over who validates what.

That is why the most robust implementations we see combine generative AI (for speed, for linguistic pattern) with rigid business rules (for compliance, for critical numbers) and with orchestrated human validation (an RPA that routes what needs reviewing, and a human who reviews in 20-45 seconds using a checklist of critical fields). This is slower than "full automation", but it is what works in regulated environments.

Generative AI does not replace compliance. It speeds up processing. Compliance is your responsibility — always.

The question you should ask before implementing

Before contracting or configuring generative AI in document management, put this question to your IT Director or consultant: "If the AI gets this task wrong, what is the cost? How long does it take to catch the error? Who is responsible?"

If the answer is "it would cost us an audit" or "it would take weeks to catch", do not use generative AI alone. Use it with mandatory human validation — and define the protocol: who validates, within what timeframe, using which checklist, and what they do if they believe the AI got it wrong.

If the answer is "it would cost us 10 minutes of processing, nothing more", then use generative AI and let a human review a batch once a week. But even in this case, implement an automatic flag if confidence drops below 85% — not because the AI is perfect above that threshold, but because it forces the human to look at cases the AI itself considers "doubtful".

Most Portuguese companies that test generative AI on documents start down the wrong path: they want full automation, with no validation. Then they discover they cannot have full automation on critical tasks. And they become frustrated with the technology, when the problem was the expectation.

Generative AI is a tool for acceleration, not for shifting responsibility. Whoever implements it has to be able to live with the error it is going to make — because it will. And they must have a plan for when it happens.

Frequently asked questions

When is it safe to use generative AI in document management?

It is safe when the task involves transforming natural language into structure, the error is quickly visible and it causes no regulatory harm. Examples: invoice triage (supplier, date), report summaries, complaint classification. The user should validate the result in under 30 seconds.

What is the risk of using generative AI on serial or batch numbers?

The AI can invent numbers that look correct but are false. If you record a non-existent batch number in a compliance dossier, you create an invisible regulatory problem. The operator has no "reference truth" to compare against, and the AI looks confident even when it is completely wrong.

Why is AI confidence more dangerous than human error?

Because it passes visual validation and reaches the records before being caught. A human operator hesitates when unsure. The AI is confident even when it hallucinates. This makes the error invisible until the audit or regulatory query.

Should I use AI for document approval or rejection decisions?

No, not without human validation. An automatic "yes" or "no" triggers downstream processes that can cause immediate operational or financial harm. The AI does not understand the domain — only statistical patterns — and cannot judge critical contexts.

How do I validate AI results in document management?

Use structured validation: create a checklist of 3-4 critical fields and review in 20-30 seconds. This works well in document triage (supplier, type, urgency) or metadata entry. If you cannot validate quickly, do not let the AI do it alone.

Can I use AI to extract contractual obligations?

Not without human validation. Deadlines, penalties and automatic renewal clauses legally bind the company. The AI may lose context or invent terms. An invisible error here causes immediate regulatory and financial harm.

What is the difference between "AI that works well" and "dangerous AI"?

It is not technological — it is about context and risk. AI summarising an email is safe because the error is obvious. AI extracting a serial number for compliance is dangerous because the error is invisible. The difference lies in who is responsible when the AI gets it wrong.