Business process automation and the failures you must notice
OperationsSeptember 1, 2026·Zorah Team

Business process automation and the failures you must notice

Business process automation runs on somebody else's infrastructure, and the division of duty is simple to state and almost never written down. The provider is answerable for the platform staying up. You are answerable for whether the process did what it was supposed to. Between those two sits a set of failures nobody is watching, because they look like nothing at all.

An automation that stops running produces no error. It produces silence, and silence reads exactly like everything being fine.

The concept the vendors keep raising

Writing on ITWeb, Christo Coetzer argues that the shared responsibility model is one of the least understood concepts in South African business today. That is a vendor's framing and it is worth treating as such.

It is also, on the operational evidence, correct. ITWeb separately reports a Cloud Security Alliance survey of 515 IT and security professionals, taken in May 2026. In it, 65% had a business-critical application go down because of a policy error in the past year, and 46% had it happen twice or more. An actual breach was the consequence for 18%.

Notice what that measures. Not an attack, and not a provider outage: a setting somebody changed, in an environment split between a provider and a customer. The useful move is to convert the framing into a list. If the model splits duties, what exactly is on your side of the line?

The four business process automation failures on your side

The provider tells you when the service is down. Every one of these can happen while the service is up.

A step stopped firing. A credential expired, a field was renamed, an integration was disconnected during a tidy-up. The automation reports success on the runs it did and says nothing about the runs it did not.

A step fired and did the wrong thing. A rule that was right in March against a price list that changed in July. The system is working exactly as configured, which is the problem.

A queue is growing. Items entering faster than they are cleared. Nothing is broken until the day somebody notices the backlog is three weeks old.

Somebody changed something. A person with admin rights adjusted a rule, and there is no record of who or when. This is the one that turns a small failure into an argument.

None of the four appear on a provider status page, because none of them are the provider's fault.

What South Africa adds to the list

Two local conditions make the silent failure more likely here than the vendor literature assumes.

Connectivity is intermittent, and load shedding takes sites offline in blocks. An automation that reads from an on-premises system, a branch server or a device at a depot will simply not run during those windows, and most of them do not queue and retry by default. The work resumes and the gap stays.

Payment terms and manual handovers do the rest. A process that depends on somebody uploading a file on a Friday fails quietly when that person is on leave. The automation cannot tell the difference between no file and no work.

The three things to log, and they are not technical

Whatever tool you use, insist on three records. They are cheap to configure and they are what turns silence into information.

  • A heartbeat. Every scheduled job reports that it ran, whether or not it did anything. An absent heartbeat is the alarm.
  • A count, with an expected range. Forty invoices a day is normal. Two is an incident and four hundred is a different incident.
  • A change log. Who edited which rule, and when. Nothing else settles the argument about when the behaviour changed.

Then name an owner. Not a department: a person, who receives the heartbeat and is expected to notice. An alert routed to a shared mailbox is an alert nobody has agreed to read.

Why this is a design decision, not a monitoring purchase

The reason most South African businesses cannot do this is that the automation was bought inside a product. It runs where the vendor put it, logs what the vendor logs, and there is no place to attach a heartbeat.

Owning the joins is what makes it possible. Zorah builds the integration layer as something the business keeps rather than something rented inside a module, which means the counts and the change log belong to you and survive changing any single vendor.

That layer is the automation and integration capability, the counts belong in the reporting layer, and the daily version of the exception list usually lands with whoever runs operations.

What to do on Monday

List every automation running in your business, including the ones inside products. For each, write down what it would look like if it stopped.

Anything where the honest answer is "we would find out from a customer" is the one to instrument first. That is not a technology problem yet, and it will not stay that way.

ShareLinkedInEmail

Comments (0)

Leave a comment