The problem
Automations are easy to build and easy to forget. Thirty-two of them run continuously: capturing leads from four sources, syncing the CRM, watching e-mail authentication, generating reports.
Each one is a small promise. When one breaks, it does not raise its hand. It just stops, and a lead goes nowhere, and nobody notices until a person asks a question that the data cannot answer.
What I built
Not a workflow. The scaffolding around them.
- One central error workflow, attached to every significant automation, that emails the failure with context.
- An evaluation harness that replays fixed test cases against the AI assistant and asserts on the answers.
- Anonymous event logging in native data tables, so usage can be measured without storing personal data.
- An automated instance-wide security audit, run and acted upon rather than filed.
- A handover runbook, because the internship ends and the workflows do not.
Select a production event below to test how the 32-workflow fleet responds: see failing upstream webhooks captured by the central dead-letter handler, and prompt regressions blocked by the eval harness.
How it works
The error workflow is the only reason silent failures became loud ones.
Under the hood
One error workflow, not scattered error handling
Every significant workflow points at the same error handler. It emails what broke, where, and with what input. It found the empty-message bug in the WhatsApp assistant minutes after it first happened, a failure that produced no client-visible symptom other than silence.
TRADE-OFF A single point of failure for failure reporting. Worth it: the alternative is thirty-two places to forget.
An eval harness for the AI assistant
Prompts are code with no compiler. Change one sentence to fix one behaviour, and you silently break a case you fixed last month.
A webhook replays a fixed set of conversations and asserts on the outcome. It runs before any prompt change ships.
TRADE-OFF Four cases is not coverage. It is the difference between zero and non-zero, which is the difference that matters.
Credentials are global, and deleting one breaks everything
A colleague reconnected a Google account by deleting the credential and creating it again. The new credential got a new internal ID. Seven workflows pointed at the old one and stopped, silently, at once.
FIX Always “Reconnect” in place, never delete-and-recreate. Written into the runbook in bold, because this is the kind of thing you learn once and only once.
Running the audit is easy; acting on it is the job
An automated audit of the instance returned hundreds of findings across secrets management, endpoint exposure and dependency freshness. It would have been very easy to file that report and move on.
Instead: keys moved behind proxies, execution history purged where a token had been persisted in plain text, retention turned off on the workflows handling personal data.
HONEST Not all of it is closed. What is left needs a server-side change and a coordinated rotation, so it is sequenced and documented for whoever comes next.
Results
Thirty-two active workflows, verified by querying the instance rather than trusting my memory of it. Error handling attached to every major one. A handover runbook so that none of this dies with my badge.
Still open. The instance is several versions behind. Tightening access to the internal endpoints has to be sequenced rather than switched on in one go, because the front ends that call them would break the moment it changed.
Stack
- n8n, self-hosted on a VPS, in Docker
- onOffice REST API: HMAC-signed, the system of record
- NocoDB and n8n Data Tables: storage, ledgers, anonymous events
- Google Workspace APIs: Drive, Docs, Gmail, Sheets
What I took away
Automation is not finished when it runs. It is finished when you find out it stopped before your manager does.
Most of what I am proud of here is unglamorous: an error workflow, four test cases, and a document explaining what will break and why.