← Work
// automation & AI system · for real-estate

32 active workflows, and how they don't fall over

Lead capture, CRM sync, monitoring, reporting. The interesting question isn't how to build them. It's how to find out when one quietly stops.

32active workflows
1central error workflow
4test cases replayed on demand

The problem

Automations are easy to build and easy to forget. Thirty-two of them run continuously: capturing leads from four sources, syncing the CRM, watching e-mail authentication, generating reports.

Each one is a small promise. When one breaks, it does not raise its hand. It just stops, and a lead goes nowhere, and nobody notices until a person asks a question that the data cannot answer.

What I built

Not a workflow. The scaffolding around them.

  • One central error workflow, attached to every significant automation, that emails the failure with context.
  • An evaluation harness that replays fixed test cases against the AI assistant and asserts on the answers.
  • Anonymous event logging in native data tables, so usage can be measured without storing personal data.
  • An automated instance-wide security audit, run and acted upon rather than filed.
  • A handover runbook, because the internship ends and the workflows do not.
LIVE 32-WORKFLOW TOPOLOGY & FAILOVER SIMULATOR
Test Global Failover, Central Error Trapping & Eval Replay

Select a production event below to test how the 32-workflow fleet responds: see failing upstream webhooks captured by the central dead-letter handler, and prompt regressions blocked by the eval harness.

TRIGGER SYSTEM EVENT:
// N8N VPS PRODUCTION FLEET (32 WORKFLOWS) 32 / 32 ACTIVE
Nominal Trapped / Retrying Error Intercepted
DISPATCHER & TRACE LOG ALL SYSTEMS NORMAL
HOSTING Docker on isolated Hetzner VPS
RECOVERY PATTERN Global Error Workflow (Zero lost leads)
REGRESSION GATE Pre-deploy synthetic conversation runner
GDPR STATUS Zero PII in execution logs

How it works

Lead capture 4 portal sources CRM sync listings · contacts Monitoring + scheduled reports ✦ CENTRAL DOCKER VPS n8n Engine self-hosted · Docker · VPS 32 active production workflows TRAP GATE Central error workflow catches every major failure CI / TEST GATE Eval harness replays synthetic test cases Security & Secret Audit credential token masking zero-retention PII enforcement

The error workflow is the only reason silent failures became loud ones.

Under the hood

DECISION

One error workflow, not scattered error handling

Every significant workflow points at the same error handler. It emails what broke, where, and with what input. It found the empty-message bug in the WhatsApp assistant minutes after it first happened, a failure that produced no client-visible symptom other than silence.

TRADE-OFF A single point of failure for failure reporting. Worth it: the alternative is thirty-two places to forget.

DECISION

An eval harness for the AI assistant

Prompts are code with no compiler. Change one sentence to fix one behaviour, and you silently break a case you fixed last month.

A webhook replays a fixed set of conversations and asserts on the outcome. It runs before any prompt change ships.

TRADE-OFF Four cases is not coverage. It is the difference between zero and non-zero, which is the difference that matters.

PITFALL

Credentials are global, and deleting one breaks everything

A colleague reconnected a Google account by deleting the credential and creating it again. The new credential got a new internal ID. Seven workflows pointed at the old one and stopped, silently, at once.

// INCIDENT POST-MORTEM · N8N GLOBAL CREDENTIAL BINDINGS
Naive Delete & Re-create
7 CASCADING FAILURES
❌ Deleting credential generates a new UUID. 7 workflows referencing old UUID break instantly.
In-Place OAuth Reconnect
ZERO DOWNTIME PRESERVED
✅ Reconnect updates the auth token while preserving identical internal ID. All 7 workflows keep running.
RUNBOOK DIRECTIVE: Never delete an active credential object in production. Always click 'Reconnect' to preserve UUID references across dependent DAG topologies.

FIX Always “Reconnect” in place, never delete-and-recreate. Written into the runbook in bold, because this is the kind of thing you learn once and only once.

DECISION

Running the audit is easy; acting on it is the job

An automated audit of the instance returned hundreds of findings across secrets management, endpoint exposure and dependency freshness. It would have been very easy to file that report and move on.

Instead: keys moved behind proxies, execution history purged where a token had been persisted in plain text, retention turned off on the workflows handling personal data.

HONEST Not all of it is closed. What is left needs a server-side change and a coordinated rotation, so it is sequenced and documented for whoever comes next.

Results

Thirty-two active workflows, verified by querying the instance rather than trusting my memory of it. Error handling attached to every major one. A handover runbook so that none of this dies with my badge.

Still open. The instance is several versions behind. Tightening access to the internal endpoints has to be sequenced rather than switched on in one go, because the front ends that call them would break the moment it changed.

Stack

  • n8n, self-hosted on a VPS, in Docker
  • onOffice REST API: HMAC-signed, the system of record
  • NocoDB and n8n Data Tables: storage, ledgers, anonymous events
  • Google Workspace APIs: Drive, Docs, Gmail, Sheets

What I took away

Automation is not finished when it runs. It is finished when you find out it stopped before your manager does.

Most of what I am proud of here is unglamorous: an error workflow, four test cases, and a document explaining what will break and why.