← Work
// automation & AI system · for real-estate

An agentic WhatsApp assistant, in production

An AI agent that answers clients on WhatsApp in three languages, searches the CRM for matching listings, and writes qualified leads back, without ever overwriting a client record.

9n8n workflows
Liveon a real WhatsApp number
DE · EN · FRlanguages handled

The problem

A large share of client conversations never reached the CRM. They happened in personal WhatsApp threads, on agents' phones, after hours. Nobody answers a listing enquiry at 11pm. And the agency's clients are international: German locals, English-speaking expats, French investors.

Every message that went unanswered was a lead quietly evaporating, with no trace anywhere that it had ever existed.

What I built

An AI agent connected to the agency's real WhatsApp number through Meta's Cloud API, live since June 2026. It holds a conversation, works out what the person actually wants, and acts on the CRM through two tools of its own.

  • Detects the language and qualifies the contact type (buyer, renter, owner, landlord), then adapts: an owner is never shown listings, they are qualified as a seller lead.
  • Searches the CRM across 14 criteria: price range, rooms, bedrooms, area, neighbourhood, energy class, elevator, cellar, balcony, floor.
  • Sends one photo per listing, with that listing's description in the caption, not a wall of text followed by loose images.
  • Creates the contact, or recognises a returning one by email, then by phone number.
  • Logs every tool call with zero personal data, and emails a weekly report every Monday morning.

How it works

USER Client WhatsApp GATEWAY Meta Cloud API inbound webhook ✦ AI AGENT · LIVE 24/7 AI agent Gemini 2.5 Flash + 12-msg context FORMATTER Composer splits the reply OUTBOUND Reply text + 1 photo / listing AGENT TOOL Tool · Search 14 query criteria SAFE WRITE Tool · CRM dedupe + activity log onOffice · REST API 400+ fields · HMAC-signed transactions Analytics & Report data table → weekly summary

Both tools log every call: search criteria and outcomes, never personal data.

Under the hood

PITFALL

“No listings found”, with twenty-four sitting in the CRM

In n8n, $fromAI('bedrooms', …, 'number') without a fourth argument marks the field as required in the schema handed to the model. The moment the model omitted an optional parameter, n8n rejected the tool call before it ever reached the sub-workflow. The agent saw an error and told the client there was nothing available.

It was intermittent, depending on what the model chose to send, so it never reproduced in testing.

FIX Give every $fromAI a default value; the field becomes optional. Coalescing with || 0 does nothing: validation happens before the expression is evaluated.

DECISION

Never overwrite a client record

An adversarial audit showed that an unverified email address could trigger an update to a third party's record. So when the agent meets a contact it already knows, it never writes to that record. It appends an entry to the CRM's activity log instead.

// ADVERSARIAL AUDIT DISCOVERY · CLIENT RECORD INTEGRITY
Standard Direct CRM Overwrite
CORRUPTION RISK
❌ Unverified inbound phone/email overwrites existing buyer profile
Hermes Append-Only Architecture
100% ISOLATED SAFE
✅ Appends timestamped log entry to CRM activity timeline. Zero destructive writes.
ARCHITECTURAL RULE: Conversational AI bots must never possess unilateral UPDATE/DELETE permissions on client database rows. Append-only ledger pattern guarantees total tamper resistance.

TRADE-OFF We lose automatic record enrichment. We gain the structural impossibility of corrupting client data.

PITFALL

The duplicate that kept coming back

The CRM's address search does not match a contact by phone number. A lead who had only ever given a phone number was recreated as a fresh duplicate every single time they came back.

FIX An exact filtered read on the phone field, fired only when no email was supplied. Email stays authoritative, so the two never cross-match.

PITFALL

The bot quoting its own outdated self

After shipping a fix that added minimum-budget search, the bot kept telling one tester it could only handle a maximum. The conversation memory still held the pre-fix message, and the model was citing itself.

FIX Verified on a fresh session, then instructed the model that its current capabilities are authoritative and it must never repeat a limitation stated earlier in the thread.

DECISION

Sort by energy rating, don't filter by it

Clients ask for energy-efficient flats. Berlin's older housing stock is overwhelmingly rated D or E. A strict A-to-C filter would have returned zero results, every time.

TRADE-OFF The agent surfaces the best available ratings first instead of promising a standard the market cannot supply.

Results

Live on the agency's real WhatsApp number since June 2026. Nine workflows: the production agent, a chat-based twin for testing, two agentic tools, a URL-sync job, an evaluation harness that replays test cases, a stats dashboard, a weekly report, and a dedicated error workflow.

That error workflow earned its keep: it surfaced a production failure within minutes, before any client noticed: the model had returned an empty string, which WhatsApp rejects.

Still open. Direct listing links depend on a bug in the agency's website that I don't control; the agent falls back to the category page rather than send a broken URL. An audit also flagged secrets management: credentials belong in environment variables, not in application code. Externalising and rotating them is a server-side change, documented for whoever picks up the instance.

Stack

  • n8n, self-hosted on a VPS: agent orchestration, tools, ops workflows
  • Gemini 2.5 Flash: reasoning, language detection, qualification
  • Meta WhatsApp Cloud API: inbound webhook, text and image sends
  • onOffice REST API: HMAC-signed listing and contact operations
  • n8n Data Tables: anonymous event logging

What I took away

The hard part of an AI agent is not the prompt. It is everything around it: what happens when the model omits a parameter, when a client replies “ok”, when two people share a phone number, when the model returns an empty string.

The prompt took an afternoon. The failure modes took weeks, plus an adversarial audit to find the ones I had talked myself out of worrying about.