HS

Himanshu Sharma

Lead Platform Engineer

← All blogs

Scaling • Observability • Messaging

Scaling Omnichannel Services to 100K Daily Interactions

Latency spikes in messaging become revenue hits. This is the architecture we used at Freshworks to keep WhatsApp, email, SMS, voice, and Teams reliable as traffic tripled.

100K+daily interactions
95 msp99 at peak
20%engagement lift
<3 mintime to detect

Plan capacity by conversation lane

Do not size “the messaging service.” Size lanes: SMS, WhatsApp, email, voice, Teams. Each lane has a different partner SLA, burst shape, and compliance boundary. Mixing them on one Kafka topic creates noisy-neighbor outages.

Clients · products · bots WebFlux ingress · per-lane connection pools SMS lane Kafka t.sms WhatsApp Kafka t.wa Email / Voice Kafka t.mail Teams bot Kafka t.teams Delivery adapters + partner circuit breakers rate-aware batching · regional failover · residency tags
Lane fan-out from ingress Active lane partition
Shared ingress, isolated lanes, shared delivery adapters with per-partner breakers.

During marketing bursts, WhatsApp lanes may 5× while SMS stays flat. Autoscaling hooks watch per-lane lag—not cluster CPU—before adding pods.

Traffic burst handling

Peak events (product launches, OTP floods) hit lanes unevenly. The control plane sheds load at ingress before partners rate-limit you.

Normal 2K/min Burst 12K/min After shed 6K/min spike back-pressure Canary rollback if p99 > SLO
Burst detection triggers shed + scale; canaries guard deploys during peaks.

Back-pressure end to end

Spring WebFlux is necessary but not sufficient. Demand must propagate from partner APIs back to ingress.

Ingress Buffer Worker Partner API Ack ← lag meters · RPS caps · canary rollback signals travel left
If the partner slows down, workers apply rate-aware batching and ingress sheds load.
  1. Rate-aware batching before third-party calls.
  2. Micrometer meters for queue depth and consumer lag in Grafana.
  3. Jenkins canaries watch those meters and auto-roll back SLO burns.

SLO dashboard contract

Every lane ships with four tiles before GA:

TileSignalPage if
Ingestion latencyp99 gateway → Kafka> 150 ms for 5 min
Downstream successpartner 2xx ratio< 99% for 10 min
Queue depthconsumer lag> 2× hourly peak
Error budgetburn rate> 2× remaining budget
Pair dashboards with synthetics that replay OTP, ticket assign, and Teams reply every minute. Humans should not be the first detector.

Design tenets and runbooks

  • Isolation: noisy neighbors cannot starve other lanes.
  • Observability: SLOs exist before GA, not after the first outage.
  • Automatability: deploy and rollback via one Jenkins job.
  • Compliance: residency tags (EU/US) travel with the payload.

Pre-built runbooks: drain a lagging partition, fail over to a warm region and rehydrate Redis, collect logs/traces for the post-mortem automatically.