Scaling • Observability • Messaging
Scaling Omnichannel Services to 100K Daily Interactions
Latency spikes in messaging become revenue hits. This is the architecture we used at Freshworks to keep WhatsApp, email, SMS, voice, and Teams reliable as traffic tripled.
Plan capacity by conversation lane
Do not size “the messaging service.” Size lanes: SMS, WhatsApp, email, voice, Teams. Each lane has a different partner SLA, burst shape, and compliance boundary. Mixing them on one Kafka topic creates noisy-neighbor outages.
During marketing bursts, WhatsApp lanes may 5× while SMS stays flat. Autoscaling hooks watch per-lane lag—not cluster CPU—before adding pods.
Traffic burst handling
Peak events (product launches, OTP floods) hit lanes unevenly. The control plane sheds load at ingress before partners rate-limit you.
Back-pressure end to end
Spring WebFlux is necessary but not sufficient. Demand must propagate from partner APIs back to ingress.
- Rate-aware batching before third-party calls.
- Micrometer meters for queue depth and consumer lag in Grafana.
- Jenkins canaries watch those meters and auto-roll back SLO burns.
SLO dashboard contract
Every lane ships with four tiles before GA:
| Tile | Signal | Page if |
|---|---|---|
| Ingestion latency | p99 gateway → Kafka | > 150 ms for 5 min |
| Downstream success | partner 2xx ratio | < 99% for 10 min |
| Queue depth | consumer lag | > 2× hourly peak |
| Error budget | burn rate | > 2× remaining budget |
Design tenets and runbooks
- Isolation: noisy neighbors cannot starve other lanes.
- Observability: SLOs exist before GA, not after the first outage.
- Automatability: deploy and rollback via one Jenkins job.
- Compliance: residency tags (EU/US) travel with the payload.
Pre-built runbooks: drain a lagging partition, fail over to a warm region and rehydrate Redis, collect logs/traces for the post-mortem automatically.