Scaling Cloud Infra for Millions

A reference for CTOs, engineering leads, and technical decision-makers evaluating
Stacknize as an engineering partner

“Scale” looks different on every platform. For a telecom operator, it’s a campaign that hits millions of subscribers at once. For OTT, it’s the first minutes of a festival release. For healthcare, it’s close to 5 million advisory messages a month alongside tele-consults and ambulance bookings that can’t fail.

Stacknize has spent fourteen years building platforms under those conditions, across telecom, media, gaming and healthcare. Here’s how we approach scaling, and what CTOs should ask any engineering partner.

1. Design for the shape of the load, not the average

Most scaling failures share one root cause: the system was sized for average traffic. Real traffic spikes with operator campaigns, festival releases, tournament start times and the morning OPD rush.

So we start with the load profile: the worst hour, how fast traffic ramps, and which actions are expensive. From that we set explicit targets for peak throughput, p95/p99 latency and recovery time. Those numbers drive the architecture.

2. The architecture patterns we rely on

  • Stateless, horizontally scaled services. Capacity grows and shrinks automatically; unavoidable session state (USSD, SIP) lives in a fast shared store.
  • Queues in front of anything that bursts. SMS, notifications, billing, transcoding and AI inference are processed at a steady rate, so a spike becomes a backlog, not an outage.
  • Idempotency wherever money or messages move. Retries never double-bill a subscriber or double-book an ambulance.
  • Cache and edge for read-heavy paths. CDNs and adaptive bitrate keep video playing on patchy mobile networks.
  • Data scaled deliberately. Read replicas, separate analytics stores, and partitioned high-volume tables.
  • Graceful degradation. A recommendations panel can disappear briefly; a payment or emergency booking cannot.
  • Documented decisions. Architecture decision records mean your team inherits the reasoning, not just the code.

3. Telecom-grade and emerging-market realities

Generic cloud playbooks assume fast networks and one region. Many of our platforms don’t have that luxury.

  • Hybrid by necessity. Telecom VAS often runs partly inside the operator network (SMSC, HLR, SIP, charging) and partly in the cloud. Latency-sensitive work sits close to the network; bursty work runs in the cloud.
  • Protocols with hard timing. USSD and SIP paths get their own latency budgets, dedicated capacity and replay-based load tests.
  • Unreliable networks. Small payloads, retry-tolerant apps, and SMS or USSD fallbacks where data is weak.
  • Data residency. Regions, backups and replication are planned around in-country rules from day one.

4. Reliability is an operating discipline, not a feature

  • Observability from day one. Logs, metrics and traces ship with the first service, tracking what users feel: session success, playback start, message delivery, booking completion.
  • SLOs agreed with the business, so alerts reflect real user impact, not noise.
  • Zero-downtime deployments. Infrastructure as code, CI/CD, rolling or blue-green releases and feature flags, with fast rollback.
  • Load testing before the peak, at and beyond expected traffic.
  • Honest incident management. Early communication, service restored first, then a blameless review with concrete fixes.
  • Code review on every change.

5. Cost and security scale too

Cost. Where revenue per user is measured in cents, a platform that scales technically but not economically has still failed. We scale down as aggressively as up, tier storage, use reserved or spot capacity where it fits, and track unit cost per subscriber, stream or consultation.

Security and compliance. Encryption, least-privilege access, secrets management and audit logs are defaults, designed around laws such as Nigeria’s NDPA, Kenya’s Data Protection Act and India’s DPDP Act.

6. What this looks like in production

  • Telecom VAS and ring back tones integrated with operator SIP, SMSC, USSD, HLR and billing. [Add: operators, subscribers, peak TPS.]
  • OTT video for 5+ clients across Africa and Asia, each tuned to its audience’s peaks. [Add: peak concurrent viewers.]
  • Gaming that orchestrates single-player, multiplayer and tournament formats from multiple providers. [Add: players.]
  • Healthcare (MediTrust): 150,000+ advisory subscribers, close to 5 million messages a month, and hundreds of tele-consults and ambulance bookings daily, live in production.
  • AI call centre with real-time voice, an LLM conversation engine and confidence-based escalation to humans. [Add: calls per day.]

7. How we work with engineering teams

Discovery first, then agreed SLO, capacity and cost targets, then small reviewed increments with CI/CD, load and failure testing before go-live, and continued involvement after launch or a documented handover. We can own a platform end to end, extend your team, or review an existing system.

Questions to ask any engineering partner

  • What was your worst production incident, and what changed afterwards?
  • How do you set and test capacity targets before launch?
  • How do you deploy without downtime, and how fast can you roll back?
  • How do you track cloud cost per user?
  • What documentation and decisions do we own at the end?

We’re happy to answer every one, with examples.

The bottom line

Scaling for millions isn’t one clever technology choice. It’s understanding how users really behave, applying a few proven patterns, and operating them with discipline. The stack gets you through launch day; the engineering practice gets you through your busiest day.

Planning a launch, migration or a platform straining under growth? Let’s talk about your load profile, and we’ll give you an honest view.


Stacknize builds production platforms across telecom and voice, OTT streaming, gaming, mHealth, and AI, for clients across Africa, Asia, and beyond. If you’re evaluating an engineering partner for your next platform, let’s talk.