
BetterCloud
Platform Engineering / Cloud Migration
Executive brief
BetterCloud’s growth pushed its Mesos/Marathon stack beyond comfortable operational limits at 400M+ events/day. Epsilon delivered a wave-based Kubernetes migration, standardized IaC and GitOps practices, and modernized CI/CD workflows. The result was a more scalable platform, faster delivery, and measurable improvements in cost and performance.
Infrastructure cost reduction
17%
post-migration
Release cycle time
3 days
after modernization
Sustained Daily Events
400M+
validated via load tests
Improved platform scalability for 400M+ events/day workloads
Faster releases through modernized pipelines
Reduced infrastructure cost and improved performance
Standardized operational guardrails with observability
Engagement shape
Duration
6 wks
Phases
4
Deliverables
5
Infrastructure cost reduction
17%
How it ran
Foundation setup + ~30-day service-by-service wave migration
BetterCloud modernized its platform by migrating from Mesos/Marathon to Kubernetes while adopting GitOps and production-grade observability.
Starting position
Rapid growth demanded a platform that could scale reliably beyond 400M events/day while reducing operational friction and improving delivery velocity.
Constraints we worked inside
3 in playEvery one of these ruled an easier option out. They are the reason the approach looks the way it does.
Production system could not experience downtime
Migration had to run in parallel with internal teams
Required staged cutovers with rollback capability
What had to be true to finish
4 goalsAgreed up front, so the engagement could be called finished on evidence instead of opinion.
Migrate from Mesos/Marathon to Kubernetes safely
Adopt GitOps and increase IaC coverage
Modernize CI/CD (updated to GitHub Actions for consistency with current messaging)
Implement guardrails and monitoring
Industry context
Platform Engineering
Cloud Migration
BetterCloud is the category founder in SaaS-Ops, helping enterprises discover, manage, and secure SaaS applications at scale.
Approach
Epsilon executed a wave-based migration to minimize risk while building a repeatable Kubernetes operating model.
Delivery track
Segment width reflects the number of workstreams inside each phase — where the engagement actually spent its effort.
Foundation: Clusters + IaC
3 workstreams
Provisioned two Kubernetes clusters using Terraform
Standardized baseline configuration (RBAC, namespaces, ingress patterns)
Established repeatable environment creation patterns
CI/CD + GitOps Operating Model
3 workstreams
Implemented GitHub Actions pipelines for build/test/deploy
Established PR-driven change control for infrastructure and manifests
Created reusable templates and conventions to reduce variance across services
Wave Migration Execution
3 workstreams
Migrated ~1 microservice/day for ~30 days
Validated each cutover using health, latency, and scaling checks
Documented patterns and trained internal teams to extend the migration
Observability + Guardrails
3 workstreams
Implemented Prometheus, Grafana, and Alertmanager
Standardized dashboards and alert routing by service/team
Established rollout/rollback runbooks and SLO-style indicators
What the team kept
Artefacts handed over at the end of the engagement — owned and operable by BetterCloud without us.
Terraform modules for cluster provisioning
GitHub Actions pipelines and shared workflow templates
GitOps conventions and repo structure
Migration playbook + service cutover checklists
Observability dashboards + alerting rules
Results
BetterCloud gained a more scalable, observable platform with faster delivery and measurable cost/performance improvements.
Outcome ledger
Every figure we measured on this engagement, including the ones that are ranges rather than headlines.
Infrastructure cost reduction
Kubernetes vs Mesos/Marathon (placeholder narrative retained from original)
17%
post-migrationRelease cycle time
Previously just over 6 days (from original text)
3 days
after modernizationSustained Daily Events
Comprehensive battery of load tests
400M+
validated via load testsWhat changed day to day
The part of the result that never shows up in a dashboard, but is the reason the numbers held.
Improved customer satisfaction within the month the engagement completed
Reduced operational friction through standardized deployment and monitoring
Faster and safer rollouts with clear guardrails and visibility
Your turn
BetterCloud modernized its platform by migrating from Mesos/Marathon to Kubernetes while adopting GitOps and production-grade observability. If that shape looks familiar, the first conversation is a working session, not a pitch.
Kubernetes migration
Infrastructure as Code
GitOps implementation
CI/CD modernization
Observability enablement
What BetterCloud got
Infrastructure cost reduction
17%
post-migration
Release cycle time
3 days
after modernization
Sustained Daily Events
400M+
validated via load tests