Logistics
Atlas Freight Cloud
Breaking up a monolith without pausing the roadmap
- Industry
- Freight & logistics
- Engagement
- Cloud & DevOps
- Team
- 4 specialists
- Duration
- 22 weeks

The problem
Atlas ran a nine-year-old monolith on oversized virtual machines. Releases happened monthly at weekends, and nobody could explain a latency spike without SSH-ing into individual boxes.
- Monthly weekend releases with a manual 40-step runbook
- Infrastructure provisioned by hand, no reproducible environments
- Cloud bill growing 8% per quarter with no attribution by service
- Mean time to diagnose incidents measured in hours
Our approach
Kubernetes · Terraform · Go
Strangler-fig extraction
We routed traffic through a gateway and peeled off the highest-churn domains first, so the monolith kept serving while services took over one at a time.
Everything in Terraform
Environments became reproducible from code, letting the team spin up an identical staging cluster in minutes and stop drift for good.
Observability before optimisation
Tracing, structured logs and SLO dashboards landed before any performance work, so every change could be measured rather than argued about.
Right-sizing and autoscaling
Workload profiling exposed nodes running at 9% utilisation; autoscaling and spot capacity did the rest of the saving.
Timeline
Architecture audit
Weeks 1-4Dependency mapping, cost analysis, migration sequencing and SLO definitions.
Platform foundation
Weeks 5-11Kubernetes clusters, Terraform modules, CI/CD pipelines and observability stack.
Service extraction
Weeks 12-19Pricing, tracking and notification domains extracted and cut over behind feature flags.
Cost & handover
Weeks 20-22Right-sizing, autoscaling policies, on-call runbooks and platform training.
Results
- 68%
- Reduction in monthly infrastructure spend
- 12x
- Deploys per week, up from one per month
- 9 min
- Mean time to diagnose, down from ~3 hours
- 99.95%
- Availability sustained through migration
- Releases ship during business hours with automated rollback
- Cost is attributed per service, so pricing decisions use real numbers
- Zero customer-facing downtime across every cutover
“Infrastructure costs dropped by two thirds and deploys went from monthly to daily. Their handover documentation was genuinely excellent.”


