Monolith to Microservices: A Step-by-Step Migration Guide from Production
I’ve led three monolith-to-microservices migrations across banking, logistics, and healthcare. Two went well. The first one taught me what not to do.
The internet is full of theoretical microservices migration guides — boxes, arrows, and phrases like “just extract a bounded context.” This isn’t that. This is the step-by-step playbook I actually follow, including the mistakes that quietly destroy real migrations.
Start with Why — and Make Sure the Answer Isn’t “Because Netflix Did It”
Before writing a line of code, answer one question honestly: What specific pain is the monolith causing?
Valid reasons to migrate:
- Deployment coupling. Changing one module requires deploying the entire application, with a 3-hour regression cycle.
- Team bottlenecks. Multiple teams editing the same codebase, stepping on each other’s changes.
- Scaling constraints. One CPU-heavy module needs 8x resources, but it’s bundled with 20 other modules that don’t.
Invalid reasons:
- “Microservices are modern.”
- “We want to use Kubernetes.”
- “Our new CTO read a blog post.”
If your monolith deploys cleanly and your team ships features without friction, keep the monolith. It’s not broken.
Phase 1: Stabilize Before You Split
This is the step most teams skip — and it’s the one that matters most.
You cannot safely decompose a system you don’t understand. Before extracting a single service:
- Add comprehensive test coverage to the areas you plan to extract. If the monolith has 20% test coverage, you’re flying blind during extraction.
- Map the dependency graph. Which modules call which? Which share database tables? Draw it. Every shared table is a migration landmine.
- Introduce a CI/CD pipeline if one doesn’t exist. You need automated builds, test runs, and deployments before you increase the number of deployable units.
Before migration:
├── 1 deployable artifact
├── 1 database
├── 0 automated tests (honesty hurts)
└── Manual deployment via SSH
After Phase 1 (still a monolith, but safer):
├── 1 deployable artifact
├── 1 database
├── 200+ automated tests covering extraction targets
└── Jenkins pipeline with automated build/test/deploy
On one banking project, Phase 1 alone took 6 weeks. It felt slow. It saved us months of debugging later.
Phase 2: Strangler Fig — Extract at the Edge
The Strangler Fig pattern is the only migration strategy I recommend for production systems. The idea is simple:
- Identify a module at the edge of the monolith — one with clear boundaries and minimal shared state.
- Build the new service alongside the monolith, not as a replacement.
- Route new traffic to the new service via an API gateway or reverse proxy.
- Keep the old code running until the new service is proven in production.
- Delete the old code only after weeks of stable parallel operation.
Choosing What to Extract First
Pick the module that has:
- The least database coupling. If it shares 15 tables with other modules, it’s not your first extraction.
- Clear API boundaries. If other modules call it through well-defined interfaces (not direct method calls across packages), it’s a good candidate.
- Independent deployment value. Extracting it should unlock something — faster deploys, independent scaling, or team autonomy.
In a logistics system, our first extraction was the notification service. It had its own tables, a clear input contract (events from ActiveMQ), and zero shared state with the core shipping engine. Perfect first candidate.
Phase 3: Database Decomposition — The Hard Part
This is where most migrations fail. The monolith’s database is a shared mutable state machine, and every table join is a hidden coupling.
Strategy: Database per Service
Each extracted service gets its own database (or schema). No cross-service joins. No shared tables.
Before:
monolith_db
├── users
├── orders
├── shipments
├── notifications
└── audit_logs
After:
order_service_db → orders, order_items
shipment_service_db → shipments, tracking
notification_service_db → notifications, templates
shared (read-only views) → users (until user service is extracted)
The Transition Pattern
During migration, you’ll inevitably have a period where both the monolith and the new service need the same data. Options:
- Database views (read-only). The new service reads from a view on the monolith’s database. Temporary, but safe.
- Event-driven sync. The monolith publishes change events (via ActiveMQ, Kafka, or a simple outbox table). The new service consumes them and maintains its own copy.
- API calls. The new service calls back to the monolith’s API for data it doesn’t own yet. Simple, but introduces runtime coupling.
I prefer option 2 for anything non-trivial. It’s more work upfront, but it eliminates runtime dependency on the monolith and sets up the event-driven pattern you’ll need long-term.
Phase 4: Inter-Service Communication
For synchronous communication between services, use REST APIs with clear contracts:
// Service A calls Service B
@FeignClient(name = "shipment-service", url = "${shipment.service.url}")
public interface ShipmentClient {
@GetMapping("/api/shipments/{id}")
ShipmentDto getShipment(@PathVariable String id);
@PostMapping("/api/shipments")
ShipmentDto createShipment(@RequestBody CreateShipmentRequest request);
}
For asynchronous communication (events, notifications, non-critical workflows), use a message broker:
// Publishing a domain event
@Service
@RequiredArgsConstructor
public class OrderService {
private final JmsTemplate jmsTemplate;
@Transactional
public Order createOrder(CreateOrderRequest request) {
Order order = orderRepository.save(mapToOrder(request));
// Publish event for downstream services
jmsTemplate.convertAndSend("order.created",
new OrderCreatedEvent(order.getId(), order.getCustomerId()));
return order;
}
}
Rule of thumb: If Service A can’t function without an immediate response from Service B, use synchronous. If Service A just needs to notify Service B that something happened, use async.
Common Mistakes That Break Real Migrations
After three migrations, these are the patterns I’ve seen cause the most damage:
1. Big Bang Extraction
Extracting 5 services at once, cutting over on a Friday. Don’t. Extract one service, stabilize it for 2–4 weeks, then extract the next.
2. Distributed Monolith
You extracted 8 services, but they all share one database and call each other synchronously for every operation. Congratulations — you now have a monolith with network latency.
3. Ignoring Data Consistency
In a monolith, a single database transaction guarantees consistency. In microservices, you need to design for eventual consistency. If your business logic requires “order created AND shipment scheduled” as an atomic operation, you need a saga pattern or an outbox table — not a distributed transaction.
4. No Observability
With one monolith, grep on one log file finds anything. With 8 services, you need centralized logging, distributed tracing (correlation IDs), and health check endpoints before you go live.
The Timeline Reality
| Phase | What | Duration |
|---|---|---|
| Phase 1 | Stabilize (tests, CI/CD, dependency mapping) | 4–8 weeks |
| Phase 2 | Extract first service (strangler fig) | 3–6 weeks |
| Phase 3 | Database decomposition for first service | 2–4 weeks |
| Phase 4 | Validate in production, extract second service | 4–6 weeks |
| Ongoing | One service at a time, 3–6 week cycles | Months to years |
A full migration of a medium-sized monolith takes 6–18 months. Anyone promising 6 weeks is selling you a rewrite, not a migration.
Final Thought
The goal of a monolith-to-microservices migration is not “microservices.” The goal is faster, safer, more independent delivery. If you can achieve that by cleaning up the monolith’s internal architecture, do that instead. Microservices are a tool, not a destination.