Effective Data Management in Microservices: Data Locality, Common Data Replication, and Remote Invocation
Introduction: The Monolithic Problem — the "God Table"
In a monolithic architecture, all customer-related data is often consolidated into a single "Customer" table. Because this table holds data from many domains (e.g., sales, marketing, reservations), it is known as a "God Table".
An example God Table
| CustomerId | Name | Gender | Social Security Number | Purchase History | Campaigns | Current Location | Discount Rate |
|---|---|---|---|---|---|---|---|
| 12345 | John Doe | Male | 123-45-6789 | ["Order1", "Order2"] | ["Campaign1"] | Seoul, South Korea | 10% |
The God Table structure has the following problems.
- Lack of domain cohesion: Distinct problem areas such as sales, marketing, and booking are mixed together in one table.
- Wider change impact: A change to one domain's data affects the other domains as well.
- Scalability problems: Because the data is managed centrally, performance degrades and bottlenecks appear.
Moving to a Microservices Architecture
To solve these problems, a microservices architecture separates data by each domain's **Bounded Context (BC)**. It does so by applying three principles: data locality, common data replication, and remote invocation.
Principle 1: Data Locality
Data locality is the principle of managing locally the data that is frequently accessed or modified within a given service's boundary (context). Each bounded context should contain only the data relevant to its own problem area, and reference data from other domains through an ID value object.
Splitting the God Table into Bounded Contexts
Below is an example of splitting the monolithic "Customer" God Table into three bounded contexts: Sales, Marketing, and Booking.
1. Sales Context (Sales BC)
The Sales BC manages data related to customer purchases and uses a CustomerId value object to reference the customer.
java
class Lead {
private CustomerId customerId; // Customer identifier
private List<String> purchaseHistory; // Purchase history
private double discountRate; // Discount rate
public Lead(CustomerId customerId) {
this.customerId = customerId;
}
// Getters and setters
}CustomerId Value Object
java
class CustomerId {
private final String id;
public CustomerId(String id) {
if (id == null || id.isBlank()) {
throw new IllegalArgumentException("Customer ID cannot be null or blank");
}
this.id = id;
}
public String getId() {
return id;
}
}2. Marketing Context (Marketing BC)
The Marketing BC manages customers' campaign participation and preferred contact channel.
java
class Subscriber {
private CustomerId customerId; // Customer identifier
private List<String> campaignParticipations; // Campaign participation history
private String preferredChannel; // Preferred contact channel (SMS, Email, etc.)
public Subscriber(CustomerId customerId) {
this.customerId = customerId;
}
// Getters and setters
}3. Booking Context (Booking BC)
The Booking BC manages customer booking data, including the customer's current location and booking history.
java
class Passenger {
private CustomerId customerId; // Customer identifier
private List<String> bookingHistory; // Booking history
private String currentLocation; // Current location
public Passenger(CustomerId customerId) {
this.customerId = customerId;
}
// Getters and setters
}Principle 2: Common Data Replication
Each aggregate references the customer through CustomerId, but the customer's immutable data (e.g., name, gender, national ID number) is replicated in each service. This reduces synchronization cost and improves performance.
Example: Common Data Replication
The CustomerId VO
The CustomerId VO contains the customer's immutable data received via events, and each bounded context keeps its own replicated copy.
java
class CustomerId {
private final String id; // Customer ID
private final String name; // Customer name
private final String gender; // Gender
private final String socialSecurityNumber; // National ID number
public CustomerId(String id, String name, String gender, String socialSecurityNumber) {
if (id == null || id.isBlank()) {
throw new IllegalArgumentException("Customer ID cannot be null or blank");
}
this.id = id;
this.name = name;
this.gender = gender;
this.socialSecurityNumber = socialSecurityNumber;
}
// Getters and equals/hashCode methods
}Principle 3: Remote Invocation (the last resort)
Remote invocation carries high network cost and is vulnerable to failures, so it should be used only as a last resort. When data changes frequently or cannot be managed locally, a value object such as CustomerId can apply the Flyweight pattern so that data is fetched only at the moment of the call.
Implementing the Flyweight pattern and a cache in CustomerId
Below is an example in which the CustomerId class fetches remote data when getCurrentLocation is called, then caches it and handles expiry.
java
import java.time.LocalDateTime;
class CustomerId {
private final String id;
private String currentLocation; // Current location (cached)
private LocalDateTime lastFetched; // Time the data was last fetched
private static final long CACHE_EXPIRY_MINUTES = 10; // Cache lifetime (10 minutes)
private final CustomerService customerService;
public CustomerId(String id, CustomerService customerService) {
if (id == null || id.isBlank()) {
throw new IllegalArgumentException("Customer ID cannot be null or blank");
}
this.id = id;
this.customerService = customerService;
}
public String getCurrentLocation() {
if (currentLocation == null || isCacheExpired()) {
currentLocation = customerService.getCustomerLocation(id);
lastFetched = LocalDateTime.now();
}
return currentLocation;
}
private boolean isCacheExpired() {
return lastFetched == null || LocalDateTime.now().isAfter(lastFetched.plusMinutes(CACHE_EXPIRY_MINUTES));
}
}CustomerService implementation
CustomerService is responsible for fetching the customer's current location through a remote call. In this example it provides a simple method that simulates the remote call.
java
class CustomerService {
public String getCustomerLocation(String customerId) {
System.out.println("Fetching location for Customer ID: " + customerId);
return "Seoul, South Korea";
}
}Conclusion
To manage data effectively in a microservices architecture, it is important to follow these principles.
- Data locality: Separate data so that each domain manages only what belongs to its own problem area.
- Common data replication: Replicate data that rarely changes in each service.
- Remote invocation (last resort): Use it only when data changes frequently or cannot be managed locally, and reduce its cost with caching and the Flyweight pattern.
This approach resolves the God Table problem and delivers both efficiency and reliability in data management.