Urgent.News

What's breaking now, across thousands of outlets.

Tech

Database Sharding para Microservicios MACH: Estrategias de Particionado a Escala Global

En arquitecturas MACH de escala global donde un microservicio individual maneja decenas de millones de registros y miles de transacciones por segundo, el escalado vertical de la base de datos (aumentar CPU y RAM del servidor) tiene un limite fisico y economico. El Database Sharding (particionado horizontal) es la tecnica que permite escalar una base de datos mas alla de las capacidades de un…

In massively scaled MACH architectures, where a single microservice manages tens of millions of records and thousands of transactions per second, vertical scaling of the database (increasing CPU and RAM) reaches physical and economic limits. Database sharding, a horizontal partitioning technique, enables database scaling beyond the capabilities of a single server by distributing data across multiple instances.

This article examines sharding strategies specifically in the context of MACH microservices, where each service owns its own database and sharding is a decision made exclusively by the team responsible for that service. When to Apply Sharding in a MACH Microservice? The first principle of sharding is not to implement it prematurely.

Sharding adds significant operational complexity and should only be justified when other scaling techniques have been exhausted. The recommended scaling order for a database in MACH is: Query optimization (correct indexes, elimination of N+1 queries, rewriting expensive queries) - This can improve performance by up to 100 times without structural changes.

Caching (Redis or Memcached in front of the most frequent queries) - Eliminates 70-90% of read query load in most workloads. Read replicas - Add read replicas to distribute read query load. Scale read capacity up to 5-10 times. Vertical scaling - Move to an instance with more CPU and RAM. A quick but costly solution with physical limits.

Sharding - When none of the previous techniques are sufficient. Indicators that a microservice needs sharding include: write latency consistently above 50ms despite correct indexes, the main database cannot absorb more read replicas without degrading replication, data volume exceeds 10TB and affects backup/recovery times, or write throughput exceeds 50,000 transactions per second sustained.

Sharding Strategies: Hash, Range and Directory There are three main sharding strategies, each with different characteristics. Hash Sharding: In hash sharding, the shard to which a record belongs is determined by applying a hash function to the shard key value. For example, if the shard key is customer_id and there are 4 shards, the corresponding shard number is: shard_number = hash(customer_id) % 4.

The advantage of hash sharding is uniform distribution of data across shards, preventing "hot" shards where one shard receives much more traffic than others. The disadvantage is that range queries (all orders between date A and date B) must consult all shards in parallel, which is less efficient than range sharding. Range Sharding: In range sharding, records are distributed across shards based on ranges of the shard key value.

For example, orders with order_id between 1 and 10 million go to shard 1, between 10 and 20 million to shard 2, etc. The advantage is that range queries are efficient because they only need to consult a specific shard. The disadvantage is hot shards: if new data always falls into the highest range (common with auto-incrementing IDs or timestamps), the last shard receives all the write load.

Directory Sharding: In directory sharding, there is a centralized lookup table that maps each shard key value to its corresponding shard. This is the most flexible approach as it allows rebalancing shards without changing application logic, but adds an additional network call for each database operation. Implementation with Citus (PostgreSQL Distributed): For MACH microservices using PostgreSQL, Citus is the most mature and widely adopted sharding extension.

Citus transforms PostgreSQL into a distributed database that implements sharding transparently to the application. With Citus, the same ORM or PostgreSQL client used by the microservice can work with a sharded database without code changes in the application. The developer defines which column is the distribution key (equivalent to the shard key) when creating the table, and Citus handles automatic data distribution and query routing to the correct shard.

The choice of distribution key is the most critical decision. It must be a field that appears in all WHERE clauses of the microservice queries, have high cardinality (many different unique values), and evenly distribute traffic across shards. In an order service, tenant_id or customer_id are good distribution keys. Cross-Shard Queries: The Major Sharding Challenge Cross-shard queries, which require data from multiple shards, are the major operational challenge of sharding and the primary reason for careful distribution key selection.

In Citus, cross-shard queries are executed in parallel across all shards, with results aggregated on the coordinator node. This is automatic but incurs latency costs proportional to the number of shards and the amount of data transferred between nodes for aggregation. The pattern to minimize cross-shard queries in MACH is co-location: when two frequently queried tables share the same distribution key, Citus ensures related records are always on the same shard.

For example, if orders and order_items use customer_id as their distribution key, all orders and their items for a customer are always on the same shard, allowing efficient JOINs without cross-shard traffic. Sharding in the MACH Ecosystem: Coordination Between Teams In a pure MACH architecture, database sharding of a microservice's database is an internal decision made by the team responsible for that service. No other team is involved.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

How I’m Making Website Feedback Easier for Developers and Agencies

Have you ever received website feedback like “this section looks wrong” without knowing which section the client was talking about? I built PinReview to solve this problem.

  • PinReview simplifies website feedback for developers and agencies
  • Users pin comments to specific page spots, turning feedback into tasks
  • Embeddable widget and browser extension collect feedback easily

More from Friday 9 October →