Микросервисы и «зоопарк» баз данных: почему маленьким командам это не всегда нужно
В конце 2000-х главным вопросом проектирования был «SQL или NoSQL». Прошло полтора десятилетия, и оказалось, что победителя нет — но не потому, что все договорились. А потому, что сам вопрос перестал иметь смысл. Исследование Datadog, построенное на анализе более 2,5 миллиона сервисов в продакшене, показывает, как микросервисная архитектура переопределила то, как компании используют базы данных —…
In the late 2000s, the primary question in designing software was "SQL or NoSQL." Fifteen years later, it appears there is no clear winner, not because everyone agreed, but because the question itself has lost its relevance. A Datadog study, based on analyzing more than 2.5 million production services, reveals how the microservices architecture has redefined how companies use database technologies and the problems that have emerged from this shift.
The study first highlights that more than half of organizations use three or more database technologies simultaneously. A quarter of customers have five or more databases, equaling the number of those still using a single database. Almost half of the companies use both SQL and NoSQL at the same time. Among nine analytical platforms, 44% of organizations utilize at least one of them. Nearly 70% of customers have implemented message queues (leaders include RabbitMQ, Kafka, and AWS SQS).
At first glance, this may seem like a success story: "the right tool for the right task." However, the hidden costs are significant. In a monolithic architecture, the database choice was strategic, as it was a single solution for the entire organization with a shared database acting as an integration point. With the transition to microservices, this choice became tactical: each team is free to choose what works best for their service. While this accelerates development, it has its downside.
Fragmented schemas emerge. A single global schema is scattered into hundreds and thousands of micro-schemas. The database ceases to be a shared artifact seen by developers and analysts and becomes a private implementation detail of a specific service. Joins become problematic, as they once resided within the database and were resolved by a transaction.
Now, data is spread across various storage systems with different schemas, and aggregating it requires work at the application level. Analytics becomes more complex, as data is scattered, and reports must be built by reassembling it in a data warehouse. GraphQL comes into play here, creating a data integration layer that pulls data from multiple stores and APIs and assembles it into a single schema. Logic that once resided inside the monolithic database's joins now moves up to GraphQL resolvers.
The data shows that 55% of services performing GraphQL queries have more than 10 child resolve spans, and some handle over 100 resolves in a single query. The median time for `graphql.execute` is about 200 milliseconds, while `graphql.resolve` takes approximately 40 milliseconds. This difference represents the cost of assembling data from various sources instead of a single join.
A company that has dispersed its data across multiple databases ultimately ends up having to build a separate layer to gather this data again, a task that used to be handled by a single database.
The issue is particularly acute for small companies. Those using five or more databases are mostly medium to large organizations. Medium-sized companies often started as microservices from scratch, while large companies migrated from monoliths and optimize different workloads with various tools. A small team finds this zoo of three to five databases problematic: each developer has to monitor, backup, update, and tune a different database.
This maintenance consumes more resources than the flexibility it provides. Joins no longer exist; they have moved up and become more expensive, replaced by chains of resolver functions, caches, and batching.
The question is whether multiple databases are justified at this scale. Is a multi-database setup truly necessary or just a price for scaling that a small team may not yet need? Practical takeaways include: don't start with a zoo of databases. A well-chosen single database, often PostgreSQL, can cover 90% of tasks for a small team.
Consider the cost of integration upfront. Each new database adds another join through an application, leading to a GraphQL layer or ETL pipeline, which is not free. GraphQL is a symptom, not a solution. If you need a layer that stitches data from multiple stores, the question is whether all those databases were necessary in the first place.
Datadog's findings suggest that "microservices have made the SQL vs NoSQL debate irrelevant." However, for small teams, the more nuanced truth is that database diversity is not an end in itself, but a cost of scaling that the team may not yet require. Before moving joins from the database to GraphQL resolvers, ask: why are the data dispersed there, and why do we need to bring them back?
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.