Urgent.News

What's breaking now, across thousands of outlets.

Tech

Introduction to Data Lakes Part 2

En el post anterior exploramos qué es un Data Lake y por qué son tan importantes en el ecosistema de datos actual. Ahora es momento de ensuciarnos las manos y ver exactamente qué servicios de AWS necesitamos para construir un Data Lake completamente serverless y cómo orquestarlos. Los Servicios Fundamentales Un Data Lake serverless en AWS se construye sobre cinco pilares fundamentales que…

Translated from Spanish Read in Spanish

The article discusses how to build a completely serverless Data Lake in AWS. It is based on five key pillars: Storage, Processing, Catalog, Security, and Exploitation. Amazon S3 is used for storage, with a recommended folder structure, and key configurations include versioning, lifecycle policies, server-side encryption, and cross-region replication.

AWS Glue is used for data transformation, with components including Glue Jobs, Glue Catalog, and Glue Crawlers. Amazon Athena allows for querying data directly from S3 using SQL, with benefits including pay-per-query and integration with Glue Catalog. AWS Lambda is used for automation, responding to events and orchestrating workflows.

Written by urgent.news from Dev.to's report — not a translation of it. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

Building Distributed Systems in Elixir: Part 6 — Named Processes

In the previous part of this series, we built a tiny supervisor from scratch. When a worker crashed, the supervisor started a replacement.

  • Named processes provide discoverable addresses for workers in distributed systems.
  • Process.register/2 associates a PID with a local name on a BEAM node.
  • Process.whereis/1 retrieves the current PID for a registered local name.

More from Wednesday 19 August →