Moving Your Local Airflow to GCP for under $150 a month
My team runs ML training and prediction pipelines on Vertex AI. For a long time, the thing telling those pipelines when to run was Apache Airflow, installed by hand on a few on-prem Linux boxes. From there it orchestrated our hybrid cloud setup from one place. Next to Airflow sat our own configuration management system. It lets people change a pipeline's settings on demand, and every change is…
My team utilizes ML training and prediction pipelines on Vertex AI. Previously, we relied on Apache Airflow, which was installed manually on a few on-prem Linux servers. This orchestrator handled our hybrid cloud setup. Alongside Airflow, we employed a configuration management system for changing pipeline settings on demand, with all changes tracked and logged.
We needed to retire these on-prem systems as our organization shifts to a fully cloud-native setup. We are currently in the process of migrating our on-prem pipelines to various clouds, including GCP. Given that the code resides in GCP, it made sense to place the orchestrator there as well. Moving the orchestrator to GCP should not compromise safety or reliability.
This article shares our approach, broken down step by step. The code is available in the self-managed-airflow-on-gcp repository, and everything below is a simplified, functional version of it.
In Chapter 1, we initially explored Cloud Composer, which is a managed version of Airflow. Google manages the scheduler, database, and workers, allowing you to drop DAGs into a bucket. We set it up using Terraform, pointed our DAGs at it, and it worked. However, the cost was prohibitive. A small Composer environment costs around $350 a month, as the pricing calculator shows.
Composer 3 charges for compute units every hour the environment is active, even if no DAG runs. Our pipelines do not require a cluster running continuously to schedule a few jobs, since the heavy lifting occurs in Vertex AI. Therefore, we continued searching for a more cost-effective solution.
Chapter 2 introduced self-managed Airflow on a VM. Our Airflow does not handle the heavy lifting; it only determines when work occurs and forwards it to Vertex AI or BigQuery. After setting up a Compute Engine VM with a self-managed Airflow 3.x instance, we used Postgres for the metadata database. Terraform was used to provision the platform, with per-environment setups (dev, test, prod) managed through a Fabric script for installation, deployment, and daily maintenance.
The cost estimate for this setup was under $150 a month. Almost $20 of this expense goes toward the HTTPS Load Balancer in front of our configuration app (GCP charges a flat $0.025/hour for the first five forwarding rules, roughly $18/month, plus data). The remaining cost is allocated to the VM, its disk, and Cloud NAT. While this price is lower than Cloud Composer, it's not the final cost.
With Cloud Composer, Google handles Airflow upgrades and patches. In contrast, on a VM, we are responsible for these tasks. The rest of this article discusses how we achieved this cost-saving while maintaining safety and reliability.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.