Urgent.News

What's breaking now, across thousands of outlets.

Tech

How Databricks Serverless Compute Cost My Team $14k in One Weekend

It’s Sunday, 2:14 AM. The PagerDuty alert hits my phone with that specific, jarring frequency that makes your stomach drop before you’ve even opened your eyes. My Databricks billing alert wasn’t a standard "usage threshold reached" notification; it was the "you’ve hit 80% of your monthly cloud spend in 48 hours" panic text. I sat up, opened the Databricks console, and stared at the Billing page.…

On a Sunday at 2:14 AM, a PagerDuty alert woke the reporter with a jarring frequency. Instead of a standard usage notification, the alert read: "you've hit 80% of your monthly cloud spend in 48 hours." The reporter opened the Databricks console and saw that the sql_warehouse_prod_v2 was burning DBU (Databricks Units) like a crypto-mining operation.

The team had recently shipped a new pipeline on Friday, and while everything looked fine, the bill was skyrocketing. The dashboard showed a flat line for three months, followed by a steep spike. The reporter first assumed a runaway loop in a Python job, but found no evidence. Checking the spark_query_history yielded no unusual patterns.

The issue lay in the Auto-stop setting: 10 minutes. The reporter realized the warehouse wasn't idling but being kept alive by a ghost. The heartbeat pings from a BI tool, configured with broad CAN USE permissions, were causing the warehouse to remain active. The Auto-stop only triggered when the warehouse was truly idle. With the heartbeat hitting the warehouse every 8 minutes, the 10-minute timer never reset.

The warehouse was being held hostage by a silent, low-latency heartbeat. The team fixed the issue by killing the connection, reducing the warehouse size, updating the connection string, and implementing a Tag policy. They also introduced a Budget Alarm to monitor daily burn rates.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

I Built a Multi-Region Pilot Light on AWS. The Diagram Was the Easy Part.

Most disaster recovery write-ups stop at the diagram. You get two boxes, an arrow labelled "replicate", and an RTO that was never measured.

  • Route 53 monitors health every 10 seconds, switches DNS to secondary region on failure
  • Failover process promotes read replica database, scales Auto Scaling group in both regions

More from Monday 21 September →