Urgent.News

What's breaking now, across thousands of outlets.

AI

Giving AI Agents Lakehouse Memory: How We Built an Apache Iceberg Connector for Cognee

Giving AI Agents Lakehouse Memory: How We Built an Apache Iceberg Connector for Cognee Author: Soumyajit Ghosh ( @somuai ) Target Publication: Dev.to / Substack Hackathon: Mergetober (WeMakeDevs x Cognee) Pull Request: topoteretes/cognee-community#348 (Closes topoteretes/cognee#5553 ) 1. The Blind Spot in Enterprise AI: Lakehouse Memory Modern autonomous AI agents are rapidly moving from toy chat…

Title: Enabling AI Agents with Lakehouse Memory: Building an Apache Iceberg Connector for Cognee

Modern autonomous AI agents are becoming critical operational tools, handling complex tasks like diagnosing ETL pipelines and answering developer queries about data schemas. However, most LLM applications lack awareness of the underlying data lakehouse infrastructure. Apache Iceberg, a standard open table format, provides essential features like ACID transactions, scalable metadata, and snapshot time travel. To bridge this gap, I developed an Apache Iceberg connector for Cognee as part of the Mergetober Hackathon.

Cognee is an open-source memory engine designed for AI agents. Unlike traditional LLM applications, Cognee ingests data through resilient pipelines, runs entity-extraction and graph-construction processes, and stores the resulting knowledge in a Knowledge Graph. This approach enables AI agents to reason across evolving lakehouse schemas by traversing complex relationships and recalling historical context.

The connector architecture focuses on two main design decisions: metadata-first semantic ingestion and cognee document-mode routing. By treating metadata as architectural prose, Cognee can extract semantic relationships between tables, columns, and partitions, storing them as explicit graph nodes. This approach prevents the hallucination of tables that have been decommissioned, ensuring the AI agent maintains a consistent understanding of the data lakehouse.

A key challenge in building the Iceberg connector was implementing a full-snapshot sync and forget-on-delete strategy. This approach ensures that the knowledge graph remains up-to-date by replacing staged data with the current set of tables in the catalog. When a table is dropped upstream, the connector automatically removes it from the knowledge graph and vector indices, preventing ghost knowledge and accidental data deletion.

The connector was thoroughly tested in a live execution against an Iceberg REST catalog with multiple namespaces and partitioned tables. Upon simulated table decommissioning, the connector successfully updated the knowledge graph, demonstrating its ability to maintain a consistent understanding of the data lakehouse. This solution provides AI agents with a robust lakehouse memory, enabling them to perform complex tasks and make informed decisions without relying on outdated or incomplete information.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

More from Friday 9 October →