Vector database showroom. Part 2: Chroma โ The E-Scooter That Starts in 90 Second
๐ด Chroma โ The E-Scooter That Starts in 90 Seconds pip install chromadb โ and you're already riding. No server, no ports, no API keys. Under the hood: an open-source embedding database (Apache 2.0, Python-first, Rust core since v0.4). Embedded mode is the default: use PersistentClient to write your data to SQLite + Parquet right on disk (while the basic Client() is ephemeral and loses data onโฆ
Chroma is an e-scooter that starts up in just 90 seconds and is a vector database designed for prototyping RAG (retrieval-augmented generation). Installing Chroma with `pip install chromadb` is all it takes to start riding. The database features an open-source embedding database with a Rust core since version 0.4 and Apache 2.0 licensing.
By default, Chroma is embedded, meaning data is written directly to SQLite and Parquet on disk, with the option to turn it into an HTTP server or container. It requires no API keys, servers, or ports, and offers first-class LangChain and LlamaIndex integrations. Chroma supports metadata filtering, document content filtering, and cosine similarity instead of the default L2 distance.
The HNSW index lives in RAM, with 1 million 1536-dimensional vectors taking up approximately 6GB of data. While the scooter is convenient and free under the Apache 2.0 license, it has limitations such as no distributed mode, single-node design, and lacks advanced features like BM25 hybrid, quantization, and RBAC. Additionally, the storage format has changed between major versions, and there is no built-in alarm system or token authentication.
Chroma is best suited for prototyping and personal tools that won't be exposed to the internet, but for production use with paying users, a more robust solution is recommended.
Written by urgent.news from Dev.to's reporting โ not their text. Machine-written โ may contain errors; check the original before relying on it.