Using ChromaDB with Python: Build a Lightweight Local Vector Database for AI Applications

Artificial Intelligence tutorial - IT technology blog
Artificial Intelligence tutorial - IT technology blog

At 2 AM, the PagerDuty alert went off relentlessly. Opening the terminal to check the production server, the top and dmesg commands immediately showed a familiar message: Out of memory: Killed process (python3). The internal document search bot had completely crashed. The culprit? A custom-built vector module that consumed all 16GB of RAM on the backend node.

Production Outage: When Naive Vector Storage Backfires

Earlier, the team had rushed to build a search feature to meet a delivery deadline. Whenever a user submitted a query, the application loaded the entire NumPy array of 50,000 embedding vectors (768 dimensions each) directly into RAM and computed cosine similarity. Everything worked fine until the scheduled data sync last night. The document count doubled to 100,000 records. RAM hit the 16GB ceiling, swap space filled up completely, and the Linux OOM killer terminated the Python process instantly.

When building small-to-medium-scale AI applications, engineering teams often fall into two extremes:

  • Too naive: Storing embeddings in pickle files or in-memory NumPy arrays. This approach works quickly for prototyping, but the CPU bottlenecks severely and memory exhausts once the dataset reaches tens of thousands of records.
  • Too heavy: Setting up a standalone Milvus or Elasticsearch cluster. You will burn an extra 4–8GB of RAM just to keep the infrastructure running, along with the maintenance overhead of networking and cross-service synchronization.

The Root Cause: Why Linear Search Crashes RAM and CPU

In semantic search, each text chunk is encoded into a high-dimensional vector. If you use a NumPy loop to match a query against $N$ vectors, the algorithmic complexity is $O(N)$.

With $N = 500$, the latency is virtually zero. But at $N = 100,000$, the CPU must compute millions of dot products per query. Worse, keeping the entire data matrix in RAM without paging mechanisms, spatial indexing (HNSW), or disk persistence turns the application into a ticking time bomb.

Evaluating Solutions: Faiss, Milvus, or ChromaDB?

During that night shift, three options were placed on the table:

1. Optimize the NumPy Array or Use Pure FAISS

FAISS provides blazing-fast C++ vector computations. However, pure FAISS is merely an algorithmic library, not a database. You still have to write custom code to map IDs to raw text, manage metadata, and handle atomic disk writes across restarts. The risk of data inconsistency remains high.

2. Deploy a Standalone Vector Database Cluster (Milvus, Qdrant)

These solutions are built for enterprise systems managing tens of millions of vectors. However, for an internal bot or a RAG application handling fewer than 500,000 documents, provisioning a Docker cluster with Zookeeper or MinIO only adds unnecessary complexity.

3. Embed ChromaDB Directly into the Python Process

ChromaDB works like SQLite for vector databases. It runs directly within the application process, automatically builds HNSW indexes, maps metadata accurately, and persists data to disk with just a few lines of configuration.

Production-Ready Deployment Guide for ChromaDB PersistentClient

Migrating the vector store to ChromaDB using persistent disk storage (PersistentClient) was the fastest way to rescue the system. Here are the 5 steps to implement it directly in your codebase.

Step 1: Install the Library

Install ChromaDB via pip:

pip install chromadb

Step 2: Initialize Persistent Client for On-Disk Storage

Do not use the default chromadb.Client() in production because data will be lost when the app restarts. Explicitly define a local storage path instead:

import chromadb
from chromadb.config import Settings

# Initialize a persistent client saving data to a local directory
client = chromadb.PersistentClient(
    path="./chroma_data",
    settings=Settings(allow_reset=True, anonymized_telemetry=False)
)

# Create or load an existing collection
# Change the default distance metric from L2 to Cosine
collection = client.get_or_create_collection(
    name="tech_docs",
    metadata={"hnsw:space": "cosine"}
)

Step 3: Load Documents and Generate Embeddings Automatically

ChromaDB comes with the built-in all-MiniLM-L6-v2 model powered by onnxruntime. Simply pass the raw text along with its metadata:

documents = [
    "Guide to configuring high-load Nginx reverse proxy and static caching.",
    "How to resolve Out of Memory (OOM) errors on Linux servers running Python.",
    "Setting up SSH key authentication and disabling password login on Ubuntu 22.04.",
    "Optimizing SQL queries with indexes and analyzing execution plans in PostgreSQL."
]

metadatas = [
    {"category": "devops", "priority": 1},
    {"category": "troubleshooting", "priority": 1},
    {"category": "security", "priority": 2},
    {"category": "database", "priority": 2}
]

ids = ["doc_001", "doc_002", "doc_003", "doc_004"]

# Add batch data to the collection
collection.add(
    documents=documents,
    metadatas=metadatas,
    ids=ids
)

print(f"Total existing records: {collection.count()}")

Step 4: Semantic Query with Metadata Filtering

When a user query arrives, ChromaDB automatically vectorizes the question and returns the most relevant document chunks:

# Search for documents related to server memory overflow errors
results = collection.query(
    query_texts=["How to fix a crashed process caused by server RAM overflow?"],
    n_results=2,
    where={"category": "troubleshooting"}  # Filter accurately by metadata
)

for i, doc in enumerate(results["documents"][0]):
    doc_id = results["ids"][0][i]
    distance = results["distances"][0][i]
    print(f"[{doc_id}] Cosine distance: {distance:.4f} -> {doc}")

Step 5: Perform Update and Delete (CRUD) by ID

When document content changes, you can update or delete it directly by ID without rebuilding the entire index:

# Update with new content and metadata
collection.update(
    ids=["doc_002"],
    documents=["Detailed guide to resolving Linux OOM killer issues and configuring swap files."],
    metadatas=[{"category": "troubleshooting", "priority": 1, "updated": "2026-10-02"}]
)

# Delete obsolete document
collection.delete(ids=["doc_001"])

Practical Troubleshooting When Deploying ChromaDB

Here are the two most common issues you will encounter when bringing ChromaDB to production:

  • SQLite Version Incompatibility on Linux (< 3.35.0): On older operating systems such as CentOS 7 or older Ubuntu releases, you may encounter RuntimeError: Your system has an unsupported version of sqlite3. The cleanest fix is to install pysqlite3-binary and add this override snippet at the very top of your entry point file:
    __import__('pysqlite3')
    import sys
    sys.modules['sqlite3'] = sys.modules.pop('pysqlite3')
    
  • Vector Dimension Mismatch: If you switch from the default embedding model (384 dimensions) to OpenAI text-embedding-3-small (1536 dimensions), never write directly into the existing collection. Create a new collection instead to avoid corrupting the HNSW index structure.

Results after the refactor: Backend node RAM dropped from 16GB to just 280MB. Query latency for 100,000 records stabilized at 8–12ms, permanently resolving the crashing issue.

Share: