Setting Up and Managing Milvus with Attu: Deploying a Vector Database for Large-Scale RAG

Artificial Intelligence tutorial - IT technology blog
Artificial Intelligence tutorial - IT technology blog

Quick start: Set up Milvus Standalone and Attu GUI in 5 minutes

When your RAG knowledge base exceeds 5–10 million vectors, in-memory libraries like FAISS or Chroma start consuming excessive RAM and become difficult to scale across clusters. Milvus thoroughly solves this problem with its distributed architecture. You can quickly spin up the entire environment using Docker Compose.

Create a docker-compose.yml file containing all 4 components: Milvus server, MinIO (stores index files), Etcd (stores metadata), and Attu (web GUI):

version: '3.5'

services:
  etcd:
    container_name: milvus-etcd
    image: quay.io/coreos/etcd:v3.5.5
    environment:
      - ETCD_AUTO_COMPACTION_MODE=revision
      - ETCD_AUTO_COMPACTION_RETENTION=1000
      - ETCD_QUOTA_BACKEND_BYTES=4294967296
      - ETCD_SNAPSHOT_COUNT=50000
    volumes:
      - ./volumes/etcd:/etcd
    command: etcd -advertise-client-urls=http://127.0.0.1:2379 -listen-client-urls=http://0.0.0.0:2379 --data-dir=/etcd

  minio:
    container_name: milvus-minio
    image: minio/minio:RELEASE.2023-03-20T20-16-18Z
    environment:
      MINIO_ACCESS_KEY: minioadmin
      MINIO_SECRET_KEY: minioadmin
    volumes:
      - ./volumes/minio:/minio_data
    command: minio server /minio_data
    healthcheck:
      test: ["CMD", "curl", "-f", "http://localhost:9000/minio/health/live"]
      interval: 30s
      timeout: 20s
      retries: 3

  standalone:
    container_name: milvus-standalone
    image: milvusdb/milvus:v2.3.4
    command: ["milvus", "run", "standalone"]
    environment:
      ETCD_ENDPOINTS: etcd:2379
      MINIO_ADDRESS: minio:9000
    volumes:
      - ./volumes/milvus:/var/lib/milvus
    ports:
      - "19530:19530"
      - "9091:9091"
    depends_on:
      - "etcd"
      - "minio"

  attu:
    container_name: milvus-attu
    image: zilliz/attu:v2.3.4
    ports:
      - "8000:3000"
    environment:
      MILVUS_URL: milvus-standalone:19530
    depends_on:
      - "standalone"

Start the entire stack with a single command:

docker compose up -d

Wait about 30 seconds for the services to stabilize, then open your browser at http://localhost:8000. Enter the host milvus-standalone:19530 to log in to the Attu dashboard.

Milvus Architecture and the Role of Attu

1. How Milvus Works

Milvus completely decouples compute from storage. Thanks to this architecture, you can independently scale out Query Nodes when search traffic surges:

  • Etcd: Stores collection metadata, schemas, partition details, and cluster state.
  • MinIO / S3: Stores raw vectors, log files, and index files. Even if compute nodes crash and restart, data loss is prevented.
  • Query Node & Data Node: Query Nodes load indexes into RAM to deliver millisecond-level search latencies. Data Nodes batch write (insert) data into immutable segments.

2. How Attu Helps with Operations

Attu is the official GUI developed by Zilliz. This interface eliminates the need to write manual scripts every time you inspect data:

  • Visually create, edit, and delete collections and partitions.
  • Check the in-memory loading status (Loaded/Unloaded) of each collection.
  • Quickly test vector search queries and scalar metadata filtering directly from the web interface.
  • Monitor entity counts and segment sizes in real time.

Data Operations with PyMilvus and Index Optimization

Two key factors determine RAG performance: designing a flexible schema with Dynamic Fields and choosing the right indexing algorithm.

1. Initializing a Collection for RAG

Below is a Python script to create a collection storing document chunks and 1536-dimensional vectors (standard for OpenAI text-embedding-3-small):

from pymilvus import connections, FieldSchema, CollectionSchema, DataType, Collection

# Connect to Milvus
connections.connect("default", host="localhost", port="19530")

collection_name = "rag_enterprise_docs"

# Define Schema
fields = [
    FieldSchema(name="id", dtype=DataType.INT64, is_primary=True, auto_id=True),
    FieldSchema(name="doc_id", dtype=DataType.VARCHAR, max_length=64),
    FieldSchema(name="content", dtype=DataType.VARCHAR, max_length=4096),
    FieldSchema(name="embedding", dtype=DataType.FLOAT_VECTOR, dim=1536)
]

# Enable enable_dynamic_field to flexibly store arbitrary JSON metadata
schema = CollectionSchema(fields, description="RAG Knowledge Store", enable_dynamic_field=True)
collection = Collection(name=collection_name, schema=schema)

print(f"Collection {collection_name} is ready!")

2. Choosing an Index: HNSW or IVF_FLAT?

Without an index, Milvus must perform a brute-force scan (Flat search). On 1 million vectors, queries will suffer latency in the hundreds of milliseconds. Practical options include:

  • HNSW: Extremely fast search speed (typically < 5ms for 1 million vectors) with 98–99% recall. Trade-off: high RAM consumption for graph construction.
  • IVF_FLAT: Clusters vectors around centroids. It consumes less RAM than HNSW, making it suitable for moderate server hardware.

Create an HNSW index with Cosine distance metric:

index_params = {
    "metric_type": "COSINE",
    "index_type": "HNSW",
    "params": {"M": 16, "efConstruction": 200}
}

collection.create_index(field_name="embedding", index_params=index_params)

# Load collection into RAM after building the index
collection.load()
print("HNSW index built successfully; collection loaded into RAM.")

3. Hybrid Search: Combining Metadata Filtering with Vector Similarity

In practice, you often need to find relevant chunks that belong to a specific document or department. Use the expr expression for filtering:

query_vector = [0.015] * 1536

search_params = {"metric_type": "COSINE", "params": {"ef": 64}}

results = collection.search(
    data=[query_vector],
    anns_field="embedding",
    param=search_params,
    limit=3,
    expr='doc_id == "finance_q1_2026"', # Filter metadata before computing similarity
    output_fields=["doc_id", "content"]
)

for hits in results:
    for hit in hits:
        print(f"Score: {hit.distance:.4f} | Content: {hit.entity.get('content')}")

Production Best Practices for Milvus

  • Always call collection.load(): Milvus can only query segments that have been loaded into RAM. If queries return empty results, open Attu to check whether the collection status is Loaded.
  • Estimate RAM requirements for HNSW: Use the formula: Number of vectors * Dimensions * 4 bytes * 1.5. For example, 5 million 1536-dimensional vectors require approximately 5,000,000 * 1536 * 4 * 1.5 ≈ 46 GB RAM. Provision a server with at least 64 GB of RAM.
  • Limit the number of partitions: Keep fewer than 64 partitions per collection. Too many small partitions cause memory fragmentation on Query Nodes. For granular filtering, rely on dynamic metadata and expr filter expressions.
  • Enable periodic Etcd compaction: Failing to clean up snapshots will cause Etcd storage to exceed its 2GB/4GB quota and crash the cluster. Always configure ETCD_AUTO_COMPACTION_RETENTION=1000.
  • Monitor via Prometheus: Milvus exposes a metrics endpoint on port 9091/metrics. Scrape this into Grafana to monitor QPS, query latency, and unindexed segment counts.
Share: