Quick start: Set up Milvus Standalone and Attu GUI in 5 minutes
When your RAG knowledge base exceeds 5–10 million vectors, in-memory libraries like FAISS or Chroma start consuming excessive RAM and become difficult to scale across clusters. Milvus thoroughly solves this problem with its distributed architecture. You can quickly spin up the entire environment using Docker Compose.
Create a docker-compose.yml file containing all 4 components: Milvus server, MinIO (stores index files), Etcd (stores metadata), and Attu (web GUI):
version: '3.5'
services:
etcd:
container_name: milvus-etcd
image: quay.io/coreos/etcd:v3.5.5
environment:
- ETCD_AUTO_COMPACTION_MODE=revision
- ETCD_AUTO_COMPACTION_RETENTION=1000
- ETCD_QUOTA_BACKEND_BYTES=4294967296
- ETCD_SNAPSHOT_COUNT=50000
volumes:
- ./volumes/etcd:/etcd
command: etcd -advertise-client-urls=http://127.0.0.1:2379 -listen-client-urls=http://0.0.0.0:2379 --data-dir=/etcd
minio:
container_name: milvus-minio
image: minio/minio:RELEASE.2023-03-20T20-16-18Z
environment:
MINIO_ACCESS_KEY: minioadmin
MINIO_SECRET_KEY: minioadmin
volumes:
- ./volumes/minio:/minio_data
command: minio server /minio_data
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:9000/minio/health/live"]
interval: 30s
timeout: 20s
retries: 3
standalone:
container_name: milvus-standalone
image: milvusdb/milvus:v2.3.4
command: ["milvus", "run", "standalone"]
environment:
ETCD_ENDPOINTS: etcd:2379
MINIO_ADDRESS: minio:9000
volumes:
- ./volumes/milvus:/var/lib/milvus
ports:
- "19530:19530"
- "9091:9091"
depends_on:
- "etcd"
- "minio"
attu:
container_name: milvus-attu
image: zilliz/attu:v2.3.4
ports:
- "8000:3000"
environment:
MILVUS_URL: milvus-standalone:19530
depends_on:
- "standalone"
Start the entire stack with a single command:
docker compose up -d
Wait about 30 seconds for the services to stabilize, then open your browser at http://localhost:8000. Enter the host milvus-standalone:19530 to log in to the Attu dashboard.
Milvus Architecture and the Role of Attu
1. How Milvus Works
Milvus completely decouples compute from storage. Thanks to this architecture, you can independently scale out Query Nodes when search traffic surges:
- Etcd: Stores collection metadata, schemas, partition details, and cluster state.
- MinIO / S3: Stores raw vectors, log files, and index files. Even if compute nodes crash and restart, data loss is prevented.
- Query Node & Data Node: Query Nodes load indexes into RAM to deliver millisecond-level search latencies. Data Nodes batch write (insert) data into immutable segments.
2. How Attu Helps with Operations
Attu is the official GUI developed by Zilliz. This interface eliminates the need to write manual scripts every time you inspect data:
- Visually create, edit, and delete collections and partitions.
- Check the in-memory loading status (Loaded/Unloaded) of each collection.
- Quickly test vector search queries and scalar metadata filtering directly from the web interface.
- Monitor entity counts and segment sizes in real time.
Data Operations with PyMilvus and Index Optimization
Two key factors determine RAG performance: designing a flexible schema with Dynamic Fields and choosing the right indexing algorithm.
1. Initializing a Collection for RAG
Below is a Python script to create a collection storing document chunks and 1536-dimensional vectors (standard for OpenAI text-embedding-3-small):
from pymilvus import connections, FieldSchema, CollectionSchema, DataType, Collection
# Connect to Milvus
connections.connect("default", host="localhost", port="19530")
collection_name = "rag_enterprise_docs"
# Define Schema
fields = [
FieldSchema(name="id", dtype=DataType.INT64, is_primary=True, auto_id=True),
FieldSchema(name="doc_id", dtype=DataType.VARCHAR, max_length=64),
FieldSchema(name="content", dtype=DataType.VARCHAR, max_length=4096),
FieldSchema(name="embedding", dtype=DataType.FLOAT_VECTOR, dim=1536)
]
# Enable enable_dynamic_field to flexibly store arbitrary JSON metadata
schema = CollectionSchema(fields, description="RAG Knowledge Store", enable_dynamic_field=True)
collection = Collection(name=collection_name, schema=schema)
print(f"Collection {collection_name} is ready!")
2. Choosing an Index: HNSW or IVF_FLAT?
Without an index, Milvus must perform a brute-force scan (Flat search). On 1 million vectors, queries will suffer latency in the hundreds of milliseconds. Practical options include:
- HNSW: Extremely fast search speed (typically < 5ms for 1 million vectors) with 98–99% recall. Trade-off: high RAM consumption for graph construction.
- IVF_FLAT: Clusters vectors around centroids. It consumes less RAM than HNSW, making it suitable for moderate server hardware.
Create an HNSW index with Cosine distance metric:
index_params = {
"metric_type": "COSINE",
"index_type": "HNSW",
"params": {"M": 16, "efConstruction": 200}
}
collection.create_index(field_name="embedding", index_params=index_params)
# Load collection into RAM after building the index
collection.load()
print("HNSW index built successfully; collection loaded into RAM.")
3. Hybrid Search: Combining Metadata Filtering with Vector Similarity
In practice, you often need to find relevant chunks that belong to a specific document or department. Use the expr expression for filtering:
query_vector = [0.015] * 1536
search_params = {"metric_type": "COSINE", "params": {"ef": 64}}
results = collection.search(
data=[query_vector],
anns_field="embedding",
param=search_params,
limit=3,
expr='doc_id == "finance_q1_2026"', # Filter metadata before computing similarity
output_fields=["doc_id", "content"]
)
for hits in results:
for hit in hits:
print(f"Score: {hit.distance:.4f} | Content: {hit.entity.get('content')}")
Production Best Practices for Milvus
- Always call
collection.load(): Milvus can only query segments that have been loaded into RAM. If queries return empty results, open Attu to check whether the collection status is Loaded. - Estimate RAM requirements for HNSW: Use the formula:
Number of vectors * Dimensions * 4 bytes * 1.5. For example, 5 million 1536-dimensional vectors require approximately5,000,000 * 1536 * 4 * 1.5 ≈ 46 GB RAM. Provision a server with at least 64 GB of RAM. - Limit the number of partitions: Keep fewer than 64 partitions per collection. Too many small partitions cause memory fragmentation on Query Nodes. For granular filtering, rely on dynamic metadata and
exprfilter expressions. - Enable periodic Etcd compaction: Failing to clean up snapshots will cause Etcd storage to exceed its 2GB/4GB quota and crash the cluster. Always configure
ETCD_AUTO_COMPACTION_RETENTION=1000. - Monitor via Prometheus: Milvus exposes a metrics endpoint on port
9091/metrics. Scrape this into Grafana to monitor QPS, query latency, and unindexed segment counts.

