LightRAG with Python Tutorial: Building Fast, Cost-Optimized Graph RAG

Artificial Intelligence tutorial - IT technology blog
Artificial Intelligence tutorial - IT technology blog

Quick start: Build Graph RAG with LightRAG in 5 Minutes

If you’ve ever experimented with Microsoft’s GraphRAG, you know the biggest hurdles are API costs and speed. Indexing a 100-page PDF can easily cost tens of dollars in tokens and take over an hour. LightRAG solves this bottleneck with a streamlined entity extraction mechanism and a dual-level retrieval architecture.

Install the library via pip:

pip install lightrag-hku openai

Create a quickstart.py file to index documents and test queries:

import os
from lightrag import LightRAG, QueryParam
from lightrag.llm import gpt_4o_mini_complete, openai_embedding

# Set up storage directory for the knowledge graph
WORKING_DIR = "./lightrag_storage"
os.makedirs(WORKING_DIR, exist_ok=True)

rag = LightRAG(
    working_dir=WORKING_DIR,
    llm_model_func=gpt_4o_mini_complete,
    embedding_func=openai_embedding,
)

# Sample data to index
sample_text = """
Nguyen Du was born in 1765 in Thang Long and is a great Vietnamese literary figure.
He is the author of The Tale of Kieu (Duan Truong Tan Thanh).
The Tale of Kieu was written in Nom script using the Luc Bat poetic form with 3,254 verses.
"""

# 1. Insert data (Insert & Graph Indexing)
rag.insert(sample_text)

# 2. Query data using hybrid mode
query = "What did Nguyen Du write, and what are its notable characteristics?"
response = rag.query(query, param=QueryParam(mode="hybrid"))
print(response)

Run the script:

export OPENAI_API_KEY="sk-proj-..."
python quickstart.py

Immediately upon completion, LightRAG automatically extracts entities such as Nguyen Du, The Tale of Kieu, and Nom script along with their relationships. The entire graph is stored locally in the ./lightrag_storage directory.

Why LightRAG Handles High-Level Overview Questions Better Than Naive RAG

Inherent Limitations of Plain Vector Search

Naive RAG breaks text into fixed chunks of roughly 500 to 1,000 tokens. This approach works well for specific pinpoint queries. However, when faced with a question like: “Summarize all architectural risks mentioned across 50 documentation files?”, Vector Search falls short. The answer is scattered across multiple files rather than contained within a single chunk.

Dual-Level Retrieval Mechanism in LightRAG

LightRAG combines a knowledge graph with two flexible retrieval levels:

  • Low-level retrieval: Traverses entity nodes and neighboring edges. Ideal for retrieving specific parameters, error codes, or precise facts.
  • High-level retrieval: Aggregates related topic clusters across the entire graph. Perfect for synthesizing broad overviews and analyzing overarching trends.

The library provides 4 query modes to cater to different needs:

  1. naive: Standard vector search over text chunks.
  2. local: Focuses on neighboring entities and localized context.
  3. global: Scans high-level relationships across the entire graph.
  4. hybrid: Combines both local and global to deliver answers that are both detailed and comprehensive.

Advanced: Running Local LLMs with Ollama and Custom Storage

To protect sensitive internal data and cut down on API costs, you can switch to open-source models using Ollama.

import asyncio
from lightrag import LightRAG, QueryParam
from lightrag.llm import ollama_model_complete, ollama_embedding
from lightrag.utils import EmbeddingFunc

WORKING_DIR = "./local_rag_storage"

async def main():
    rag = LightRAG(
        working_dir=WORKING_DIR,
        llm_model_func=ollama_model_complete,
        llm_model_name="qwen2.5:7b",
        llm_model_kwargs={"host": "http://localhost:11434"},
        embedding_func=EmbeddingFunc(
            embedding_dim=768,
            max_token_size=8192,
            func=lambda texts: ollama_embedding(
                texts, 
                embed_model="nomic-embed-text", 
                host="http://localhost:11434"
            )
        )
    )

    # Asynchronously index large documents
    with open("system_spec.txt", "r", encoding="utf-8") as f:
        await rag.ainsert(f.read())

    # Query in global mode
    res = await rag.aquery(
        "Summarize the network architecture and identify potential bottlenecks?",
        param=QueryParam(mode="global")
    )
    print(res)

if __name__ == "__main__":
    asyncio.run(main())

LightRAG natively supports both synchronous and asynchronous interfaces (ainsert, aquery), making integration with async backends like FastAPI or Sanic seamless.

Production Best Practices for LightRAG

When deploying LightRAG into real-world production search systems, keep the following tips in mind to optimize costs and stability:

1. Choose the Right Entity Extraction Model

The indexing phase consumes the most tokens because the LLM must analyze each text chunk to extract entities and relationships. Prioritize smaller models with solid structured JSON output capabilities, such as gpt-4o-mini or gemini-1.5-flash. This can cut API costs by 80–90% compared to frontier models while maintaining high graph quality.

2. Leverage Incremental Updates

LightRAG supports granular, incremental graph updates. You don’t need to rebuild the entire graph from scratch whenever new documents arrive:

# Simply ingest new text; LightRAG will merge it into the existing graph
rag.insert("Additional documentation regarding security policy v2.0...")

The system automatically links new entities into the existing network. User search queries continue running smoothly without interruption.

3. Route Query Modes by User Intent

Avoid overusing hybrid mode for every request, as it calls the LLM twice to aggregate context. Classify user intent before triggering queries:

  • Looking up definitions, error codes, or technical specs: Use local for fast responses and low token consumption.
  • Comparing architectures, evaluating risks, or analyzing trends: Use global.
  • Vague questions or complex multi-faceted inquiries: Use hybrid.

4. Back Up and Decouple Graph Storage

By default, LightRAG saves the graph structure to graph_chunk_entity_relation.graphml alongside key-value JSON files inside working_dir. When containerizing with Docker, mount this directory to an external Persistent Volume to prevent data loss across restarts or redeployments.

Share: