Getting Started with Letta (MemGPT) in Python: Build AI Agents with Self-Managing Long-Term Memory

Artificial Intelligence tutorial - IT technology blog
Artificial Intelligence tutorial - IT technology blog

2:00 AM and a VIP Bot Crash Caused by Hitting the Context Window Limit

At exactly 2:15 AM, the on-call Slack channel blew up with a stream of bright red alerts: HTTP 400 - ContextWindowExceededError. The cause? A VIP customer had been messaging continuously for three weeks, accumulating over 180 custom technical queries. The bot collapsed instantly because the prompt blew past gpt-4o’s 128k token limit.

This is a classic production challenge. LLMs are stateless by nature: each request is completely independent of the past. Conversely, users always expect the bot to recall every little detail across days or weeks of conversation.

Weighing 3 Memory Architectures for AI Agents

That very night, our team had to dissect three common approaches to find a way out:

  • 1. Sliding Window: Retain only the last k chat turns (typically 10-20 messages) and stuff them straight into the prompt. All older messages get truncated.
  • 2. Vector DB + RAG (Semantic Search History): Store every conversation snippet in a vector database (such as Qdrant or pgvector). When a new question comes in, the system embeds it and pulls the top-k most similar snippets into the context.
  • 3. OS-Style Tiered Memory (Letta / MemGPT): Mimics computer hardware memory hierarchies. Core Memory acts like RAM, while Archival and Recall Memory function like an SSD. The agent autonomously decides whether to read, write, delete, or update memory via internal tool calls.

Battle-Tested Comparison: Flaws of RAG vs. Strengths of Letta

1. Sliding Window: Cheap and Easy to Build, but Prone to Amnesia

This approach is lightning-fast to implement—all you need is a List or Queue in Redis. The downside is that after 15-20 exchanges, crucial initial context like project budget, tech stack, or shipping address vanishes from the context window. The bot starts asking naive, repetitive questions, frustrating users.

2. Vector DB RAG: Finds the Right Keywords, but Blind to Timelines

RAG can store massive volumes of data. However, it handles state mutation terribly. Take a real-world scenario:

  • Last week, the user said: “I’m running PostgreSQL 14 on AWS RDS.”
  • Today, the user announces: “Our system just migrated over to CockroachDB!”

Standard RAG vector search will retrieve both chunks with nearly identical similarity scores. The LLM receives a prompt containing conflicting facts and starts hallucinating, completely confused about what the current reality is.

3. Letta: A Self-Managing Three-Tier Memory Architecture

Letta (originally born out of UC Berkeley’s MemGPT research project) tackles this problem at its root by partitioning memory into three distinct tiers:

  • Core Memory (RAM): Comprises the human block (user profile) and the persona block (bot persona). This block stays directly inside the context window on every request and is capped in size (usually a few thousand tokens) to optimize costs.
  • Recall Memory: A chronological log of all past messages, with built-in filtering and pagination support.
  • Archival Memory (Disk Storage): An unlimited long-term repository powered by semantic search via embeddings.

The core distinction: Agents in Letta possess Memory Editing Tools. When a user reports a stack change, the agent immediately invokes core_memory_replace to overwrite the outdated data in Core Memory. Developers don’t have to write extra cron jobs or SQL update statements.

Hands-On Implementation: Building a Letta Agent with Python

Follow the steps below to build an agent with self-updating memory from local development to production.

Step 1: Install the Letta SDK and Set Up Your Environment

Create a clean virtual environment and install the letta package:

python3 -m venv letta-env
source letta-env/bin/activate

pip install letta

Step 2: Configure API Keys and Launch the Letta Server

Letta operates on a client-server model. The Letta Server manages state, connects to the database (defaults to local SQLite or PostgreSQL in production), and orchestrates background tasks:

export OPENAI_API_KEY="sk-proj-your-api-key-here"

# Launch the local background server (default port 8283)
letta server

Step 3: Initialize the Agent and Seed Initial Core Memory

Create an agent_demo.py file. Connect to the server and configure initial state for your agent:

from letta import create_client

# Connect to the Letta server running on localhost:8283
client = create_client()

# Create agent with two Core Memory blocks: persona and human
agent_state = client.create_agent(
    name="devops_support_agent",
    memory={
        "persona": "I am an SRE infrastructure support engineer. My communication style is concise, prioritizing actionable commands and root-cause analysis from logs.",
        "human": "Customer's name is Nam, DevOps Lead, managing a 40-node production Kubernetes cluster on GCP."
    }
)

print(f"Agent initialized successfully! ID: {agent_state.id}")

Step 4: Verify Autonomous Memory Overwrites

Now, send two consecutive messages to verify whether the agent detects changes and autonomously updates its internal RAM:

from letta import create_client

client = create_client()
# Replace with the ID obtained from Step 3
agent_id = "agent-xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"

# Message 1: Supply updated infrastructure details
response = client.send_message(
    agent_id=agent_id,
    role="user",
    message="My team just completely decommissioned Kubernetes; all services have migrated over to HashiCorp Nomad!"
)

for msg in response.messages:
    # Print the LLM's internal reasoning when it decides to call a memory tool
    if hasattr(msg, 'message_type') and msg.message_type == 'reasoning_message':
        print(f"[LLM Reasoning]: {msg.reasoning}")
    elif hasattr(msg, 'text') and msg.text:
        print(f"[Agent]: {msg.text}")

# Message 2: Put the bot's memory to the test
response2 = client.send_message(
    agent_id=agent_id,
    role="user",
    message="What does my current core infrastructure look like?"
)

for msg in response2.messages:
    if hasattr(msg, 'text') and msg.text:
        print(f"[Agent]: {msg.text}")

Step 5: Inspect Core Memory Directly to Verify

Beyond checking the text responses, query the API directly to inspect the actual values stored in the database:

current_memory = client.get_core_memory(agent_id=agent_id)
print("\n=== CORE MEMORY DATA IN DB ===")
for block in current_memory.blocks:
    print(f"[{block.label}]: {block.value}")

Looking at the console logs, you will see that the human block has been updated with HashiCorp Nomad. The entire operation happened seamlessly via an under-the-hood core_memory_replace tool call. The context window for each request sent to OpenAI remains steady around 2,000–3,000 tokens—completely eradicating context overflow errors while slashing API token costs by 70–80% compared to appending the full chat history.

Share: