Context: The Selector Maintenance Nightmare
You’ve just finished writing a crawler to extract 10,000 e-commerce products using BeautifulSoup or Selenium. The next morning, the entire pipeline crashes. The reason is all too familiar: the site’s frontend team just deployed a redesign, renaming class="product-price" to a randomized hash like class="_2rK9a".
Auditing and updating broken XPath or CSS selectors manually consumes 30% to 50% of routine crawler maintenance time. This remains the inherent weakness of traditional web scraping methods.
ScrapeGraphAI solves this issue once and for all. Instead of relying on a rigid DOM tree, the library utilizes graph-based processing pipelines (Graph pipelines) combined with Large Language Models (LLMs). LLMs understand web page context much like a human eye does. You simply provide a URL along with a prompt describing the data you want to extract. ScrapeGraphAI analyzes the page automatically and returns clean, accurate JSON.
Environment Setup
ScrapeGraphAI requires Python 3.10 or higher. Running within a virtual environment is strongly recommended to avoid package conflicts:
# Initialize and activate virtualenv
python3 -m venv venv
source venv/bin/activate
# Install ScrapeGraphAI and Playwright for dynamic JS rendering
pip install scrapegraphai playwright
# Download browser binaries
playwright install
The library features out-of-the-box support for multiple LLM backends: OpenAI, Google Gemini, Groq, Azure OpenAI, or local models powered by Ollama.
Configuring Practical Pipeline Graphs
At the core of ScrapeGraphAI are its pre-built pipeline graphs. The two most common classes are SmartScraperGraph (single-page extraction) and SearchGraph (search engine integration).
1. Data Extraction with SmartScraperGraph and GPT-4o-mini
The following example extracts an article listing from a tech news site at an approximate cost of just $0.002 per request:
import json
import os
from scrapegraphai.graphs import SmartScraperGraph
graph_config = {
"llm": {
"api_key": os.getenv("OPENAI_API_KEY"),
"model": "openai/gpt-4o-mini",
"temperature": 0,
},
"headless": True,
"verbose": False,
}
prompt = """
Extract the list of articles on the page with the following fields:
- title: Article title (string)
- author: Author name (string, null if not available)
- points: Upvote score (integer)
- comments_count: Number of comments (integer)
"""
smart_scraper = SmartScraperGraph(
prompt=prompt,
source="https://news.ycombinator.com",
config=graph_config
)
result = smart_scraper.run()
print(json.dumps(result, indent=2, ensure_ascii=False))
2. Running Local LLMs with Ollama (Zero API Costs, Complete Privacy)
If you need to scrape proprietary internal data or operate at a scale of hundreds of thousands of pages per day, point to a local Ollama cluster:
from scrapegraphai.graphs import SmartScraperGraph
local_config = {
"llm": {
"model": "ollama/qwen2.5:7b",
"base_url": "http://localhost:11434",
"temperature": 0,
},
"embeddings": {
"model": "ollama/nomic-embed-text",
"base_url": "http://localhost:11434",
},
"headless": True
}
scraper = SmartScraperGraph(
prompt="Extract product name, original price, and discounted price",
source="https://example-shop.com/flash-sale",
config=local_config
)
data = scraper.run()
3. Token Optimization and Cost Reduction
Raw HTML on modern e-commerce sites often weighs between 2MB and 5MB (packed with thousands of SVG tags, inline CSS, and boilerplate scripts). Pushing this full HTML directly to an LLM drastically inflates token usage and increases latency.
ScrapeGraphAI automatically parses and compresses HTML before sending it to the LLM. Furthermore, you should set a strict max_tokens limit and tune chunk_size in the config when handling massive DOM trees.
Production Deployment & Monitoring
When packaging your scraper into a Celery worker or a Kubernetes CronJob, solid exception handling and token usage monitoring are essential.
1. Implementing Retry and Fallback Logic
Prevent intermittent network glitches or 429 rate limits from bringing down an entire batch job by wrapping the execution inside a retry function:
import logging
import time
from scrapegraphai.graphs import SmartScraperGraph
logging.basicConfig(level=logging.INFO, format="%(asctime)s - %(levelname)s - %(message)s")
logger = logging.getLogger(__name__)
def fetch_with_retry(url: str, prompt: str, config: dict, max_retries: int = 3, backoff_factor: int = 2):
for attempt in range(1, max_retries + 1):
try:
logger.info(f"Crawling: {url} (Attempt {attempt}/{max_retries})")
scraper = SmartScraperGraph(prompt=prompt, source=url, config=config)
output = scraper.run()
if output:
return output
except Exception as err:
logger.warning(f"Error on attempt {attempt}: {err}")
if attempt == max_retries:
logger.error(f"All {max_retries} attempts failed for: {url}")
raise
time.sleep(backoff_factor ** attempt)
return None
2. Monitoring Token Usage
The get_execution_info() method returns input/output token counts and latency metrics for each graph node. Ship these metrics to Prometheus or Grafana to monitor API budgets in real time:
execution_info = smart_scraper.get_execution_info()
logger.info(f"Total tokens used: {execution_info.get('total_tokens', 0)}")
logger.info(f"Execution time: {execution_info.get('execution_time', 0):.2f}s")
3. Three Core Production Best Practices
- Bypassing Anti-Bot Systems: When scraping sites protected by Cloudflare, configure a rotating Proxy Pool through Playwright settings instead of issuing raw HTTP requests.
- Structured Prompts: Always define JSON schema keys in English within your prompt (e.g.,
price,sku,stock_status) to ensure consistent output formatting and prevent backend serialization errors. - Hybrid Architecture: For the 90% of static pages whose layouts rarely change, leverage lightweight parsers (Scrapy, Selectolax) to maintain 100+ req/s throughput. Trigger ScrapeGraphAI only as a fallback layer when traditional selectors return empty payloads.

