Load Testing Systems with Locust in Python: Real-World Scenarios and API Metric Analysis

Python tutorial - IT technology blog
Python tutorial - IT technology blog

Comparing Current Load Testing Solutions: Apache JMeter, k6, and Locust

Before major sales events or new feature rollouts, how much concurrent traffic can your backend handle? Without load testing in advance, your system can easily crash right as traffic peaks. Choosing the right testing tool helps your team save dozens of debugging hours.

Currently, there are 3 popular solutions backend engineers frequently consider:

  • Apache JMeter: A veteran tool built in Java. Its strengths lie in a drag-and-drop GUI and an extensive plugin ecosystem, but XML-based configuration is bulky and difficult to track with Git.
  • k6 (Grafana): A modern choice written in Go. k6 uses JavaScript (ES6) for scripting, is extremely lightweight, and was built from the ground up for CI/CD pipeline integration.
  • Locust: An open-source framework written 100% in Python. You define user behavior entirely in standard Python code, complete with an intuitive Web UI for real-time metric monitoring.

Pros and Cons of Each Tool

Each tool offers distinct advantages depending on your team’s infrastructure and workflow:

1. Apache JMeter

  • Pros: Supports almost all network protocols (HTTP, JDBC, FTP, TCP, LDAP). Rich documentation and a large community.
  • Cons: The 1-thread-per-user model consumes significant RAM/CPU when simulating 2,000 to 5,000+ users. Clunky XML files make team collaboration and code merging difficult.

2. k6

  • Pros: Extremely fast load generation thanks to an optimized Go runtime. Readable JS scripts with seamless metric streaming to Prometheus and Grafana.
  • Cons: Does not run on a full Node.js runtime, so you cannot freely npm install arbitrary third-party packages. The open-source version lacks a built-in local Web UI.

3. Locust

  • Pros: Pure Python code. You can freely use requests, faker, redis, or any package in the Python ecosystem. Thanks to a coroutine architecture based on gevent, a 4-core dev machine can easily spawn 5,000 – 10,000 virtual users (VUs). Comes with a built-in Web UI displaying real-time charts without extra infrastructure setup.
  • Cons: Raw throughput on a single core is lower than k6. However, Locust neatly solves this via a Master-Worker distributed mode using just a single CLI flag.

Why Python Teams Should Choose Locust

If your primary tech stack is Python, Locust delivers the smoothest experience. You don’t need to learn a complex GUI or master a new domain-specific language.

Every complex scenario is handled directly in code: from logging in to obtain JWT tokens, querying databases for test data preparation, to generating mock data with Faker. for loops, if/else branches, exception handling — it’s all familiar Python.

Step-by-Step Guide to API Load Testing with Locust

Step 1: Install Locust

Create a dedicated virtual environment and install Locust via pip:

# Create and activate a virtual environment
python3 -m venv venv
source venv/bin/activate

# Install Locust
pip install locust

# Verify successful installation
locust --version

Step 2: Write the API Test Script (locustfile.py)

Create a locustfile.py file in your project directory. We will simulate an e-commerce flow: log in to fetch a JWT token, then browse the product list and view the personal user profile.

import json
from locust import HttpUser, task, between

class EcommerceUser(HttpUser):
    # Random think time between requests: 1 to 3 seconds
    wait_time = between(1, 3)
    token = None

    def on_start(self):
        """Runs once when each Virtual User starts (used for login)"""
        payload = {
            "username": "test_user",
            "password": "Secret@123"
        }
        headers = {"Content-Type": "application/json"}
        
        response = self.client.post("/api/v1/auth/login", json=payload, headers=headers)
        if response.status_code == 200:
            self.token = response.json().get("access_token")
        else:
            response.failure("Login failed, could not retrieve token!")

    @task(3)
    def get_products(self):
        """Task to browse products (weight 3: executed 3x more frequently than profile task)"""
        headers = {"Authorization": f"Bearer {self.token}"} if self.token else {}
        with self.client.get("/api/v1/products?page=1&limit=20", headers=headers, catch_response=True) as res:
            if res.status_code == 200:
                res.success()
            else:
                res.failure(f"Failed to fetch product list: {res.status_code}")

    @task(1)
    def get_user_profile(self):
        """Task to view user profile (weight 1)"""
        headers = {"Authorization": f"Bearer {self.token}"} if self.token else {}
        self.client.get("/api/v1/users/me", headers=headers, name="/api/v1/users/me")

Key highlights in the script above:

  • HttpUser: Class representing a virtual user, equipped with a built-in HTTP client managing sessions and cookies.
  • wait_time = between(1, 3): Simulates realistic user think-time, preventing virtual users from unrealistically spamming requests continuously.
  • on_start: Lifecycle hook that runs before tasks begin, ideal for login and access token retrieval flows.
  • @task(weight): Defines the traffic distribution ratio. A weight of @task(3) receives 75% of total requests compared to 25% for @task(1).
  • catch_response=True: Allows manually marking a request as failed if returned data violates business logic, even when the HTTP status code is 200 OK.

Step 3: Run the Test and Control via Web UI

Launch Locust with the following terminal command:

locust -f locustfile.py --host http://localhost:8000

Open http://localhost:8089 in your browser and configure the parameters:

  • Number of users: 200 (total virtual users to simulate).
  • Ramp-up (users started/second): 10 (spawns 10 new users per second until reaching 200).
  • Host: Target backend base URL for the load test.

Click Start swarming to generate load and observe real-time charts.

Step 4: Run Headless Tests in CI/CD

When integrating into GitHub Actions or GitLab CI pipelines, run in headless mode without a UI:

locust -f locustfile.py \
    --headless \
    --users 100 \
    --spawn-rate 10 \
    --run-time 3m \
    --host http://localhost:8000 \
    --html report.html

The command above runs a test with 100 users for 3 minutes and automatically exports an intuitive HTML report to report.html.

How to Analyze API Performance Metrics

While the test is running, focus on these 3 key metrics:

1. Requests Per Second (RPS / Throughput)

This metric measures the number of requests successfully processed by the backend per second. If the user count increases but RPS plateaus or drops, the backend has hit a bottleneck — often due to exhausted database connection pools or 100% server CPU utilization.

2. Response Time (Latency Percentiles: 50%, 95%, 99%)

Never rely solely on Average Response Time. This average is often skewed by lightweight 10ms requests, masking heavy requests that hang for 5 seconds.

  • 50th Percentile (Median): The experience of 50% of regular users.
  • 95th Percentile: 95% of requests complete faster than this threshold. This is the gold standard for evaluating SLAs (e.g., guaranteeing p95 < 250ms).
  • 99th Percentile: Measures worst-case scenarios, often occurring during database table locks or triggered Garbage Collection pauses.

3. Failure Rate (% Error)

The percentage of requests returning 5xx status codes or timing out. In standard load testing, an acceptable failure rate should be under 1%. When this figure spikes to 5-10%, you have found your system’s breaking point.

Pre-Load Test Optimization Checklist

  • Isolate the test runner from the server: Host Locust and the backend server on separate machines. Running them together causes Locust to compete for server CPU, skewing test results.
  • Test in a production-like environment: Run tests on Staging/UAT environments with RAM, CPU, and database volumes sized as close to production as possible. Avoid testing locally over localhost due to virtual I/O bottlenecks.
  • Monitor server resources: Keep monitoring tools (such as htop, Prometheus, Datadog) active to track CPU, RAM, disk IOPS, and database connection pools throughout the load generation phase.
Share: