1. Quick start: Deploy Apache Doris and query data in 5 minutes
If you want to quickly test scanning tens of millions of rows in under a second, Docker all-in-one is the fastest way to get started.
Create a docker-compose.yml file with the following configuration:
version: '3.8'
services:
doris:
image: apache/doris:doris-all-in-one-2.1.0
container_name: apache-doris-standalone
ports:
- "8030:8030" # FE HTTP Server
- "9030:9030" # FE MySQL Server Port
- "8040:8040" # BE HTTP Server
environment:
- FE_SERVERS=fe1:127.0.0.1:9010
- BE_SERVERS=be1:127.0.0.1:9050
volumes:
- doris_fe:/opt/apache-doris/fe/doris-meta
- doris_be:/opt/apache-doris/be/storage
volumes:
doris_fe:
doris_be:
Run the container in the background:
docker compose up -d
Doris communicates via standard MySQL protocol. Simply use your familiar MySQL Client to connect (default user root, no password):
mysql -h 127.0.0.1 -P 9030 -u root
Initialize the database and sample table for web access analytics:
CREATE DATABASE analytics_db;
USE analytics_db;
CREATE TABLE site_access_log (
event_date DATE NOT NULL,
site_id INT NOT NULL,
user_id VARCHAR(64) NOT NULL,
page_url VARCHAR(255),
pv INT SUM DEFAULT "1"
)
AGGREGATE KEY(event_date, site_id, user_id, page_url)
DISTRIBUTED BY HASH(site_id) BUCKETS 10
PROPERTIES("replication_num" = "1");
INSERT INTO site_access_log VALUES
('2026-03-01', 101, 'usr_9921', '/home', 1),
('2026-03-01', 101, 'usr_9921', '/home', 2),
('2026-03-01', 102, 'usr_1024', '/checkout', 1);
SELECT site_id, SUM(pv) as total_views FROM site_access_log GROUP BY site_id;
2. In-depth: RDBMS bottlenecks and Doris’s MPP architecture
When traditional RDBMS struggles with analytics
Imagine your orders or tracking_logs table in MySQL or PostgreSQL ballooning from 5 million to 80 million rows. At this point, a single SUM with GROUP BY or a COUNT(DISTINCT user_id) query is enough to peg your I/O. Server CPU hits 100%, while your Metabase or Grafana dashboards spin indefinitely, waiting 30 to 40 seconds for a response.
The root cause lies in row-oriented storage. Even if you only need to calculate the total amount across two columns, order_date and amount, the database is still forced to pull entire rows—including heavy columns like shipping_address or JSON payloads—from disk into RAM.
How does Apache Doris solve this?
Doris is a modern Data Warehouse built on a Massively Parallel Processing (MPP) architecture. It combines columnar storage with vectorized execution (SIMD) to drive query latencies down to sub-second levels.
- Columnar Storage & High Compression: Doris reads only the columns involved in a query. Leveraging LZ4/ZSTD algorithms, real-world compression ratios reach 5:1 to 8:1, saving up to 75% of disk storage.
- Two-Tier Independent Architecture: A Doris cluster consists only of Frontend (FE – manages metadata, parses SQL, distributes plans) and Backend (BE – stores tablets, handles scans, and performs distributed computation). The system runs autonomously, eliminating the operational overhead of ZooKeeper or Hadoop HDFS.
- Full MySQL Wire Protocol Compatibility: BI teams can directly connect Superset, Metabase, Tableau, or PowerBI to Doris without installing third-party drivers.
3. Advanced: Data model design and real-time Stream Load
Choosing the right data model
Doris provides three core table models tailored to different use cases:
- Aggregate Model: Automatically aggregates data based on predefined keys during ingestion (the top choice for dashboard metrics and PV/UV analytics).
- Unique Key Model: Supports high-speed upserts and deletes (ideal for real-time CDC synchronization from MySQL/PostgreSQL).
- Duplicate Model: Retains raw log records as-is, supporting sort keys for blazing-fast range scans.
Here is an example of creating a Unique Key table with Merge-on-Write enabled, optimized for high-frequency updates:
CREATE TABLE ecom_orders (
order_id BIGINT NOT NULL,
user_id INT NOT NULL,
order_status VARCHAR(32),
total_amount DECIMAL(12, 2),
updated_at DATETIME
)
UNIQUE KEY(order_id)
DISTRIBUTED BY HASH(order_id) BUCKETS 16
PROPERTIES(
"enable_unique_key_merge_on_write" = "true",
"replication_num" = "1"
);
Ingesting data via HTTP Stream Load with Python
Doris features a built-in HTTP Stream Load mechanism, allowing micro-batches of JSON/CSV data to be loaded directly into target tables at throughputs of 30,000 to 50,000 records per second per single BE node.
Python script for streaming data in real-time via Stream Load:
import requests
import json
doris_host = "http://127.0.0.1:8030"
db = "analytics_db"
table = "ecom_orders"
url = f"{doris_host}/api/{db}/{table}/_stream_load"
headers = {
"format": "json",
"strip_outer_array": "true",
"Expect": "100-continue"
}
data = [
{"order_id": 10001, "user_id": 55, "order_status": "PAID", "total_amount": 1450000.00, "updated_at": "2026-03-01 10:20:00"},
{"order_id": 10002, "user_id": 89, "order_status": "SHIPPED", "total_amount": 320000.00, "updated_at": "2026-03-01 10:21:15"}
]
response = requests.put(
url,
data=json.dumps(data),
headers=headers,
auth=("root", "")
)
print(response.json())
4. Battle-tested best practices for production
When our team’s tracking system reached 250 million records per month, migrating the entire reporting workload to Apache Doris slashed CPU load on our primary MySQL cluster by 75%. Here are four practical production takeaways to keep in mind:
- Size Buckets Properly: The ideal size for a tablet (compressed bucket) is between 100MB and 1GB. Avoid over-partitioning (such as allocating 100 buckets for a table with only 2 million rows), as this causes metadata fragmentation and increases FE overhead.
- Leverage Rollup Indexes: Detail tables store logs at second-level granularity, but reports mostly aggregate by day. Creating background Rollup Indexes can accelerate trend chart queries by 5x to 10x without modifying existing SQL queries.
- Choose the Right Backend (BE) Hardware: Invest in NVMe SSDs. Because Doris heavily leverages SIMD instruction sets (AVX2/AVX-512) for vectorized computation, CPUs with higher single-core clock speeds (3.2GHz+) significantly outperform high-core-count, low-frequency alternatives.
- Configure FE JVM Heap Memory: Doris keeps all metadata entirely in FE memory. When your dataset scales to hundreds of millions of partitions, allocate at least 8GB to 16GB of JVM Heap (the
JAVA_OPTS="-Xmx8192m"parameter inconf/fe.conf) to prevent OutOfMemory errors.

