A Guide to git merge-base: Finding the Common Ancestor Commit for Accurate Diffing in CI/CD

Git tutorial - IT technology blog
Git tutorial - IT technology blog

The Nightmare of “Redundant Test Runs” in CI/CD

A typical monorepo or medium-sized backend project often contains 3,000 to 5,000 source files. Running linters and tests across the entire repository on every Pull Request (PR) can easily eat up 15 to 20 minutes. This bottleneck clogs the CI queue and tests the entire team’s patience. The most straightforward solution: scan only the files that were newly added or modified in the PR.

The real challenge lies in the comparison step. How does the runner know exactly which files you modified compared to the main branch? Pick the wrong reference commit, and your pipeline will miss critical bugs. Even worse, you could be unfairly blamed when linters fail on files a colleague merged just five minutes earlier.

Three Ways to List Changed Files: Which Approach Is Best?

Developers usually start with one of the following three approaches when writing file-filtering scripts for CI.

Approach 1: Compare with the Immediately Preceding Commit (HEAD~1)

The most intuitive way is to inspect only the latest commit pushed to the branch:

git diff --name-only HEAD~1 HEAD

Approach 2: Direct Branch Tip Comparison (Double-dot syntax)

Many developers reach for the double-dot syntax to compare the feature branch directly against main:

git diff --name-only origin/main..HEAD

Approach 3: Find the Common Ancestor Using git merge-base

The git merge-base command determines the best common ancestor commit between two branches. Using this historical reference point, we can diff against the tip of the current branch:

BASE_COMMIT=$(git merge-base origin/main HEAD)
git diff --name-only $BASE_COMMIT HEAD

Git provides a convenient shorthand for this logic via the triple-dot syntax:

git diff --name-only origin/main...HEAD

Weighing the Pros and Cons of Each Approach

Approach 1 (HEAD~1): Fast, but Fundamentally Flawed

  • Pros: Blazing fast. The runner does not need to fetch the history of any other branch.
  • Cons: Breaks down as soon as a feature branch has two or more commits. Pushed four commits to your PR? CI will only scan the files in the fourth commit, completely ignoring the previous three.

Approach 2 (origin/main..HEAD): A Common Trap

  • Pros: Short command, easy to type.
  • Cons: Prone to cascading issues. Suppose you branch off main on Monday morning. By the afternoon, teammates have merged eight other PRs into main. The origin/main..HEAD syntax factors in all changes on main that your branch doesn’t have yet. As a result, CI starts flagging formatting and syntax errors in dozens of files written by other people.

Approach 3 (git merge-base): 100% Accuracy

  • Pros: Filters exactly the changes you introduced since branching off. It remains completely unaffected by any new commits landed on main in the meantime. This is precisely the mechanism GitHub and GitLab use to render the “Files changed” tab.
  • Cons: The runner requires sufficient commit history. Default shallow clones may throw errors if not properly configured.

Real-World Impact of Adopting git merge-base

After implementing this mechanism across a team of 8 engineers, our test pipeline wait time plummeted from 18 minutes to under 3 minutes per commit. This led to substantial cost savings on AWS and GitHub Actions runners.

Developers complaining “I never touched this file, why is CI asking me to fix its formatting?” became a thing of the past. To ensure your pipeline accurately reflects your scope of work, git merge-base is the only reliable choice.

Step-by-Step Implementation Guide for CI Runners

Step 1: Understanding Ancestor Resolution via the Commit DAG

Consider the commit tree illustration below:

      B---C (feature - HEAD)
     /
A---M1---M2 (origin/main)

Where:

  • A is the base commit where you branched off to create the feature.
  • M1, M2 are new commits merged into main by teammates.
  • B, C are commits you just authored.

When running the following command:

git merge-base origin/main HEAD

Git traverses backwards through the DAG and returns the commit hash for A. Diffing between A and C (HEAD) will solely return the files modified in B and C.

Step 2: Resolving Shallow Clone Issues on CI Runners

CI engines such as GitHub Actions and GitLab CI default to shallow clones with depth: 1 to conserve network bandwidth. In this state, the runner lacks commit A, causing git merge-base to fail with fatal: Not a valid object name.

Fix this by fetching additional history for the target branch before computing the diff:

# Method 1: Fetch additional commits from the main branch
git fetch origin main --depth=100

# Method 2: Convert to full history if the repo size is manageable
git fetch --unshallow || true

Step 3: Writing an Automated Changed-Files Filter Script

Create a shell script named get_changed_files.sh with the following content:

#!/usr/bin/env bash
set -euo pipefail

TARGET_BRANCH="origin/main"

# Determine the best common ancestor commit
BASE_COMMIT=$(git merge-base "$TARGET_BRANCH" HEAD)
echo "Resolved base commit: $BASE_COMMIT"

# Filter modified or added files, excluding deleted files (--diff-filter=d)
CHANGED_FILES=$(git diff --name-only --diff-filter=d "$BASE_COMMIT" HEAD)

if [ -z "$CHANGED_FILES" ]; then
  echo "No changed files detected."
  exit 0
fi

echo "List of files to verify:"
echo "$CHANGED_FILES"

# Example: Only run the flake8 linter on changed Python files
PYTHON_FILES=$(echo "$CHANGED_FILES" | grep -E '\.py$' || true)
if [ -n "$PYTHON_FILES" ]; then
  echo "Starting linting for Python files..."
  echo "$PYTHON_FILES" | xargs flake8
fi

Step 4: Embedding into a Monorepo CI Pipeline

Sample configuration to selectively run tests per affected microservice:

# 1. Sync the reference for the main branch
git fetch origin main:refs/remotes/origin/main

# 2. Retrieve the common ancestor hash
MERGE_BASE=$(git merge-base origin/main HEAD)

# 3. Scan top-level module folders containing changed files
CHANGED_DIRS=$(git diff --name-only "$MERGE_BASE" HEAD | cut -d/ -f1 | sort -u)
for dir in $CHANGED_DIRS; do
  if [ -f "$dir/package.json" ]; then
    echo "Triggering tests for service: $dir"
    (cd "$dir" && npm test)
  fi
done

Mastering git merge-base gives you full control over your CI/CD pipelines, optimizes compute resources, and eliminates reliance on fragile third-party plugins.

Share: