GitHub Actions performance monitoring

Find the workflows slowing down delivery

GitHub Actions performance monitoring reveals whether feedback is consistently slow, occasionally terrible, or failing before useful work completes. Pipetrics keeps workflow, job, step, and outcome evidence together for a focused investigation.

  • Compare average, median, and p95 duration
  • Find the jobs and steps using the most time
  • Review execution count, failures, and retries

Start with a question

Questions for GitHub Actions performance monitoring

Which workflow consumes the most runner time?

Which job or step contributes to slow p95 duration?

Is the change recurring or just one unusual run?

01

GitHub Actions performance monitoring needs median and p95

An average can hide the experience developers actually have. A small number of very slow runs can pull it upward, while a fast majority can conceal a painful tail. Median duration describes the typical run. The 95th percentile shows the boundary that all but the slowest five percent complete within.

Compare these statistics over a meaningful number of executions and within similar workflow conditions. A deployment workflow that runs weekly should not be judged like a pull-request workflow that runs hundreds of times. Pipetrics shows average, median, p95, execution count, outcomes, and total runtime together so the statistic remains connected to its sample.

  • Median describes the typical execution
  • P95 reveals slow-tail developer experience
  • Execution count provides confidence context
  • Total runtime reveals organization-wide impact
The duration trend plots average and p95; the job detail view also provides median and execution count. Click to zoom.

02

Do not confuse queueing with execution time

A workflow can feel slow before its first step begins. Queue delay can point toward runner availability, concurrency limits, or workload bursts. Execution duration points toward workflow configuration, dependencies, tests, builds, and external services. Those are different problems with different fixes.

The duration chart above shows how long work took to run; it is not a queue-delay chart. If a run feels slower than its execution metrics suggest, check its start times and runner availability separately before rewriting steps or adding capacity.

  • Keep waiting time distinct from execution duration
  • Check runner availability when a run starts late
  • Use job metrics to investigate time spent executing
  • Review concurrency policy before adding capacity

03

GitHub Actions performance monitoring at job and step level

Workflow-level monitoring tells you where to look. Job-level data identifies the branch of the execution graph responsible, while step-level data shows the command or action consuming time. Rank jobs by total runtime, typical duration, p95 duration, failure count, and retries rather than opening runs one by one.

Critical-path duration and total runtime answer different questions. Parallel jobs can add substantial billed runtime without extending wall-clock feedback by the same amount. A serial job on the critical path can delay every developer even if its total cost is modest. Preserve both perspectives when prioritizing performance work.

  • Rank jobs by total and per-run duration
  • Inspect steps repeated across many executions
  • Identify jobs on the feedback critical path
  • Keep parallel usage distinct from wall-clock delay
Job analysis connects total runner time to execution statistics and the steps with the greatest impact. Click to zoom.

04

Measure failure and retry overhead

A quick failure may be inexpensive but disruptive. A late failure can consume nearly the full workflow duration before providing no usable result. Track failed runtime alongside failure rate to find workflows where reliability improvements also recover meaningful developer and runner time.

Reruns are part of real performance. Counting only the final successful attempt understates how long the team waited and how much work the runner performed. Pipetrics records valid rerun attempts while rejecting invalid timestamp combinations, including reversed completion times from recreated skipped jobs.

  • Measure failed runtime, not only failure count
  • Include valid retries and rerun attempts
  • Find jobs that fail late in the workflow
  • Reject invalid or reversed timestamp intervals

05

Investigate regressions without claiming causation

When duration changes, identify the first observed commit associated with the difference and compare adjacent execution evidence. Then inspect the workflow and repository diff for dependency changes, new test groups, cache behavior, runner selection, or configuration changes that could explain the signal.

Association is not causation. Workload, external services, cache state, and runner availability can change between commits. Pipetrics commit impact narrows the investigation by showing observed before-and-after runtime and modeled cost. Your team or coding agent reviews the corresponding changes before reaching a conclusion.

  • Compare adjacent observed commits
  • Check whether the change persists
  • Inspect workflow and dependency diffs
  • State uncertainty when external factors remain

A repeatable path from signal to fix

  1. 01

    Rank

    Find workflows with high total runtime, slow median, painful p95, or repeated failed runtime.

  2. 02

    Separate

    Determine whether slow feedback comes from execution, failure, reruns, or time spent waiting to start.

  3. 03

    Drill down

    Inspect the responsible jobs, steps, runner labels, outcomes, and observed commits.

  4. 04

    Validate

    Make a focused change and compare enough later executions against the original baseline.

Common questions

What metrics should GitHub Actions performance monitoring include?

Use execution count, average, median, p95, total runtime, queue delay, outcomes, failed runtime, retries, and job- and step-level duration. No single metric explains both developer wait and resource usage.

Why is p95 useful for GitHub Actions?

P95 exposes slow-tail executions that a median or average may hide. It helps teams understand the poor but recurring experience without letting one extreme outlier dominate the report.

How do I find a slow GitHub Actions job?

Start with a slow or high-impact workflow, rank its jobs by median, p95, and total runtime, then open the responsible job and rank its steps. Check execution count and failures before prioritizing a change.

Can Pipetrics prove that a commit caused a regression?

No. Commit impact shows an association between observed commits and measured changes. It narrows the investigation, but a person or coding agent must inspect the diff and account for workload, cache, runner, and external-service changes.

Verify workflow syntax and billing behavior against the official GitHub Actions documentation and GitHub Actions billing documentation.