TechCompare LogoTechCompare

Blacksmith runner benchmarks and hardware details

The Blacksmith sample includes two-vCPU Linux x64, Linux ARM64, and Windows runners, plus a six-vCPU macOS runner. Each CPU test uses one execution thread. Linux and Windows pin it to one logical CPU, while the macOS thread can migrate. The operating system and compiler are part of the result, so the four configurations aren't interchangeable measurements of a processor alone.

By TechCompare · Updated

Measured configurations
4
CPU execution threads
1
Five accepted samples per runner

Separate the native cache from GitHub cache

Where both backends verified all three save and restore pairs, the tool lets you inspect them separately. Transfers use a 512 MiB random payload and restore on a fresh runner job. This measures the archive service and action overhead under a controlled transfer, rather than a complete dependency installation. A real package cache can compress differently and may include many small files.

Read the runner size as a request

Blacksmith documents Linux and Windows sizes up to 32 vCPUs and macOS sizes up to 12. Its documentation also describes free capacity upgrades, so a workflow label is a requested tier rather than proof of the exact allocation during a test. The reports don't preserve that allocation. Prices are attached to the requested tier, with that limit visible in the hardware details.

Calculator

GitHub Actions runner CPU models, single thread CPU scores, disk, cache, and internet scores, and value per USD compute cost. Missing tests have no score.
RunnerCPUSingle threadCPU valueDiskCacheInternetCost/vCPU/minUSDMax vCPU
Blacksmith Linux x64

AMD EPYC (model hidden)

45,851
22.9M/USD
1,331.6
154.4
741.9
0.00232
Blacksmith macOS arm64

Apple M4 Pro

41,001
3.1M/USD
262.4
262.9
244.5
0.013312
Blacksmith Windows x64

AMD EPYC 4565P 16-Core

40,527
10.1M/USD
708.9
24.8
731.2
0.00432
Blacksmith Linux arm64

ARM64 (model undisclosed)

27,191
21.8M/USD
320.8
88.2
616.7
0.001332

Single thread CPU: CoreMark iterations/s. Disk and cache: composite MiB/s. Internet: composite Mbps. Value: score ÷ USD cost/vCPU/min. K means thousand, M means million. Higher is better. Cache uses each runner's best tested cache or artifact backend. Limits: maximum runner vCPU. Prices and value use USD. RunJob's EUR rate is converted at €1 = $1.1206 (ECB, 2026-10-09).

Open the full GitHub Actions runner benchmark tool

Measurement context

The Linux x64 guest reports AMD EPYC without a model number. Windows identifies EPYC 4565P, and macOS identifies M4 Pro. The Linux ARM guest reports only aarch64. The table preserves those differences instead of assigning the same specific processor to every Blacksmith runner.

Verdict

Read the CPU and cache results together if dependency reuse is a large part of your CI runtime. Keep a hidden CPU SKU marked as hidden. A similar score doesn't prove that two guests share the same host processor.

More runner benchmark scenarios

Frequently asked questions

Are these independent GitHub Actions runner benchmarks?
TechCompare publishes measurements from its own runner benchmark suite. The results come from the supplied published reports, while provider documentation supplies hardware mappings, prices, and available sizes. Scores from provider marketing pages aren't used as test results.
Does a CoreMark score predict my build time?
No. CoreMark is a CPU microbenchmark running one execution thread. A full build also depends on parallelism, memory, disk, dependencies, cache behavior, and queue time. This snapshot doesn't measure complete application builds or queue delays.
How are disk, cache, and internet scores calculated?
Each score is a geometric mean of that runner's own measurements. Disk combines sequential read/write MiB/s with random read/write IOPS converted to MiB/s using the 4 KiB operation size. Cache combines save and restore MiB/s from one backend. By default, each runner uses its highest verified score across tested cache and artifact backends. Hover or open details to see the backend, or select a specific backend type. Internet combines available download and upload Mbps, or uses the single recorded direction. Protocol 7 reports speeds from actual payload bytes and elapsed time, including transfers that reach the time limit. Older timeout records use recorded bytes divided by the nominal limit. These fixed formulas have no ranking reference or upper score limit, so adding a faster runner doesn't change existing scores. CPU keeps its measured CoreMark iterations per second.
How is price per minute per vCPU calculated?
A fixed runner's quoted minute rate is divided by its billed vCPU count. Prices are shown in USD and are list or overage rates before allowances and extras. RunJob publishes EUR consumed CPU rates, which are converted using the dated ECB reference rate shown on the page. Its original rate stays in the details, and memory charges are additional. GitHub uses private standard rates, and Namespace shows overage with prepaid equivalents.
What does CPU value mean?
CPU value is the measured single thread CoreMark score divided by the USD cost per vCPU per minute. More billed vCPUs don't multiply the measured score. EUR compute rates are converted before calculating value, so every provider is in one USD ranking. K means thousand and M means million. Optional disk, cache, and internet value columns use the same score divided by rate formula. These ratios don't measure full build performance or predict the final bill.
Why is an exact CPU model sometimes missing?
Some virtual machines expose only AMD EPYC or aarch64. The table uses a provider mapping where documentation identifies a family, labels that evidence, and leaves undisclosed model numbers unconfirmed. Similar benchmark scores aren't enough to identify a processor.
Are missing tests counted as zero?
No. Internet transfers with a nonzero recorded byte count can supply a speed even if they reach the time limit. A blocked request or a timeout with no recorded bytes stays unscored. Disk uses verified sample medians with all four throughput components. The updated BuildPulse ARM disk result is included using its three completed samples, with the original counts in the details. Cache requires all three verified save and restore pairs. The table shows a dash when no score is available. Hover or open the row details to inspect the measurements.