C/C++ Optimization (PGO / AutoFDO / LTO)¶
Advanced, opt-in techniques for squeezing more performance out of release C/C++ builds:
- Profile-Guided Optimization (PGO) — instrument, run a representative workload, rebuild with the collected profile.
- Sample-based PGO (AutoFDO) — PGO without instrumentation: sample a normal optimized binary (production traffic) and rebuild.
- Link-Time Optimization (LTO) — cross-module optimization (inlining, devirtualization, dead-code elimination) at link time.
PGO and LTO are independent axes and compose (e.g. --profile-use + --lto); AutoFDO is the no-instrumentation flavor of PGO. The relevant flags are listed in the command-line reference; this page is the how-and-why.
Profile-Guided Optimization (PGO)¶
Profile-Guided Optimization (PGO), also known as Feedback-Directed Optimization (FDO), is a technique that uses profile data collected from a program's real execution to guide the compiler's optimizations.
PGO is a global build mode (not a per-target attribute): you instrument the whole build, run a representative workload, then rebuild using the collected profile. It is wired for gcc, clang, and native MSVC; both phases use a dedicated build_*_pgo directory so they never clobber your normal build_* objects.
# Phase 1 — instrument, then run a representative workload to collect data
blade build //foo:server --profile-generate=/tmp/pgo
./build_release_pgo/foo/server # exercise the hot paths
# Phase 2 — rebuild optimized, using the profile
blade build //foo:server --profile-use=/tmp/pgo
Toolchain differences blade handles for you:
- gcc reads the
.gcdafiles directly and adds-fprofile-correction(tolerates the count skew of multithreaded programs). Both phases must share the build dir — gcc keys its.gcdalookup by object path — which the dedicatedbuild_*_pgodir guarantees. - clang needs a merged
.profdata: its instrumented run emits.profraw, and-fprofile-use=must point at a merged file, not a directory. Point--profile-useat the directory of.profraw(or an already-merged.profdata) and blade runsllvm-profdata mergefor you. clang's-fprofile-correctiondoes not exist, so it is not emitted. - MSVC uses whole-program (LTCG) instrumentation: blade compiles with
/GL, links the instrument phase with/LTCG /GENPROFILEand the optimize phase with/LTCG /USEPROFILE(and archives withlib /LTCG). The instrumented run writes<binary>!N.pgcnext to the<binary>.pgd, and/USEPROFILEauto-merges them on the optimize link — nopgomgrstep. The.pgdis keyed to the output name, so the sharedbuild_*_pgodir is what lets the optimize phase find the instrument phase's profile. Thepathargument is gcc/clang-specific and is not needed on MSVC (the profile lives next to the binary).
Blade owns the build flags and the clang merge step; producing a representative workload (and deciding when a profile is stale) is your job. Profiles are not portable across compilers — instrument, run, and optimize all on one toolchain.
The instrument build defines
BLADE_PGO_GENERATE(a Blade-private macro, not a compiler/industry standard), so source can flush the profile runtime in long-running or forking servers — e.g.#ifdef BLADE_PGO_GENERATE→__llvm_profile_write_file()(clang) /__gcov_dump()(gcc) /PgoAutoSweep(...)(MSVC). The use build and AutoFDO define nothing: those binaries should behave exactly like a normal release.
Sample-based PGO (AutoFDO)¶
The instrumentation-based PGO above needs two builds — instrument, then optimize — which is cumbersome, and the instrumented run carries real overhead.
AutoFDO is the no-instrumentation flavor of PGO: instead of an instrumented build, you sample a normal optimized binary and rebuild with the result. The collection runs at ~1% overhead (vs ~2× for instrumentation), so the profile can come from real production traffic. Works on gcc/clang (sample with perf) and native MSVC (sample with xperf — see SPGO below). Uses a dedicated build_*_autofdo dir.
# Phase 1 — build with AutoFDO-friendly debug info, then sample under perf
blade build //foo:server --autofdo-generate
perf record -b -- ./build_release_autofdo/foo/server # -b = LBR branch records
# Convert perf.data -> a sample profile yourself (it needs the collected binary):
# clang: llvm-profgen --perfdata=perf.data --binary=build_release_autofdo/foo/server --output=foo.prof
# gcc: create_gcov --binary=build_release_autofdo/foo/server --profile=perf.data --gcov=foo.afdo
# Phase 2 — rebuild using the converted profile
blade build //foo:server --autofdo-use=foo.prof
Typical usage — one build per release (steady state). Unlike instrumentation PGO (which always needs a separate slow instrumented build), AutoFDO's collection binary is a normal binary, so once you have a profile you combine both phases into a single build that is both optimized and collectable:
# Each release: optimize with last cycle's profile AND stay sample-able for the next.
blade build //foo:server --autofdo-generate --autofdo-use=foo.prof
# ship it -> sample it in production -> convert -> feeds the *next* release's --autofdo-use
--autofdo-generate + --autofdo-use compose on all three toolchains (gcc/clang add the debug info alongside -fprofile-sample-use/-fauto-profile; MSVC links /spgo alongside /spdin: — verified legal on MSVC 14.51, x64 + ARM64). So in steady state every build is simultaneously "optimize with last profile" and "collect for next" — no dedicated extra build. (The very first time, run --autofdo-generate alone to bootstrap a profile.) The AutoFDO debug flags are cheap, so you can also bake them into your release config and pass just --autofdo-use daily.
- clang →
-fprofile-sample-use=<profile>; the collection build adds-fdebug-info-for-profiling+-funique-internal-linkage-names(better sample-to-source mapping). - gcc →
-fauto-profile=<profile>; the collection build relies on debug line tables (the clang flags above are clang-only). - Line tables are guaranteed for the collection build. Sample-PGO maps samples back to source via debug line tables, so
--autofdo-generateneeds them. Ifdebug_info_levelemits none (e.g.no), Blade adds minimal ones for that build only —-gmlt(clang) /-g1(gcc) //Z7+/DEBUG(MSVC SPGO) — and prints a one-time notice. It never downgrades an existing-g(a later-gmltwould), so a normaldebug_info_level=midbuild keeps its full debug info. --autofdo-usetakes an already-converted profile, not a rawperf.data— the converter (llvm-profgen/create_gcov) needs the collected binary, which Blade doesn't have at build time. A rawperf.datais detected and rejected with the conversion command.- native MSVC uses its own sample-PGO — SPGO (Sample Profile Guided Optimization, VS 2022 / 2026 with MSVC 14.51+), which Blade drives from the same
--autofdo-*flags:--autofdo-generatelinks/spgo(collect build),--autofdo-use=app.spdlinks/LTCG /spdin:app.spd(both compile/GL). You sample withxperf(IP sampling on any CPU, LBR on Intel Haswell+/AMD Zen 4+/ARM64 ARMv9.2-A+) and convert withSPDConvertinto the.spd— those are your step, as with perf. clang-cl can't do sample-PGO on Windows (SPGO is cl-only; AutoFDO needsperf), so--autofdo-*is skipped there with a warning.
Platform availability (collection vs use). Sample-PGO collection needs hardware sampling: gcc/clang need perf + LBR (a bare-metal / PMU-passthrough x86_64 Linux host — ARM Linux VMs typically expose no PMU/LBR, so perf record -b fails there); native MSVC needs xperf (IP sampling works on any CPU, including ARM64). macOS has no sample-PGO collection path at all (no perf, no SPGO; Instruments/sample/dtrace don't feed llvm-profgen) — the --autofdo-use flag is still portable, so a Linux-collected profile can be applied there, but cross-OS/arch reuse is imperfect and unsupported. Where you can't collect, use instrumentation PGO (--profile-generate/--profile-use) — fully functional everywhere with no special hardware.
References: GCC AutoFDO tutorial · MSVC SPGO
Link-Time Optimization (LTO)¶
LTO performs cross-module optimization (inlining, devirtualization, dead-code elimination) at link time. Whether a project ships with LTO is a stable decision, so the primary control is the project intrinsic cc_config(lto='thin'); the --lto flags only override it per invocation.
blade build //foo:server -p release --lto # ThinLTO for this build
blade build //foo:server -p release --lto=full # monolithic LTO
blade build //foo:server -p release --lto=no # off (skip LTO's link cost while iterating)
- thin (default for
--lto) maps to clang-flto=thin/ gcc-flto=auto; clang additionally keeps a persistent ThinLTO cache under<build_dir>/.cache/thinlto/when the linker supports it (ld64 / lld / gold; GNUbfdis detected and the cache omitted). full is monolithic-flto. - Release only: debug builds never use LTO unless
--ltois given explicitly. No separate build dir (unlike PGO/coverage) — LTO ships and is stable, so it ridesbuild_release; toggling triggers a normal full rebuild. - Per-target opt-out:
lto = Falseon acc_librarykeeps it native (linked as an ordinary object alongside the bitcode) — for a TU that miscompiles under LTO, or a library that should stay native. - Toolchains: all four — gcc, clang, native MSVC
cl.exe, and clang-cl. Native cl maps to/GL+/LTCG(thin/full both →/LTCG, since MSVC LTCG is whole-program; subsumed when a PGO mode is active). clang-cl takes the LLVM path instead:-flto[=thin]→ bitcode, withlld-linkdoing the LTO (+ a/lldltocachefor thin).
Toolchain notes / robustness. clang and clang-cl thin is true ThinLTO (incremental, cached). gcc has no ThinLTO, so thin maps to gcc's parallel WHOPR (-flto=auto) — a different model with no persistent cache. Native MSVC's /GL+/LTCG keeps objects COFF (so cc_check_undefined reads them with dumpbin); clang-cl LTO emits bitcode, so the check routes to llvm-nm instead — both keep working. gcc's whole-program LTO is also less robust on large C++ binaries: gcc 15.x has been observed to ICE (internal compiler error: in odr_types_equivalent_p, at ipa-devirt.cc) when linking heavy protobuf/RPC binaries. An ICE is a compiler bug, not a code error — if you hit one, drop LTO for that binary with --lto=no (a per-TU lto=False won't help, since the failure is in the whole-program link). In short: clang ThinLTO is the well-trodden path; treat gcc full-program LTO as experimental. Runtime gains are also workload- and toolchain-dependent — concentrated on genuinely cross-module hot paths, often negligible elsewhere — so measure on your own hot paths before committing a project to it.
Self-registration / plugin patterns (a common gotcha). LTO's whole-program dead-code elimination can drop static-initializer registrations that are only reached by name at runtime — factory / plugin / dependency-injection registries, and self-registering objects in general (the [[gnu::constructor]] + register-into-a-map idiom). From LTO's whole-program view the registrar is unreferenced, so it — and its Register(...) side effect — is stripped; the runtime lookup then fails (… not found / a registry CHECK) even though the program builds and passes without LTO. (Note that "force-load the archive" tricks — link_all_symbols / -Wl,--whole-archive / -Wl,-force_load — do not save you here: they force the object to be linked, but LTO still internalizes and dead-strips the registrar. LTO does not honor force-load for symbol retention.) This is a property of the code pattern, not a blade or compiler defect, but it can fail a lot of tests at once if a project's DI mechanism uses it.
Make the registration survive LTO by keeping the registrar — mark it [[gnu::used, gnu::retain]] (used keeps it through LTO's DCE, retain through the linker's --gc-sections):
[[gnu::constructor, gnu::used, gnu::retain]] static void register_foo() { registry.Register("foo", …); }
When used + retain is not enough. Keeping the registrar only helps if the registry it writes into is a single instance. If the registry is reached through a vague-/COMDAT-linkage singleton — e.g. a template function-local static (template<class T> T& Get() { static T r; return r; }) — ThinLTO can internalize it per translation unit: the registrar runs, but registers into its TU's copy while the lookup reads a different copy → still not found. The fix then is to make the registry one strong-linkage global, defined in a single .cc (not synthesized per-TU via a template/inline local static).
If you can't (or won't) fix the source, turn LTO off for that binary — lto = False on the target, or build it with --lto=no. That is the reliable resolution when the registry can't be made LTO-safe.
Either way, run your test suite under LTO before shipping such a binary.