Performance profiling¶
Performance work begins with a reproducible scenario and a user-visible metric. Do not optimize from intuition, one debug build, or a profile captured from a different commit/configuration.
Choose the evidence¶
- Use Catch2 microbenchmarks for repeatable, isolated algorithm comparisons.
- Use stable wall/turn/frame measurements for an end-to-end scenario.
- Use the
CATA_PROFILE_*wrappers insrc/profiling.hfor Tracy scopes, frames, text, and plots. - Use a platform profiler for CPU, allocation, I/O, GPU, or Android-specific questions.
- Use diagnostic timings only to explain a live failure, not as a substitute for a benchmark.
Tracy build¶
With an installed Tracy client library, the repository contract is:
When TRACY is off, the wrappers compile to no-ops. Game code must not call Tracy macros
directly; this preserves disabled builds and profiler choice.
Reproducible comparison¶
Record commit, compiler, optimization/LTO, sanitizer, frontend, SDL version, hardware, power mode, world/save/mods, RNG seed, scenario, warmup, sample count, statistic, and raw results. Compare before/after under the same conditions and inspect correctness tests before accepting a speedup.
Hot-path rules¶
Name the owner and invalidation boundary before caching. Check complexity, allocations, string/ translation work, registry lookups, map/inventory scans, renderer stalls, and cross-language calls. Do not trade deterministic behavior, save compatibility, or bounded Lua handles for speed.
Generated artifacts¶
Profiler captures, compiler time traces, flame graphs, compile_commands.json, Doxygen, ctags,
clangd indexes, and large symbol databases are generated. Upload them as scoped CI/review
artifacts when useful; do not commit them.
Acceptance¶
State baseline and new numbers with uncertainty, correctness checks, platforms tested, and any memory/startup/build-size regression. If results are inconclusive, report that rather than claiming an optimization.