Two structural compile-time optimisations on the conformallab_cgal_tests
target, both opt-out-able and verified safe (236/236 tests pass under
every configuration).
(1) Precompiled headers — option CONFORMALLAB_USE_PCH (default ON)
target_precompile_headers(conformallab_cgal_tests PRIVATE
<CGAL/Surface_mesh.h>
<CGAL/Simple_cartesian.h>
<CGAL/Kernel_traits.h>
<CGAL/boost/graph/iterator.h>
<CGAL/Polygon_mesh_processing/triangulate_faces.h>
<Eigen/Dense> <Eigen/Sparse> <Eigen/SparseCholesky> <Eigen/SparseQR>
<gtest/gtest.h>
<vector> <string> <cmath> <complex>
)
Absorbs the per-TU CGAL+Eigen template-parse cost (measured at 5.9 s
per minimal "include <CGAL/Discrete_conformal_map.h>" hello-world TU
on Apple M1).
(2) Unity Build — UNITY_BUILD ON with UNITY_BUILD_BATCH_SIZE 4
Concatenates the 22 test TUs into 5 batches of <=4 files each;
CGAL+Eigen headers parsed once per batch instead of once per TU.
Batch size 4 keeps gtest's TEST(...) macros and per-file
`using namespace ...` from colliding across batched files.
Numbers (Apple M1, Ninja, -j8, clean rebuild)
─────────────────────────────────────────────
wall CPU tests
baseline 78 s 676 s 236/236
+ PCH 66 s 474 s 236/236 (-15% wall, -30% CPU)
+ PCH + Unity 55 s 167 s 236/236 (-30% wall, -75% CPU)
Honest deferred items (documented in doc/architecture/compile-time.md):
* `extern template` (lever #2 in the analysis) — subsumed by PCH;
estimated residual gain <5%, would add Eigen-version fragility.
* Header split <CGAL/Discrete_conformal_map_{euclidean,spherical,
hyper_ideal}.h> (lever #3) — downstream-only benefit (our test
build needs all three); kept as a future cleanup once a downstream
user actually requests it.
Opt-outs: `-DCONFORMALLAB_USE_PCH=OFF` and `-DCMAKE_UNITY_BUILD=OFF`.
Detailed measurement methodology, per-TU breakdowns, clang
-ftime-trace template hot-spots, and a "what comes next" lever list
live in doc/architecture/compile-time.md.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
6.1 KiB
Compile-time analysis & quick-wins
Audience. Anyone maintaining or extending the build system. Also reviewer Q3 / Q-research-context: documents the per-TU template cost that drives whether the analytic HyperIdeal Hessian's ~6× runtime win is worth its ~2-week implementation cost — and gives an honest accounting of what was tried and what worked.
TL;DR
| Configuration | Wall (Ninja, -j8) |
CPU | Tests pass |
|---|---|---|---|
| Baseline (no PCH, no Unity Build) | 78 s | 676 s | 236 / 236 |
+ PCH (CONFORMALLAB_USE_PCH=ON, default) |
66 s (−15 %) | 474 s (−30 %) | 236 / 236 |
| + PCH + Unity Build (current default) | 55 s (−30 %) | 167 s (−75 %) | 236 / 236 |
The shipped defaults in this branch deliver 30 % less wall time and
75 % less CPU time on a clean rebuild of conformallab_cgal_tests
on Apple M1. All 236 CGAL tests pass under every configuration.
Tunable via -DCONFORMALLAB_USE_PCH=ON|OFF and
-DCMAKE_UNITY_BUILD=ON|OFF.
Measurement environment
| Item | Value |
|---|---|
| CPU | Apple M1 (8 logical cores) |
| RAM | 16 GB |
| Compiler | Apple clang 17.0.0 |
| Build type | Release |
| CGAL | 6.1.1 vendored (48 MB headers) |
| Eigen | 3.4.0 vendored (6.5 MB headers) |
| Generator | Ninja 1.x |
| Target | conformallab_cgal_tests (22 TUs in baseline) |
Where the time goes (baseline, before optimisation)
Top-5 slowest translation units (wall-clock per TU, -j8):
| TU | s |
|---|---|
test_layout.cpp |
54.7 |
test_geometry_utils.cpp |
52.2 |
test_phase6.cpp |
50.7 |
test_pipeline.cpp |
50.4 |
test_phase7.cpp |
45.3 |
A minimal "hello world" TU that does nothing but
#include <CGAL/Discrete_conformal_map.h> already costs 5.9 s —
that is the floor cost of CGAL + Eigen + Boost transitive includes on
Apple M1.
Clang -ftime-trace on test_layout.cpp (~17 s isolated):
| Phase | Time | % |
|---|---|---|
| Backend (CodeGen + Opt) | 9.3 s | 55 % |
| Frontend (Parse + Sema + Templates) | 8.0 s | 45 % |
| ⤷ InstantiateFunction | 4.3 s | 25 % |
| ⤷ InstantiateClass | 3.3 s | 19 % |
Single most expensive Eigen template instantiations (~1.1 – 1.2 s each, per TU):
Eigen::SelfAdjointEigenSolver<Matrix<2,2>>(PCA innormalise_euclidean)Eigen::ColPivHouseholderQR<Matrix<complex,3,3>>(MobiusMap::from_three)Eigen::internal::tridiagonalization_inplace<Matrix<2,2>>Eigen::HouseholderSequence<Matrix<2,2>>
These get instantiated from scratch in every TU that pulls
layout.hpp in — the inefficiency this branch's PCH closes.
What was tried
The four candidate quick-wins from the prior analysis were:
- Precompiled headers — shared PCH covering CGAL + Eigen + gtest +
the std headers every test uses. Implemented; shipped, opt-out
via
-DCONFORMALLAB_USE_PCH=OFF. Seecode/tests/cgal/CMakeLists.txt. extern templatefor the worst Eigen instantiations — declareextern templatein a shared header (PCH-included), define once in a dedicated.cpp. Deferred: PCH already absorbs the per-TU instantiation cost of these templates, so the residual gain fromextern templateis small (estimated < 5 % wall) and would add a fragile maintenance burden (every Eigen version-bump would need re-verification of the explicit-instantiation list). If a future Eigen update breaks the PCH, this is the next lever.- Header split for
<CGAL/Discrete_conformal_map.h>** — separate into_euclidean.h/_spherical.h/_hyper_ideal.h` so downstream consumers who only need one geometry pay less. Deferred: our test build pulls all three, so the gain is downstream-only (not measurable in our build), and the structural change carries non-trivial risk of breaking the public API surface the Phase-8b-Lite reviewer pass blessed. Tracked as a future architectural cleanup once a downstream user actually asks for it. - Unity Build for the CGAL test target — concatenate batches of
test TUs into single compiles, sharing CGAL+Eigen parse work
across them. Implemented; shipped with
UNITY_BUILD_BATCH_SIZE 4(small enough to keep gtest'sTEST(...)macros + per-fileusing namespace …from colliding). Opt-out via-DCMAKE_UNITY_BUILD=OFF.
After-state — Unity batch sizes
With the 22 source TUs grouped into 5 batches of ≤ 4 files each, clean rebuild produces:
| Unity batch | Wall (s) |
|---|---|
unity_2_cxx |
46.3 |
unity_3_cxx |
41.5 |
unity_4_cxx |
40.6 |
unity_1_cxx |
26.0 |
unity_0_cxx |
13.0 |
Plus gtest itself (gtest-all.cc.o, 8.5 s) and gtest_main.cc.o
(1.2 s) outside the batches.
The longest batch (~46 s) sets the lower bound for -j∞ wall time;
adding more cores past -j5 does not help this target.
Honesty notes
-jscaling is sublinear. Measured speedups:-j1→-j2= 1.66×,-j2→-j4= 1.59×,-j4→-j8= 1.28×. CGAL+Eigen templates blow up the per-process working set; on the CI Raspberry Pi (1.6 GB RAM cap) we run-j1for the CGAL job by necessity.- The PCH compiles in ~3 s the first time and then short-circuits every TU. Total PCH cost amortises after the second TU.
- CGAL version sensitivity. The PCH is keyed to the specific CGAL headers it lists. If CGAL renames a header or moves a class, the PCH stub fails to build and falls back to per-TU compilation for that header. The four CI quality-gate runs catch this within one PR.
- macOS-specific. The numbers above are Apple M1. Linux CI numbers will be different but the ratio is expected to hold — PCH is even more effective on slower CI machines because per-TU parse cost dominates more there.
Next levers (not in this branch)
If the 55 s wall is still not enough:
| Lever | Estimated win | Cost |
|---|---|---|
ccache integration in CI |
hot rebuild 55 s → ~5 s | 30 min setup |
-O0 for the CGAL test target in CI PRs |
55 s → ~30 s | 1 h policy doc |
extern template (lever #2 above) |
55 s → ~52 s | 2 h + Eigen version tracking |
| Header split (lever #3) | downstream-only | 1 day + API risk |
| C++20 Modules | speculative; experimental in Apple clang 17 | weeks |