Conversation
tuhaihe
force-pushed
the
meson-adapt
branch
2 times, most recently
from
September 1, 2026 06:15
8cc05ea to
ff5685d
Compare
The PG16 merge brought in 266 meson.build files verbatim from upstream, with no Cloudberry awareness at all: no GP options, no GP source files, no GP catalogs, and several upstream targets that Cloudberry does not build. This wires up enough of it to configure, build, install and initdb a working single-node Cloudberry with meson. Build options and configuration: - 24 GP options in meson_options.txt, defaults matching configure.ac - GP_VERSION / GP_MAJORVERSION / GP_VERSION_NUM computed via getversion, without writing a VERSION file into the source tree - ~20 GP config macros (USE_ORCA, USE_INTERNAL_FTS, FAULT_INJECTOR, ...) - libcurl, libbz2 and libuv detection Backend: - 21 new meson.build files for cdb, fts, task, crypto, the AO/AOCS/bitmap access methods, resgroup and the other GP utils subdirectories - GP sources added to 21 existing meson.build files - 40+ GP catalog headers and 5 .dat files registered; gp_version_at_initdb.dat generated from its template next to a build-dir copy of its header, since genbki.pl derives each .dat path from its .h path - system_views_gp.sql generated, plus cdb_schema.sql and the GP catalog SQL - libpq frontend sources compiled into the backend for QD<->QE communication - libpqwalreceiver built as a normal backend object, not a separate module Divergences from upstream that meson was still assuming: - geqo is not built by Cloudberry (not in optimizer/Makefile SUBDIRS) - gen_node_support.pl is unused; NodeTag is hand-maintained in nodes.h and copyfuncs.funcs.c / copyfuncs.switch.c are committed to the repository - ecpg is not built (commented out of src/interfaces/Makefile SUBDIRS) - man pages cannot be built; stylesheet-man.xsl is absent from the tree - libpq needs an explicit -DFRONTEND, as its sources include c.h rather than postgres_fe.h so the same files can also be compiled into the backend Extensions: - contrib/interconnect, required as a preload library by initdb - the eight always-on gpcontrib extensions array_userfuncs.c included ../../catalog/pg_type_d.h, which only resolves in an in-place build; corrected to the normal include path. ORCA, gpfdist, pxf, gpMgmt, the flag-gated gpcontrib extensions and the test suites are not wired up yet. Assisted-by: Claude Code
ORCA is 945 C++ translation units spread over 34 near-identical leaf Makefiles that only emit objfiles.txt. In meson the whole tree collapses to two files, built as static libraries and linked whole into the backend: - C++ is enabled before the LLVM block, gated on the orca option, because upstream only enables it for the JIT - xerces-c is probed with a library check plus a header check honouring extra_include_dirs / extra_lib_dirs, and scoped to the ORCA targets rather than leaked into the global LIBS the way configure does - cpp_std=c++14 as an option, where configure appends -std=c++14 to CXX - the strict -Werror -Wextra -Wpedantic triple is kept on Linux and dropped on darwin, matching gporca.mk - postgres gets link_language: 'cpp', since src/backend/Makefile links it with $(CXX) and meson would otherwise pick the C driver libpq_binddomain() used ldir without declaring it. That code only compiles under ENABLE_NLS, which Cloudberry does not build, so it had never been reached; the declaration is restored from upstream anyway. Verified: 3048 targets build, initdb succeeds, gp_opt_version() reports "GPOPT version: 4.0.0, Xerces version: 3.3.0", and a query planned by ORCA returns the expected rows. Assisted-by: Claude Code
…sions - gpfdist, with apr located through apr-1-config rather than pkg-config. configure uses the config tool (--with-apr-config) for a reason: the system apr-1.pc on macOS points at an include directory that does not exist. Adds an apr_config option mirroring --with-apr-config. libevent is required, libyaml optional and gating -DGPFXDIST, and gfile.c keeps its per-file -DFRONTEND. - pxf_fdw, with libcurl as a dependency instead of a hardcoded -lcurl - the debug-extension set (gp_debug_numsegments, gp_inject_fault, gp_replica_check including its python script, reject_partition_fullscan), zstd (MODULE_big name differs from the directory name) and orafce src/Makefile.global.in's Cloudberry @variables@ are now filled in. They are defined by configure via AC_SUBST and consumed by recursive make -- for example gpcontrib/Makefile reads enable_gpcloud / enable_pxf / with_diskquota / enable_debug_extensions, and src/bin/gpfdist/Makefile reads have_yaml / EVENT_LIBS / apr_*. The meson build does not use them itself, but Makefile.global is installed as part of PGXS, so out-of-tree extensions built with USE_PGXS=1 still read them; leaving them empty would silently change their behaviour. meson setup no longer warns about missing substitutions. Assisted-by: Claude Code
Builds Cloudberry with meson in the same rocky9 build container the autoconf workflows use, installs it, and runs a smoke test. This does not replace build-cloudberry.yml: until the two paths are proven equivalent the autoconf build stays authoritative, and this exists to keep the meson files from rotting and to surface divergence early. The workflow header maps each flag of the reference configure line in coverity.yml onto its meson option, and records which ones are not reachable from meson yet (gpcloud, pax, mapreduce, diskquota, gp_stats_collector) and that --with-pythonsrc-ext was dropped deliberately. The smoke test lives in devops/build/automation/cloudberry/scripts so it can be run by hand against any install prefix. It checks the installed binaries and GP extensions, that the generated catalog data is present, that initdb succeeds, and then in single-user mode that the GP catalogs exist, that the gp_* views generated from system_views_gp.in were created, that gp_opt_version() reports a working ORCA, and that an append-only table actually stores and returns rows. Assisted-by: Claude Code
Four levels of recursive make that exist almost entirely to copy files collapse into per-directory install_data calls. The 0755 / 0644 split is inconsistent in the autoconf build (base.py is installed as a script while its siblings are data, and so on), so it is reproduced file by file rather than with a blanket install_subdir. - 32 utilities into bindir, 11 into sbindir, gppylib and its eight subpackages into libdir/python/gppylib, the gpcheckcat / gpconfig / gpssh module directories and bin/lib into bindir - the two C programs: stream into bindir/lib with the backward-compatible bindir/stream/stream symlink, and ifaddrs into libexecdir, which is where gp_bash_functions.sh looks for it - the six gpdemo scripts, which live in gpAux/gpdemo but are installed from gpMgmt/bin - gp_bash_version.sh generated from its template; configure.ac does this with an inline sed because the placeholder is ##version## rather than @Version@, and writes the result into the source tree putversion rewrites $Revision$ in ~48 files *after* installing them, which meson cannot do. subst_version.py performs the same substitution on the way into the build directory instead, so the installed file is already correct. It is applied to exactly the files the Makefiles apply putversion to; the two that are left untouched here are untouched by the autoconf build as well. Not ported: --with-pythonsrc-ext (dropped by decision - it downloads psutil, PyYAML and PyGreSQL from PyPI during install), and the behave test targets. Assisted-by: Claude Code
Everything left that the autoconf build produces and meson did not: the frontend binaries under src/bin (gpfts, gpnetbench, pg_alterckey), and the gpcontrib components that need more than a PGXS-shaped extension -- gpmapreduce, gp_stats_collector, gpcloud with gpcheckcloud, and diskquota. Three needed something other than a straight translation. gpmapreduce's generated scanner and parser hardcode their own output names with %output= and %option outfile=, which override the -o meson passes, so bison and flex run with the build directory as their working directory. gp_stats_collector generates C++ from protobuf, so protoc runs from the extension root: its -I has to be a textual prefix of every input path, and the sources import "protos/xxx.proto". diskquota is a cmake project in tree. Rather than shell out to cmake from meson -- two build systems, two configurations, one of them invisible to ninja -- its sources are listed directly, which is what the rest of this port does with everything else. libpostgres.so comes with them, and with the symbol map the Makefiles generate: the shared backend must not re-export libpq's symbols, or an extension linked against libpq gets the backend's copies. gen_symbol_map.py derives the list from src/interfaces/libpq/exports.txt, as the Makefile does, and emits a version script on Linux and an -unexported_symbols_list on darwin. Assisted-by: Claude Code
PAX ships a standalone CMake project driven from its Makefile. The meson source list mirrors the groups in src/cpp/cmake/pax.cmake for the default configuration (USE_PAX_CATALOG=ON, USE_MANIFEST_API=OFF, VEC_BUILD off), which is 93 files. Note the mixed extensions: one source is .cpp where the rest are .cc, and missing it produced a module that built cleanly but failed to load with an undefined pax::tools::OrcDumpReader::Dump(). Two pieces of code generation: - protoc, declared in a nested storage/proto/meson.build so the generated headers land in a storage/proto directory inside the build tree, which is how the sources include them. CMake instead writes them back into the source tree. - the extension SQL, which cmake/pax.cmake produces by compiling tools/gen_sql.c and redirecting its stdout into pax-cdbinit--1.0.sql in the source tree; meson captures the output instead. yyjson is not needed: it is only linked when USE_MANIFEST_API is on and USE_PAX_CATALOG is off, which is the opposite of the defaults. tabulate is needed unconditionally (used by a UDF), so PAX requires that submodule. pg_waldump also needs PAX when it is enabled: rmgrdesc.c includes "paxc_desc.h" under USE_PAX_STORAGE, and the Makefile symlinks paxc_desc.[ch] in from contrib/pax_storage. Verified end to end: the pax access method is registered, and a table created with USING pax stores and returns rows. CI now enables gpcloud, mapreduce, diskquota and pax as well, and the smoke test checks the additional binaries and the generated PAX SQL. gp_stats_collector remains off there for the reason recorded in the workflow. Assisted-by: Claude Code
Two places where the meson build did more than configure.ac asks for. NLS: upstream's meson exposes an nls option; Cloudberry's configure.ac has had --enable-nls deliberately removed since 2015, and the translation catalogues have gone stale in the years since. The option now defaults to disabled and errors if turned on, rather than offering a build nobody maintains. Warnings: configure.ac keeps -Wdeclaration-after-statement commented out, because "GPDB code is full of declarations after statement", and adds -Wno-unused-but-set-variable. meson had neither, which is the whole difference between 2263 warnings and 10 on the same tree. The meson floor moves to 0.61 here, for install_symlink. That turns out to be the wrong trade and is settled later in the series; see "meson: fixes from installing and running the result". Assisted-by: Claude Code
Mirrors build-cloudberry.yml's shape -- check-skip, build, test, report -- so
that test jobs can be added alongside cluster-test later.
The build matrix is the part that earns its keep today. Most of the bugs found
while porting were in option combinations rather than in any single
configuration: C++ being enabled only for ORCA, PAX requiring the shared
backend, extensions gated on flags never exercised together. Three
configurations cover that surface:
full every component this branch supports
minimal everything optional off, which catches conditional
guards that only work when enabled
no-shared-backend -Dshared_postgres_backend=false, a path nothing else
covers; PAX comes off with it since PAX links against
libpostgres.so
cluster-test goes past what the smoke test can reach. meson-smoke-test.sh only
runs single-user mode, so it never starts a segment. The new script brings up a
real multi-segment cluster with gpinitsystem and then checks that data actually
distributes (1000 rows across segments, not just 1000 rows), that a plan
contains a Motion node, that ORCA plans a distributed join, and that ao_row and
pax tables work distributed.
That also makes it the first test of two things the build alone cannot prove:
the gpMgmt install is complete enough to run gpinitsystem, and dropping
--with-pythonsrc-ext is workable -- psutil, PyGreSQL and PyYAML are installed
from python-dependencies.txt instead of being vendored.
Two checks are new in the build job: meson_version feature warnings are fatal,
since they are cheap to fix and easy to let accumulate, and the compiler
warning count is reported to the step summary, because a jump there usually
means the meson flags have drifted from configure.ac.
Regression suites are still not run. src/test/regress and friends are
make-driven and include src/Makefile.global from the source tree, which only a
configure run produces, so they cannot execute against a meson-only tree.
meson also now installs $prefix/cloudberry-env.sh, which gpMgmt/Makefile
generates and which demo_cluster.sh sources; without it the cluster test could
not run, and an install was not usable the way an autoconf one is.
Assisted-by: Claude Code
None of these were visible from a build. Every one came out of installing the tree and starting a real cluster on it, and each would have shipped a build that compiles and links and then does not work. libdir. meson derives it from the platform, which is lib64 on RHEL. cloudberry-env.sh hardcodes PYTHONPATH=$GPHOME/lib/python and LD_LIBRARY_PATH=$GPHOME/lib, so gppylib and libpostgres.so landed where nothing looks for them. Every file present, install unusable. PG_VERSION_STR. gpMgmt parses select version() and insists on "(Apache Cloudberry <version> build <build>)". With upstream's string, gpstop dies on a healthy cluster with "too many tokens in version". gpMgmt files. Ten the autoconf build installs and this port did not, nine of them Python modules, so nothing failed to build: gppylib/commands/base.py, all of gppylib/util, gppylib/programs/gppkg.py, gppylib/system/ComputeCatalogUpdate.py, gppylib/gpMgmttest, lib's crashreport.gdb, and the foreign-key JSON gpcheckcat reads, which is generated from the same headers as postgres.bki. gpMgmt Python modules. --with-pythonsrc-ext downloads psutil, PyGreSQL and PyYAML mid-build, which has no honest meson equivalent, so it was dropped. What should not have gone with it is the install: the Makefile's half of that option only copies what is already in gpMgmt/bin/ext, and without the same conditional copy there is no supported way to get the modules into an install at all. meson floor. Raising it to 0.61 for install_symlink turned roughly thirty-five deprecations in upstream's own meson files into warnings on every configure -- meson only reports a deprecated API once the project claims a version that has it deprecated. Three symlinks are not worth that, nor worth diverging from upstream on the one line every future merge touches. The floor goes back to upstream's value and the symlinks are made by an install script. Per-directory warning flags. Five directories relax a warning in their own Makefile. Two need more than an extra c_args: meson compiles the whole backend as one target, so src/backend/task moves into a static library that carries the flag, and plpgsql, where the Makefile scopes -Wno-array-bounds to pl_exec.o alone, gets a library for that one file. Assisted-by: Claude Code
The workflow that came with the port built one configuration and ran a fixed smoke test. Everything here is a response to something the first real runs turned up. An install-parity job. The port lists installed files by hand, so a file added to a Makefile and not to the meson.build beside it is silently dropped. Ten such files went unnoticed until gpstop died on a cluster that had come up cleanly; no build catches that, and a test catches only the fraction it exercises. The check is static, finishes in seconds, and reports before the build it would not have been caught by. A configuration-aware smoke test. The old one asserted the full configuration's binaries and GUCs, so every other matrix entry failed for doing what it was told -- missing bin/gpmapreduce, and a FATAL from optimizer=on in a build without ORCA. The expected set now comes from meson-info, so one test serves the whole matrix, which is what makes having a matrix worth anything. gpcheckcat in the cluster test. It is the only thing that reads the foreign-key JSON, so an install missing that file is silent everywhere else. The PyPI downloads move to the front of the job and are cached. They were happening after a forty-minute build, which put the network on the critical path at the worst moment; the Makefile's download target also pip-installs wheel and Cython when absent, and they are absent from the build image, so those go in with meson and ninja instead. Nothing between meson setup and the end of the build reaches the network now, and an offline build only has to stage the three tarballs. Building src/common's frontend library on its own, before the full build. That directory is compiled both ways, and a missing generated-header dependency there fails or not depending on scheduling -- one red matrix entry and one green one on the same commit. Also: the matrix drops to two opposite configurations rather than three, and the workflow triggers on main only, pg16-ci having been a branch in a personal fork. Assisted-by: Claude Code
curl without -f treats an HTTP error as success: it writes the error body to
the output file and exits 0. The guard around each download then only asks
whether the file exists, so a single 404 or bad gateway leaves a broken
tarball that every later run skips over as already downloaded. The build
fails much later, in tar, with nothing pointing at the cause.
$ curl -sSL .../psutil-99.99.99.tar.gz -o t # 404
$ echo $?; wc -c < t
0
0
Add -f so the request fails, --retry 3 so a transient one does not, and make
each guard -s rather than -f so a zero-byte leftover is retried instead of
trusted. curl leaves the partial file behind, and --remove-on-error needs a
curl newer than rocky8 ships, so the guard is the portable place to fix this.
Verified in the build container: a 404 now stops the target with curl's exit
status and leaves no file, and with no route to PyPI a zeroed tarball fails
the target rather than silently persisting.
Assisted-by: Claude Code
src/common is compiled twice, once with -DFRONTEND. percentrepl.c and
kmgr_utils.c include utils/builtins.h unguarded, and that pulls in
utils/fmgrprotos.h, which is generated for the backend and is not a declared
dependency of anything frontend. Nothing orders the two, so ninja is free to
compile the frontend copy first:
../src/include/utils/builtins.h:21:10: fatal error:
utils/fmgrprotos.h: No such file or directory
Which is exactly what happened, and only sometimes: on one commit the full
matrix entry went green and the minimal one went red, on the same runner
image, differing only in how much else there was to schedule. make never
showed it because every Makefile there depends on submake-generated-headers,
so the ordering is forced.
Everything the offending headers declare -- GpIdentity, terminal_fd, pg_ltoa
-- is used only from #ifndef FRONTEND blocks, so the includes belong inside
one. kmgr_utils.c also included postgres.h unconditionally, which a FRONTEND
translation unit must not do at all; it uses nothing from either header.
Verified as a deterministic before/after in a cold build directory, building
just the two frontend objects: both fail on the old sources, both compile on
the new ones.
Assisted-by: Claude Code
Three things landed on main that the meson build had to catch up with. Node support functions are generated now (apache#1823). The meson port had dropped gen_node_support.pl because Cloudberry maintained NodeTag by hand; main deleted the hand-written copyfuncs.funcs.c and friends, so the generator is back, fed with Cloudberry's header list. That list is copied from src/backend/nodes/Makefile in the same order, because the order numbers the NodeTag enum: a different order is a meson build and an autoconf build of one release disagreeing on node tags. Cloudberry's generator also writes the outfast/readfast pair used to serialise plans for dispatch, so outfast.c and readfast.c move next to copyfuncs.c, where the generated files are included from. gp_relaccess_stats joined gpcontrib's unconditional list. --enable-datalake-fdw appeared in Makefile.global.in. The meson build does not build datalake_fdw in tree -- neither does the autoconf CI, which builds it with PGXS because it needs Arrow and Parquet -- so PGXS sees `no`. Assisted-by: Claude Code
Cloudberry's configure.ac defaults --with-blocksize and --with-wal-blocksize to 32; meson_options.txt kept upstream's 8. Nothing about the build notices: it compiles, installs, initialises and serves queries. But block size is on-disk format, so a data directory made by one build cannot be opened by the other, and a meson build and an autoconf build of the same release could not share a cluster, a basebackup or pg_upgrade. The regression suite noticed at once, from the first run against the meson build: an 8 kB page where the expected output has 32 kB, a 32 kB row that is "too big: maximum size 8160", and the page counts and costs of nineteen planner tests off accordingly. With the defaults matched, installcheck-good passes those unchanged. Assisted-by: Claude Code
Every one of these built cleanly and passed the smoke and cluster tests. The regression suites, run against a meson install, found them. pageinspect.so without bmfuncs.c. Cloudberry added bitmap-index page functions to the Makefile's OBJS; the meson.build is upstream's list. The module loads, and CREATE EXTENSION fails: could not find function "bm_metap". installcheck-good creates it before running anything, so the whole suite bailed out on it. Five contrib modules not built at all. The top-level GNUmakefile builds a list of contrib modules on every make, and five are Cloudberry's own: extprotocol, formatter, formatter_fixedwidth, indexscan and try_convert. No meson.build existed for any of them. zstd without its catalog entry. gpcontrib/zstd's Makefile installs zstd_compression.sql into cdb_init.d, which initdb runs to register the compressor in pg_compression; the meson.build had an install_data() with no files in it. Every AO table asking for compresstype=zstd failed with "unknown compress type". gpconfig's list of GUCs that may not go in postgresql.conf. gpMgmt/bin's install target generates share/greenplum/gucs_disallowed_in_file.txt by scanning guc.c; without it gpconfig warns on every call and checks nothing. Generated here with the same parser, and checked by the smoke test, since no static check can see a file a script writes. pg_buffercache's 1.4.1 upgrade script. Cloudberry's control file names default_version 1.4.1 and its Makefile installs pg_buffercache--1.4--1.4.1.sql; the meson.build stopped at 1.4, so CREATE EXTENSION pg_buffercache failed with "no installation script". The cdb/ and task/ headers. src/include/Makefile installs them with the rest, and nodes/primnodes.h includes cdb/cdbpathlocus.h, so without them no extension that includes a node header -- datalake_fdw, built with PGXS -- could be compiled against a meson install. --enable-ic-udp2 is off by default and builds contrib/udp2 through its own CMake project, which meson does not do yet. Turning the option on used to define ENABLE_IC_UDP2 for a server with no udp2 module; it now refuses. Assisted-by: Claude Code
The suites cannot run without their harness, and the meson build had only upstream's part of it: pg_regress and a regress.so missing Cloudberry's functions. pg_regress compares results with gpdiff.pl, found beside its own executable, and refuses to start unless `gpdiff.pl --version` prints the server's GP_VERSION -- which comes from GPTest.pm, a file configure writes from GPTest.pm.in into the source tree. So GPTest.pm is generated, and it, gpdiff.pl and the other scripts the Makefile installs into pgxs/src/test/regress are installed there and copied beside the build tree's pg_regress. Without them, no extension could run `make USE_PGXS=1 installcheck` against a meson install either. regress.so gains regress_gp.c and libpq, and is installed into pkglibdir as the Makefile does. The tests' client programs are built beside it -- twophase_pqexecparams, extended_protocol_resqueue -- as are the hooktest and query_info_hook_test modules, in subdirectories of those names, since the tests LOAD them by path. isolation2's pg_isolation2_regress, module and clients get a meson.build; the autoconf build installs none of it, and neither does this. gpmapreduce's test module is built for the same reason. contrib/interconnect's Makefile includes Makefile.interconnect, which configure writes from Makefile.interconnect.in to learn enable_ic_proxy; it is rendered here from the values Makefile.global is. Everything lands at the same path relative to the build directory as it does in an in-tree build, which is what lets the next commits run the Makefiles' own test targets on it. extended_protocol_resqueue.c had an unused variable; the Makefile compiles it without CFLAGS, so only meson's -Wall saw it. Assisted-by: Claude Code
Cloudberry's configure.ac defines a handful of macros upstream's configure
does not, and the meson build, being upstream's, left them out. Each
difference below is one the code tests; each built, installed and started
without complaint.
USE_FLOAT4_BYVAL. configure defines it unconditionally ("always defined in
GPDB"). indexam.c still frees the previous float4 ORDER BY distance when it
is not defined, and a float4 is passed by value either way -- so on the
meson build every GiST index scan ordered by a float4 distance pfree()d a
number and took the backend down: btree_gist's and pg_trgm's installcheck.
FLOAT4PASSBYVAL, USE_FLOAT8_BYVAL and FLOAT8PASSBYVAL come along; the
headers define the last three the same way already.
HAVE_LIBRT. --with-rt is on by default, and the interconnect's
retransmission timer uses CLOCK_MONOTONIC only when it is defined. Without
it, both interconnects timed packets by gettimeofday(), which steps when
the wall clock does.
PACKAGE_NAME, PACKAGE_VERSION, PACKAGE_STRING, PACKAGE_BUGREPORT,
PACKAGE_URL, PACKAGE_TARNAME. meson hardcoded upstream's, so postgres
--help and initdb --help sent bug reports to pgsql-bugs and pointed at
postgresql.org. They are read from configure.ac's AC_INIT, so the two
builds cannot drift apart.
PG_VERSION_NUM and PG_MINORVERSION_NUM. configure takes the minor version
from PACKAGE_VERSION, which is Cloudberry's own, so a 16.9 code base says
160000. That looks like an oversight in configure.ac, but a meson build and
an autoconf build of one release telling extensions different things would
be worse; meson computes it the same way, and a comment says why.
PG_KRB_SRVTAB. Upstream's meson.build formats it as
'FILE:/@0@/krb5.keytab)', with a stray parenthesis and a relative
sysconfdir, so the default krb_server_keyfile named a file that cannot
exist. It is now FILE:$(sysconfdir)/krb5.keytab, as configure has it.
These were found by comparing the two pg_config.h files; a later commit
makes that comparison part of CI.
Assisted-by: Claude Code
Upstream compiles C++ only for the LLVM JIT, so its meson.build gives C++ the common flags only when LLVM is found. Cloudberry compiles ORCA, gpcloud and gp_stats_collector as C++ with or without LLVM, and configure gives them -fno-strict-aliasing and -fwrapv in CXXFLAGS; the meson build gave them neither. Those two change what the compiler may assume, so the same C++ compiled to different code under the two build systems. C++ now gets them whenever there is a C++ compiler, with the warning flags configure's CXXFLAGS has: the C list without -Werror=vla, which configure keeps to C and which gpopt does not build under. PAX is C++ as well, but the autoconf build compiles it with its own CMake project, not with configure's CXXFLAGS -- no -fwrapv, no -fno-strict-aliasing, none of configure's warnings. Its target undoes the two -f flags, and the one warning its sources trip, so that PAX too compiles to the same code under both. The rendered Makefile.global had the same blind spot, and PGXS builds extensions from it: CXX and CXXFLAGS were empty without LLVM, so no C++ extension -- datalake_fdw is one -- could be built against a meson install at all. CXX carries -std=c++14, as configure's AX_CXX_COMPILE_STDCXX puts it there. CFLAGS had no optimisation level, no -g and no -Wall. meson passes those from buildtype and warning_level rather than from the flag lists Makefile.global is written from, so every PGXS extension compiled at -O0. They are spelled out for PGXS now, which also makes pg_config --cflags report what the server was built with. The one C++ warning the build now shows, a -Wcast-function-type cast in nodeFuncs.h included from gpopt, is the one the autoconf build shows too. Assisted-by: Claude Code
Upstream's 9244c11 added a check for X509_get_signature_info() to configure.ac and configure, and its template line to pg_config.h.in. Cloudberry has the check but lost the template line when the PostgreSQL 16 history was merged in -- and config.status only writes a macro whose #undef the template has. So configure ran the check, found the function, and defined nothing. The backend's and libpq's OpenSSL code then fall back to X509_get_signature_nid() for tls-server-end-point channel binding, which does not work for RSA-PSS certificates; that case is what the upstream commit fixed. Found by comparing configure's pg_config.h with the meson build's, which defines it. Assisted-by: Claude Code
Each gap the regression suites turned up was one the static parity check
could have reported in seconds, had it looked. It looks now.
install also follows install's helper targets -- `install: install-data`
is where gpcontrib/zstd kept its INSTALL_DATA -- reads PGXS-style
Makefiles, which install without an install target at all, reads
GNUmakefile as well as Makefile, reads the DATA and SCRIPTS an
extension installs, and covers all of contrib -- whose meson.build
files list scripts by hand, as Cloudberry's own do -- and
src/test/regress, whose Makefile installs the whole gpdiff
toolchain. pg_buffercache's 1.4.1 script was the gap in contrib.
source new: every object a Makefile compiles into something meson also
builds must have its source named in a meson.build at or above it.
pageinspect's bmfuncs.c, and nothing else in the tree.
contrib new: every contrib module the top-level GNUmakefile builds by
default is subdir()'d by contrib/meson.build. The five that were
not.
options new: the options that fix the on-disk format, and the port, have
configure's default. The block sizes.
Each was checked against the tree before its fix and reports that gap; the
current tree is clean. The whole run takes a third of a second.
Assisted-by: Claude Code
The workflow built, installed, smoke-tested and started a cluster, and ran no tests. Now it runs build-cloudberry.yml's suites -- installcheck-good with ORCA off and on, isolation2 and its crash, hot-standby, expand/shrink and parallel-retrieve-cursor schedules, the parallel and fixme schedules, the recovery TAP tests, and the contrib and gpcontrib installchecks, including the three stats extensions that need preloading -- as the same Makefile targets, through the same demo-cluster script, test runner and result parser, in that workflow's entry format. The suites stay defined where they are. They expect an in-tree configured build, and every piece of one is a meson build product: the build job packs them -- the rendered Makefile.global, pg_regress, regress.so, the clients -- and meson-test-harness.sh lays them into the test job's checkout, then turns each dir:target into make arguments that keep make from rebuilding the harness meson built, or reinstalling the EXTRA_INSTALL modules the meson install already has. contrib and gpcontrib run in-tree as well, as they do in build-cloudberry.yml: a dozen of their Makefiles point --init-file at the source tree's init_file, which no install carries, so USE_PGXS=1 would be a different run from the one the project relies on. The script's header explains the details, including why extraction resets the timestamps. Porting the suites to `meson test` would give each a second definition to keep in step with the Makefile, and a green run of the copy would not show that the meson build passes the tests the project actually uses. Five of build-cloudberry.yml's entries are not moved yet, and the workflow says why for each: PAX's suites and diskquota's need CMake-built pieces the harness does not carry, datalake_fdw needs Arrow and a PGXS build step, singlenode needs directories meson does not build yet, and resgroup needs cgroups in the container. Also: gp_stats_collector is enabled in the full configuration, as configure-cloudberry.sh enables it -- it builds on Linux without a warning -- and the actions move to the versions the rest of the repository uses. Assisted-by: Claude Code
The same source compiled against two different pg_config.h files is two different programs, and nothing at build time says so. That is how the meson build had 8 kB pages, freed float4 datums in GiST index scans and timed the interconnect by the wall clock: in each case a macro configure defines and meson did not, and a build that compiled, installed and started. meson-config-parity.py compares two pg_config.h files and fails on a macro that differs -- present in one and not the other, or with another value -- when some source file uses it as code, outside comments and strings. The two build systems do not define the same set: autoconf carries boilerplate like STDC_HEADERS that upstream's meson build dropped, and records library checks nothing reads. Those 31 are listed but cannot change the build. The four that are referenced and intended -- CONFIGURE_ARGS, PG_VERSION_STR, PG_KRB_SRVTAB and restrict -- are in the script, each with its reason. The full build carries the configure line it stands for, so the workflow's option table is now checked rather than only documented. The build runs configure with it in a worktree of the same commit and compares, before it compiles. configure runs in-tree because a VPATH configure writes an empty GP_VERSION: it runs ./getversion from its working directory. The step adds about a minute. Assisted-by: Claude Code
tuhaihe
force-pushed
the
meson-adapt
branch
from
September 29, 2026 16:57
ff5685d to
5a7f1a8
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #ISSUE_Number
What does this PR do?
Type of Change
Breaking Changes
Test Plan
make installcheckmake -C src/test installcheck-cbdb-parallelImpact
Performance:
User-facing changes:
Dependencies:
Checklist
Additional Context
CI Skip Instructions