Repository navigation
build and publish ml_metadata_store_server container image for ARM64 #10308
Description
Activity
/cc @chensun @zijianjoy
For those who want to test, I have made a forked repo in the deployKF org with the ARM versions of the
gcr.io/tfx-oss-public/ml_metadata_store_serverimage. You can test a patched version ofml-metdataversion1.14.0by using the following container:ghcr.io/deploykf/ci/ml_metadata_store_server:sha-cad0c56ghcr.io/deploykf/ml_metadata_store_server:1.14.0-deploykf.0- (EDIT: use this one now that we have a proper release)
Note, building under emulation on GitHub actions took about 5 hours:
Reacted by xixiciGreat. I pull this image and run it correctly. Then, I am finding
gcr.io/ml-pipeline/metadata-writerandgcr.io/ml-pipeline/metadata-envoywith ARM version.@xixici can you confirm what you are saying?
Because
gcr.io/ml-pipeline/metadata-writer:2.0.5andgcr.io/ml-pipeline/metadata-envoy:2.0.5(and all other versions) are only published for ADM64.I assume you mean that they work via Rosetta Emulation on a MacBook?
This issue has been automatically marked as stale because it has not had recent activity. It will be closed if no further activity occurs. Thank you for your contributions.
- addedlifecycle/staleThe issue / pull request is stale, any activities remove this label.The issue / pull request is stale, any activities remove this label.
on Mar 13, 2024 This issue has been automatically closed because it has not had recent activity. Please comment "/reopen" to reopen it.
/reopen
@thesuperzapper: Reopened this issue.
Details
In response to this:
/reopen
Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes/test-infra repository.
Clearly, this is going to take some time, so I will prevent the bot from closing it.
/lifecycle frozen
- added and removedlifecycle/staleThe issue / pull request is stale, any activities remove this label.The issue / pull request is stale, any activities remove this label.
on Apr 5, 2024 I have successfully completed metadata envoy 2.0.5 and built it in an ARM environment
The following are the construction steps:
docker pull --platform linux/arm64 envoyproxy/envoy:v1.16.0
In the directory of pipelines/third_party/metadata_envoy
1、modify Dockerfile and config proxy info
FROM envoyproxy/envoy:v1.16.0 RUN apt-get -o Acquire::http::proxy="http://proxy:port" update -y && \ apt-get -o Acquire::http::proxy="http://proxy:port" install --no-install-recommends -y -q gettext openssl COPY third_party/metadata_envoy/envoy.yaml /etc/envoy.yaml # Copy license files. #RUN mkdir -p /third_party COPY third_party/metadata_envoy/license.txt /third_party/license.txt ENTRYPOINT ["/usr/local/bin/envoy", "-c"] CMD ["/etc/envoy.yaml"]2、modify envoy.yaml
admin: access_log_path: /tmp/admin_access.log address: socket_address: { address: 0.0.0.0, port_value: 9901 } static_resources: listeners: - name: listener_0 address: socket_address: { address: 0.0.0.0, port_value: 9090 } filter_chains: - filters: - name: envoy.http_connection_manager typed_config: "@type": type.googleapis.com/envoy.config.filter.network.http_connection_manager.v2.HttpConnectionManager codec_type: auto stat_prefix: ingress_http route_config: name: local_route virtual_hosts: - name: local_service domains: ["*"] routes: - match: { prefix: "/" } route: cluster: metadata-cluster max_grpc_timeout: 0s cors: allow_origin_string_match: - exact: "*" allow_methods: GET, PUT, DELETE, POST, OPTIONS allow_headers: keep-alive,user-agent,cache-control,content-type,content-transfer-encoding,custom-header-1,x-accept-content-transfer-encoding,x-accept-response-streaming,x-user-agent,x-grpc-web,grpc-timeout max_age: "1728000" expose_headers: custom-header-1,grpc-status,grpc-message http_filters: - name: envoy.grpc_web - name: envoy.cors - name: envoy.router clusters: - name: metadata-cluster connect_timeout: 30.0s type: logical_dns http2_protocol_options: {} lb_policy: round_robin load_assignment: cluster_name: metadata-cluster endpoints: - lb_endpoints: - endpoint: address: socket_address: address: metadata-grpc-service port_value: 80803、finally,build image
In the directory of pipelines/backend/Makefile.PHONY: metadata_envoy metadata_envoy: cd $(MOD_ROOT) && docker build -t registry.cnbita.com:5000/kubeflow-pipelines/metadata-envoy:2.0.5-arm -f third_party/metadata_envoy/Dockerfile .run make metadata_envoy
Clearly, this is going to take some time, so I will prevent the bot from closing it.
/lifecycle frozen
ml_metadata_store_serverdoes build on arm64. I got it working and opened a PR for the first piece: google/ml-metadata#247.I mention this because the discussion here has assumed the blocker is deep, and it turned out to be a hardcoded build config rather than a portability problem in the C++ code.
What actually blocks it
Building
//ml_metadata/metadata_store:metadata_store_serveron aarch64 hits exactly two things.1. The Dockerfile hardcodes the x86_64 Bazel installer.
ml_metadata/tools/docker_server/Dockerfilefetchesbazel-$BAZEL_VERSION-installer-linux-x86_64.sh, so the image can only be built on x86_64. Bazel does publishlinux-arm64(for both 5.3.0 and master's 7.7.0), just as a plain binary rather than an installer script. That is what google/ml-metadata#247 fixes.2. Abseil emits an ARMv8.3 instruction the assembler rejects at the default baseline. On current master,
absl/debugging/stacktrace.ccemitsxpaclri(pointer authentication):external/abseil-cpp/absl/debugging/stacktrace.cc [for tool] failed /tmp/ccRV18O3.s:166: Error: selected processor does not support `xpaclri'Adding
--copt=-march=armv8.3-a --host_copt=-march=armv8.3-ato thebazel buildinvocation clears it. Note--host_coptis genuinely required alongside--copt— the failing target compiles[for tool], i.e. in the exec configuration, which--coptdoes not affect. With--coptalone the error is byte-identical, which makes it look like the flag did nothing.-march=armv8-a+pauthwould be preferable to raising the whole baseline, sincexpaclriis HINT-space and a no-op on older cores, whereasarmv8.3-amakes the binary require ARMv8.3 hardware (Graviton2 is ARMv8.2). But GCC 9 on theubuntu:20.04builder rejects it (invalid feature modifier 'pauth'), so a newer builder base image would be the cleaner fix.Result
With those two changes, master (
be943b8) builds clean on aarch64 with zero errors:arch=arm64, 131 MB $ docker run --entrypoint /bin/metadata_store_server ... --help metadata_store_server: Warning: SetUsageMessage() never called Flags from external/com_github_gflags_gflags/src/gflags.cc: ...Built on an NVIDIA GB10 (aarch64, Ubuntu 24.04, Docker 28.5.1).
It unblocks Kubeflow Pipelines on ARM
I deployed the resulting image into a Kubeflow Pipelines 2.16.1 install on an arm64 kind cluster (alongside the KFP images rebuilt for arm64 — details in #10309).
metadata-grpc-deploymentgoes to1/1 Running, andml-pipeline-persistenceagentandml-pipeline-scheduledworkflow, which had been crash-looping against the unavailable metadata service, recover with it.Before the rebuild, the stock image on that same node gave the expected:
exec /bin/metadata_store_server: exec format errorOne note on older release branches
On
v1.14.0a third change was also needed, which does not apply to master.ml_metadata/postgresql.BUILDgeneratespg_config.hfrom a fixed list of lines captured from aconfigurerun on x86_64 —PG_VERSION_STRstill records"PostgreSQL 12.1 on x86_64-apple-darwin19.2.0"— asserting x86-only CPU capabilities with noselect()on architecture anywhere. WithHAVE__GET_CPUIDdefined,src/port/pg_bitutils.cincludes<cpuid.h>, which only exists on x86:external/postgresql/src/port/pg_bitutils.c:16:10: fatal error: cpuid.h: No such file or directoryOn master
HAVE__GET_CPUIDis already/* #undef */, and my master build confirms the remaining unconditional#define HAVE_X86_64_POPCNTQ 1does not break aarch64 on its own — so nopg_config.hchange is needed there. Recording it only because it still affects release branches, and because the underlying pattern (aconfiguresnapshot hardcoded for one CPU) could resurface.Scope
google/ml-metadata#247 covers only the Dockerfile change, since the
-marchflags need to be conditional on target CPU (they are invalid on x86_64) and the baseline-vs-builder-image tradeoff is a maintainer call. I offered that as a follow-up there.Publishing a multi-arch image is a separate step from making the build work, and that is the part that would actually close this issue. Happy to help if it is useful.
Cross-linking: #13957 tracks the remaining work to publish arm64 for Kubeflow Pipelines, with this issue listed as the external ML Metadata dependency. google/ml-metadata#247 is the open PR for the first half of it.
🏁 flagged for closure
Closing as no longer planned for the KFP 3.0 architecture. #13986 removed MLMD, including the metadata gRPC server, metadata-writer, and metadata envoy, from the active KFP deployment and image inventories on master. Publishing an ARM64 MLMD server is therefore no longer a prerequisite for KFP on ARM.
This does not mean the MLMD image was made multi-architecture or that supported ARM execution is complete. The remaining native ARM CI, installation/pipeline smoke tests, release validation, and supported-profile decisions remain tracked in #13957.
The release-2.18 branch retains MLMD and its ARM limitations; this closure does not establish ARM support for 2.18. Thank you to everyone who investigated and contributed upstream ARM build work.
Metadata
Metadata
Assignees
Type
Projects
- StatusShow more project fieldsNeeds triage

Description
Right now, the
gcr.io/tfx-oss-public/ml_metadata_store_servercontainer image is theonly image used in Kubeflow which is not published for both. This means that Kubeflow 1.8 still can not properly run on ARM clusters.amd64ANDarm64I have made a PR upstream in
google/ml-metadatato get the builds working for ARM64:ml_metadata_store_serverimage on ARM64 google/ml-metadata#188We need to work with the
ml-metadatateam to review/merge it and then set up a process to also push the ARM version of that image to GCR.EDIT: I was incorrect about this being the "only one" but I think this must be the only one that does not work at all under Rosetta emulation (but either way, we need to fix this one too as we also push native arm images for the others). I have raised a separate issue to track fixing the other images:
Love this idea? Give it a 👍.