Skip to content

build and publish ml_metadata_store_server container image for ARM64 #10308

Description

@thesuperzapper

Description

Right now, the gcr.io/tfx-oss-public/ml_metadata_store_server container image is the only image used in Kubeflow which is not published for both amd64 AND arm64. This means that Kubeflow 1.8 still can not properly run on ARM clusters.

I have made a PR upstream in google/ml-metadata to get the builds working for ARM64:

We need to work with the ml-metadata team to review/merge it and then set up a process to also push the ARM version of that image to GCR.

EDIT: I was incorrect about this being the "only one" but I think this must be the only one that does not work at all under Rosetta emulation (but either way, we need to fix this one too as we also push native arm images for the others). I have raised a separate issue to track fixing the other images:


Love this idea? Give it a 👍.

Activity

  1. thesuperzapper commented on Dec 12, 2023

    @thesuperzapper
    MemberAuthor
  2. thesuperzapper commented on Dec 13, 2023

    @thesuperzapper
    MemberAuthor

    For those who want to test, I have made a forked repo in the deployKF org with the ARM versions of the gcr.io/tfx-oss-public/ml_metadata_store_server image. You can test a patched version of ml-metdata version 1.14.0 by using the following container:

    Note, building under emulation on GitHub actions took about 5 hours:

  3. xixici commented on Dec 13, 2023

    @xixici

    Great. I pull this image and run it correctly. Then, I am finding gcr.io/ml-pipeline/metadata-writer and gcr.io/ml-pipeline/metadata-envoy with ARM version.

  4. thesuperzapper commented on Dec 13, 2023

    @thesuperzapper
    MemberAuthor

    @xixici can you confirm what you are saying?

    Because gcr.io/ml-pipeline/metadata-writer:2.0.5 and gcr.io/ml-pipeline/metadata-envoy:2.0.5 (and all other versions) are only published for ADM64.

    I assume you mean that they work via Rosetta Emulation on a MacBook?

  5. github-actions commented on Mar 13, 2024

    @github-actions

    This issue has been automatically marked as stale because it has not had recent activity. It will be closed if no further activity occurs. Thank you for your contributions.

  6. added
    lifecycle/staleThe issue / pull request is stale, any activities remove this label.
    on Mar 13, 2024
  7. github-actions commented on Apr 3, 2024

    @github-actions

    This issue has been automatically closed because it has not had recent activity. Please comment "/reopen" to reopen it.

  8. thesuperzapper commented on Apr 5, 2024

    @thesuperzapper
    MemberAuthor

    /reopen

  9. google-oss-prow commented on Apr 5, 2024

    @google-oss-prow

    @thesuperzapper: Reopened this issue.

    Details

    In response to this:

    /reopen

    Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes/test-infra repository.

  10. thesuperzapper commented on Apr 5, 2024

    @thesuperzapper
    MemberAuthor

    Clearly, this is going to take some time, so I will prevent the bot from closing it.

    /lifecycle frozen

  11. added and removed
    lifecycle/staleThe issue / pull request is stale, any activities remove this label.
    on Apr 5, 2024
  12. MouseSun846 commented on Jul 18, 2024

    @MouseSun846

    I have successfully completed metadata envoy 2.0.5 and built it in an ARM environment

    The following are the construction steps:

    docker pull --platform linux/arm64 envoyproxy/envoy:v1.16.0

    In the directory of pipelines/third_party/metadata_envoy

    1、modify Dockerfile and config proxy info

    FROM envoyproxy/envoy:v1.16.0
    
    RUN apt-get -o Acquire::http::proxy="http://proxy:port" update -y && \
      apt-get -o Acquire::http::proxy="http://proxy:port" install --no-install-recommends -y -q gettext openssl
    
    COPY third_party/metadata_envoy/envoy.yaml /etc/envoy.yaml
    
    # Copy license files.
    #RUN mkdir -p /third_party
    COPY third_party/metadata_envoy/license.txt /third_party/license.txt
    
    ENTRYPOINT ["/usr/local/bin/envoy", "-c"]
    CMD ["/etc/envoy.yaml"]
    

    2、modify envoy.yaml

    admin:
      access_log_path: /tmp/admin_access.log
      address:
        socket_address: { address: 0.0.0.0, port_value: 9901 }
    
    static_resources:
      listeners:
        - name: listener_0
          address:
            socket_address: { address: 0.0.0.0, port_value: 9090 }
          filter_chains:
            - filters:
                - name: envoy.http_connection_manager
                  typed_config:
                    "@type": type.googleapis.com/envoy.config.filter.network.http_connection_manager.v2.HttpConnectionManager
                    codec_type: auto
                    stat_prefix: ingress_http
                    route_config:
                      name: local_route
                      virtual_hosts:
                        - name: local_service
                          domains: ["*"]
                          routes:
                            - match: { prefix: "/" }
                              route:
                                cluster: metadata-cluster
                                max_grpc_timeout: 0s
                          cors:
                            allow_origin_string_match:
                              - exact: "*"
                            allow_methods: GET, PUT, DELETE, POST, OPTIONS
                            allow_headers: keep-alive,user-agent,cache-control,content-type,content-transfer-encoding,custom-header-1,x-accept-content-transfer-encoding,x-accept-response-streaming,x-user-agent,x-grpc-web,grpc-timeout
                            max_age: "1728000"
                            expose_headers: custom-header-1,grpc-status,grpc-message
                    http_filters:
                      - name: envoy.grpc_web
                      - name: envoy.cors
                      - name: envoy.router
      clusters:
        - name: metadata-cluster
          connect_timeout: 30.0s
          type: logical_dns
          http2_protocol_options: {}
          lb_policy: round_robin
          load_assignment:
            cluster_name: metadata-cluster
            endpoints:
            - lb_endpoints:
              - endpoint:
                  address:
                    socket_address:
                      address: metadata-grpc-service
                      port_value: 8080   
    

    3、finally,build image
    In the directory of pipelines/backend/Makefile

    .PHONY: metadata_envoy
    metadata_envoy:
    	cd $(MOD_ROOT) && docker build -t registry.cnbita.com:5000/kubeflow-pipelines/metadata-envoy:2.0.5-arm  -f third_party/metadata_envoy/Dockerfile .
    

    run make metadata_envoy

  13. MouseSun846 commented on Jul 18, 2024

    @MouseSun846

    Clearly, this is going to take some time, so I will prevent the bot from closing it.

    /lifecycle frozen

    #10308 (comment)

  14. MouseSun846 commented on Jul 18, 2024

    @MouseSun846

    It seems that kubeflow v2 does not use the metadata writer component. My cluster has not installed it, but I can still use the pipeline normally

    image

  15. juliusvonkohout commented on Aug 20, 2025

    @juliusvonkohout
    Member
  16. hsinhoyeh commented on Aug 3, 2026

    @hsinhoyeh
    Contributor

    ml_metadata_store_server does build on arm64. I got it working and opened a PR for the first piece: google/ml-metadata#247.

    I mention this because the discussion here has assumed the blocker is deep, and it turned out to be a hardcoded build config rather than a portability problem in the C++ code.

    What actually blocks it

    Building //ml_metadata/metadata_store:metadata_store_server on aarch64 hits exactly two things.

    1. The Dockerfile hardcodes the x86_64 Bazel installer. ml_metadata/tools/docker_server/Dockerfile fetches bazel-$BAZEL_VERSION-installer-linux-x86_64.sh, so the image can only be built on x86_64. Bazel does publish linux-arm64 (for both 5.3.0 and master's 7.7.0), just as a plain binary rather than an installer script. That is what google/ml-metadata#247 fixes.

    2. Abseil emits an ARMv8.3 instruction the assembler rejects at the default baseline. On current master, absl/debugging/stacktrace.cc emits xpaclri (pointer authentication):

    external/abseil-cpp/absl/debugging/stacktrace.cc [for tool] failed
    /tmp/ccRV18O3.s:166: Error: selected processor does not support `xpaclri'
    

    Adding --copt=-march=armv8.3-a --host_copt=-march=armv8.3-a to the bazel build invocation clears it. Note --host_copt is genuinely required alongside --copt — the failing target compiles [for tool], i.e. in the exec configuration, which --copt does not affect. With --copt alone the error is byte-identical, which makes it look like the flag did nothing.

    -march=armv8-a+pauth would be preferable to raising the whole baseline, since xpaclri is HINT-space and a no-op on older cores, whereas armv8.3-a makes the binary require ARMv8.3 hardware (Graviton2 is ARMv8.2). But GCC 9 on the ubuntu:20.04 builder rejects it (invalid feature modifier 'pauth'), so a newer builder base image would be the cleaner fix.

    Result

    With those two changes, master (be943b8) builds clean on aarch64 with zero errors:

    arch=arm64, 131 MB
    $ docker run --entrypoint /bin/metadata_store_server ... --help
    metadata_store_server: Warning: SetUsageMessage() never called
      Flags from external/com_github_gflags_gflags/src/gflags.cc: ...
    

    Built on an NVIDIA GB10 (aarch64, Ubuntu 24.04, Docker 28.5.1).

    It unblocks Kubeflow Pipelines on ARM

    I deployed the resulting image into a Kubeflow Pipelines 2.16.1 install on an arm64 kind cluster (alongside the KFP images rebuilt for arm64 — details in #10309). metadata-grpc-deployment goes to 1/1 Running, and ml-pipeline-persistenceagent and ml-pipeline-scheduledworkflow, which had been crash-looping against the unavailable metadata service, recover with it.

    Before the rebuild, the stock image on that same node gave the expected:

    exec /bin/metadata_store_server: exec format error
    

    One note on older release branches

    On v1.14.0 a third change was also needed, which does not apply to master. ml_metadata/postgresql.BUILD generates pg_config.h from a fixed list of lines captured from a configure run on x86_64 — PG_VERSION_STR still records "PostgreSQL 12.1 on x86_64-apple-darwin19.2.0" — asserting x86-only CPU capabilities with no select() on architecture anywhere. With HAVE__GET_CPUID defined, src/port/pg_bitutils.c includes <cpuid.h>, which only exists on x86:

    external/postgresql/src/port/pg_bitutils.c:16:10: fatal error: cpuid.h: No such file or directory
    

    On master HAVE__GET_CPUID is already /* #undef */, and my master build confirms the remaining unconditional #define HAVE_X86_64_POPCNTQ 1 does not break aarch64 on its own — so no pg_config.h change is needed there. Recording it only because it still affects release branches, and because the underlying pattern (a configure snapshot hardcoded for one CPU) could resurface.

    Scope

    google/ml-metadata#247 covers only the Dockerfile change, since the -march flags need to be conditional on target CPU (they are invalid on x86_64) and the baseline-vs-builder-image tradeoff is a maintainer call. I offered that as a follow-up there.

    Publishing a multi-arch image is a separate step from making the build work, and that is the part that would actually close this issue. Happy to help if it is useful.

  17. hsinhoyeh commented on Aug 3, 2026

    @hsinhoyeh
    Contributor

    Cross-linking: #13957 tracks the remaining work to publish arm64 for Kubeflow Pipelines, with this issue listed as the external ML Metadata dependency. google/ml-metadata#247 is the open PR for the first half of it.

  18. rvolchek commented on Aug 15, 2026

    @rvolchek

    🏁 flagged for closure

  19. jeffspahr commented on Sep 15, 2026

    @jeffspahr
    Collaborator

    Closing as no longer planned for the KFP 3.0 architecture. #13986 removed MLMD, including the metadata gRPC server, metadata-writer, and metadata envoy, from the active KFP deployment and image inventories on master. Publishing an ARM64 MLMD server is therefore no longer a prerequisite for KFP on ARM.

    This does not mean the MLMD image was made multi-architecture or that supported ARM execution is complete. The remaining native ARM CI, installation/pipeline smoke tests, release validation, and supported-profile decisions remain tracked in #13957.

    The release-2.18 branch retains MLMD and its ARM limitations; this closure does not establish ARM support for 2.18. Thank you to everyone who investigated and contributed upstream ARM build work.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions