Skip to content

wasm: use vqtbl1q_s8 for swizzle on A64V8 for SIMD and relaxed SIMD - #1450

Merged
mr-c merged 1 commit into
simd-everywhere:masterfrom
jawj:a64v8-swizzle-vqtbl1q_s8
Sep 14, 2026
Merged

mr-c merged 1 commit into
simd-everywhere:masterfrom
jawj:a64v8-swizzle-vqtbl1q_s8

Conversation

@jawj

@jawj jawj commented Sep 14, 2026

Copy link
Copy Markdown
Contributor

This patch has simde_wasm_i8x16_swizzle and simde_wasm_i8x16_relaxed_swizzle use vqtbl1q_s8 for SIMDE_ARM_NEON_A64V8_NATIVE.

I haven't done comprehensive benchmarking, but some arguments in favour:

  • It compiles to one instruction, not several (see below).
  • There's a precedent in SIMDe: _mm_shuffle_epi8 in ssse3.h already uses vqtbl1q_s8.
  • Results are identical: the native tests pass in both C and C++, and the instruction returns 0 for indices ≥ 16, as WASM's swizzle requires.
  • Real-world performance: some hex encoding/decoding functions with wasm2c on my M3 Pro Mac show 25 – 45% higher throughput (that's the whole functions: the swizzle bump is presumably larger).

In terms of what's emitted, here's Apple Clang 17.0.0 (clang-1700.6.4.2) with -O3 before —

       0:      	tbl.8b	v2, { v0 }, v1
       4:      	ext.16b	v1, v1, v1, #0x8
       8:      	tbl.8b	v0, { v0 }, v1
       c:      	mov.d	v2[1], v0[0]
      10:      	mov.16b	v0, v2
      14:      	ret

— and after —

       0:      	tbl.16b	v0, { v0 }, v1
       4:      	ret

@mr-c mr-c left a comment •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@mr-c
mr-c enabled auto-merge (rebase) September 14, 2026 13:49
@mr-c
mr-c merged commit cc45f1d into simd-everywhere:master Sep 14, 2026
145 of 146 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants