Inspecting AArch64 disassembly of zlib-ng's level-9 deflate path showed that the rolling-hash insert loop emitted a redundant store of the running hash to memory on every byte processed. The cause was visible in the C source and the fix is small and self-contained.
| /* benchmark_dist1.cc -- compare strategies for the dist=1 path in CHUNKMEMSET */ | |
| #include <benchmark/benchmark.h> | |
| extern "C" { | |
| # include "zbuild.h" | |
| # include "zutil.h" | |
| } | |
| #if defined(__ARM_NEON) || defined(__ARM_NEON__) |
| diff --git a/deflate_quick.c b/deflate_quick.c | |
| index 6b84388e..41086525 100644 | |
| --- a/deflate_quick.c | |
| +++ b/deflate_quick.c | |
| @@ -49,13 +49,19 @@ Z_INTERNAL block_state deflate_quick(deflate_state *s, int flush) { | |
| unsigned char *window; | |
| unsigned last = (flush == Z_FINISH) ? 1 : 0; | |
| + /* Hold strstart and lookahead in registers across the inner loop. | |
| + * Sync to s before fill_window (which mutates them) and quick_*_block |
PR: zlib-ng/zlib-ng#2281 — [CI] Cache Ubuntu .deb packages to speed up installing dependencies.
Follow-up to https://gist.github.com/nmoinvaz/978f248ea7d528c954e3d52fb8dc99c0 — measuring the v2 design after the author refactored into composite actions and adopted the matrix.packages gate.
v2 is a clean win across the board. Total apt-install time across Ubuntu jobs is now −71% vs. develop, up from −47% in v1. Default-only jobs no longer pay cache overhead, and warm cache hits skip the write-back step entirely.
PR: zlib-ng/zlib-ng#2281 — [CI] Cache Ubuntu .deb packages to speed up installing dependencies.
The cache works: it cuts total apt-install time across Ubuntu jobs roughly in half. But cmake.yml enables the cache on every Ubuntu job, including the 23 jobs that only install the default libgtest-dev libbenchmark-dev package set. For those, cache-action overhead exceeds the savings. configure.yml already gates on matrix.packages; mirroring that gate in cmake.yml is a one-line fix.
Validity-check benchmark for the zng_check_lens(lens, codes)
function proposed in the PR #2267 discussion. All three variants
scan lens[0..codes-1] and return -1 if any entry exceeds MAX_BITS
(15). Input is all-valid (random values in [0, 15]) so the worst
case — a full scan with no early exit — is measured.
Variants:
Investigation spun out of the PR #2267 discussion on zlib-ng: can the
SIMD paths in count_lengths (inftrees.c) be replaced with a SWAR
implementation using zng_memread_8?
Mirror the pair-interleaved 8-bit-lane structure of the active SIMD
path. Two pairs of uint64_t accumulators (s1_lo/s1_hi,
| /* SIMD validity check for a Huffman code-length buffer. | |
| * | |
| * Returns 0 if every entry in lens[0..codes-1] is <= MAX_BITS, | |
| * otherwise returns -1. Called from zng_inflate_table before | |
| * count_lengths to guard against the out-of-bounds read of one[] | |
| * described in zlib-ng issue #2266. | |
| * | |
| * Main loop scans 8 uint16_t per iteration via 128-bit vector | |
| * compares; a scalar tail handles the remaining 0-7 entries. No | |
| * assumptions about buffer padding or caller identity. |
| #!/bin/bash | |
| # Toy-test cross-compile: compile a minimal C file (see | |
| # zlib-ng-pr2261-visibility-test.c) with every GCC cross compiler we can | |
| # reasonably get, twice per architecture (visibility("hidden") and | |
| # visibility("internal")), and diff the assembly output — ignoring the | |
| # .hidden/.internal pseudo-op so any real codegen difference surfaces. | |
| # | |
| # Run via Docker: | |
| # docker run --rm -v /path/to/scripts:/work -w /work \ | |
| # debian:trixie-slim bash /work/zlib-ng-pr2261-visibility-runall.sh |
| /* Minimal reproducer: MSVC v142 _mm_set_epi64x miscompile, 32-bit x86. | |
| * | |
| * cl /O2 /arch:SSE2 repro_v142.c && repro_v142.exe | |
| * | |
| * v142 Win32: FAIL (wrong result due to register corruption) | |
| * v143+ Win32: PASS | |
| * | |
| * https://developercommunity.visualstudio.com/t/10853479 | |
| */ | |
| #include <stdio.h> |