Skip to content

Instantly share code, notes, and snippets.

View nmoinvaz's full-sized avatar

Nathan Moin Vaziri nmoinvaz

  • Los Angeles, California
View GitHub Profile
@nmoinvaz
nmoinvaz / benchmark_dist1.cc
Last active May 10, 2026 06:51
zlib-ng PR #2286: microbenchmark of memset replacements for the dist=1 path in CHUNKMEMSET
/* benchmark_dist1.cc -- compare strategies for the dist=1 path in CHUNKMEMSET */
#include <benchmark/benchmark.h>
extern "C" {
# include "zbuild.h"
# include "zutil.h"
}
#if defined(__ARM_NEON) || defined(__ARM_NEON__)
@nmoinvaz
nmoinvaz / deflate_quick.patch
Last active May 8, 2026 07:55
zlib-ng: hoist strstart and lookahead in deflate_quick (~12% level 1)
diff --git a/deflate_quick.c b/deflate_quick.c
index 6b84388e..41086525 100644
--- a/deflate_quick.c
+++ b/deflate_quick.c
@@ -49,13 +49,19 @@ Z_INTERNAL block_state deflate_quick(deflate_state *s, int flush) {
unsigned char *window;
unsigned last = (flush == Z_FINISH) ? 1 : 0;
+ /* Hold strstart and lookahead in registers across the inner loop.
+ * Sync to s before fill_window (which mutates them) and quick_*_block
@nmoinvaz
nmoinvaz / zlib-ng-insert-string-roll-spill.md
Created May 8, 2026 06:21
zlib-ng: eliminate s->ins_h spill in insert_string_roll inner loop

zlib-ng: eliminate s->ins_h spill in insert_string_roll inner loop

Background

Inspecting AArch64 disassembly of zlib-ng's level-9 deflate path showed that the rolling-hash insert loop emitted a redundant store of the running hash to memory on every byte processed. The cause was visible in the C source and the fix is small and self-contained.

Root cause

@nmoinvaz
nmoinvaz / zlib-ng-pr2281-cache-v2-results.md
Last active May 7, 2026 23:58
zlib-ng PR #2281 — apt-cache v2 results: composite actions + matrix.packages gate

zlib-ng PR #2281 — apt-cache v2 results

PR: zlib-ng/zlib-ng#2281 — [CI] Cache Ubuntu .deb packages to speed up installing dependencies.

Follow-up to https://gist.github.com/nmoinvaz/978f248ea7d528c954e3d52fb8dc99c0 — measuring the v2 design after the author refactored into composite actions and adopted the matrix.packages gate.

TL;DR

v2 is a clean win across the board. Total apt-install time across Ubuntu jobs is now −71% vs. develop, up from −47% in v1. Default-only jobs no longer pay cache overhead, and warm cache hits skip the write-back step entirely.

@nmoinvaz
nmoinvaz / zlib-ng-pr2281-cache-default-packages.md
Created May 7, 2026 18:47
zlib-ng PR #2281 — apt-cache analysis: skip cache on default-package Ubuntu jobs

zlib-ng PR #2281 — apt-cache analysis and recommendation

PR: zlib-ng/zlib-ng#2281 — [CI] Cache Ubuntu .deb packages to speed up installing dependencies.

TL;DR

The cache works: it cuts total apt-install time across Ubuntu jobs roughly in half. But cmake.yml enables the cache on every Ubuntu job, including the 23 jobs that only install the default libgtest-dev libbenchmark-dev package set. For those, cache-action overhead exceeds the savings. configure.yml already gates on matrix.packages; mirroring that gate in cmake.yml is a one-line fix.

Methodology

@nmoinvaz
nmoinvaz / zlib-ng-check-lens-bench-results.md
Created April 23, 2026 00:58
zlib-ng zng_check_lens: SIMD vs SWAR vs scalar benchmark (spun out of PR #2267)

zng_check_lens: SIMD vs SWAR vs scalar

Validity-check benchmark for the zng_check_lens(lens, codes) function proposed in the PR #2267 discussion. All three variants scan lens[0..codes-1] and return -1 if any entry exceeds MAX_BITS (15). Input is all-valid (random values in [0, 15]) so the worst case — a full scan with no early exit — is measured.

Variants:

@nmoinvaz
nmoinvaz / zlib-ng-count-lengths-swar-results.md
Last active April 23, 2026 01:18
zlib-ng count_lengths: SWAR vs SIMD benchmark investigation (spun out of PR #2267)

count_lengths: SWAR vs SIMD

Investigation spun out of the PR #2267 discussion on zlib-ng: can the SIMD paths in count_lengths (inftrees.c) be replaced with a SWAR implementation using zng_memread_8?

SWAR design

Mirror the pair-interleaved 8-bit-lane structure of the active SIMD path. Two pairs of uint64_t accumulators (s1_lo/s1_hi,

@nmoinvaz
nmoinvaz / zlib-ng-pr2267-check-lens.c
Last active April 23, 2026 00:46
zlib-ng PR #2267: SIMD lens[] validity prescan with zero-init invariant
/* SIMD validity check for a Huffman code-length buffer.
*
* Returns 0 if every entry in lens[0..codes-1] is <= MAX_BITS,
* otherwise returns -1. Called from zng_inflate_table before
* count_lengths to guard against the out-of-bounds read of one[]
* described in zlib-ng issue #2266.
*
* Main loop scans 8 uint16_t per iteration via 128-bit vector
* compares; a scalar tail handles the remaining 0-7 entries. No
* assumptions about buffer padding or caller identity.
@nmoinvaz
nmoinvaz / zlib-ng-pr2261-visibility-runall.sh
Last active April 18, 2026 04:00
zlib-ng PR #2261 — visibility("internal") vs visibility("hidden") for Z_INTERNAL. Four-way audit: GCC source, linker source (mold/lld/BFD/gold/ld.so), community prior art (LLVM issue #9555 aliased internal→hidden in 2015), and empirical cross-builds (301 zlib-ng objects across 8 arches, 0 disassembly differences). Consensus: internal is dead-end…
#!/bin/bash
# Toy-test cross-compile: compile a minimal C file (see
# zlib-ng-pr2261-visibility-test.c) with every GCC cross compiler we can
# reasonably get, twice per architecture (visibility("hidden") and
# visibility("internal")), and diff the assembly output — ignoring the
# .hidden/.internal pseudo-op so any real codegen difference surfaces.
#
# Run via Docker:
# docker run --rm -v /path/to/scripts:/work -w /work \
# debian:trixie-slim bash /work/zlib-ng-pr2261-visibility-runall.sh
@nmoinvaz
nmoinvaz / repro_v142.c
Created April 17, 2026 20:14
zlib-ng: Minimal reproducer for MSVC v142 _mm_set_epi64x miscompile on 32-bit x86
/* Minimal reproducer: MSVC v142 _mm_set_epi64x miscompile, 32-bit x86.
*
* cl /O2 /arch:SSE2 repro_v142.c && repro_v142.exe
*
* v142 Win32: FAIL (wrong result due to register corruption)
* v143+ Win32: PASS
*
* https://developercommunity.visualstudio.com/t/10853479
*/
#include <stdio.h>