Skip to content

Instantly share code, notes, and snippets.

@raeq
raeq / README.md
Last active September 23, 2026 19:11
epubveri 0.16.0: four behaviour-preserving changes that cut CPU time to 39% (patch, benchmark, numbers)

epubveri 0.16.0: four changes that cut CPU time to 39%

epubveri-0.16.0-perf.patch applies to the v0.16.0 tag with git apply or patch -p1. It changes no finding: over a 2,798-book Apple Books library, every report matches 0.16.0's (message, severity, file and line) on every book. The crate's own tests pass, 711 of 711, and clippy reports no new warning.

Where the time went

The slowest book on the shelf, Python (521 XHTML files, 108 s of CPU), sampled with macOS

@raeq
raeq / README.md
Created September 16, 2026 19:35
epubcheck message ids: 21 of 315 use an underscore (repro)

epubcheck message ids: 21 of 315 use an underscore

repro.sh builds a minimal fixed-layout EPUB with two faults and runs epubcheck 5.3.0 on it. Output:

ERROR    HTM_056    Viewport metadata has no "height" dimension (both "width" and "height"
ERROR    RSC-012    Fragment identifier is not defined.
@raeq
raeq / weaponizing_unicode_uts39_overlap.py
Created September 4, 2026 11:26
raeq/disarm#977: weaponizing-unicode against the normative confusables.txt — corpus composition and detector coverage, reproduction
"""weaponizing-unicode against the normative confusables.txt: what the corpus is, and what each detector covers.
Reads the paper's `new_predicted_homoglyphs.txt` from $DISARM_META_WEAPONIZING or the meta-benchmark
cache, and UTS #39 `confusables.txt` from $DISARM_META_CONFUSABLES or `data/confusables.txt` in the
checkout. Needs disarm and confusable-homoglyphs importable. Expected output on disarm 0.15.0 at
fba73bc5 with confusables 17.0.0 and confusable-homoglyphs 3.3.1:
scored 7402; UTS #39 sources 586; single-code-point prototype 39; of those in another script 35
multi-code-point-prototype sources by script: [('ARABIC', 193), ('CANADIAN', 129), ('HANGUL', 108), ('LATIN', 20)]
scored n=7402 disarm 236 confusable-homoglyphs 362
@raeq
raeq / emoji_delimiter_ladder.py
Created September 4, 2026 07:52
raeq/disarm#972: every disarm surface on the Emoji Attack construction (emoji inside a word) — the meta-benchmark ladder, reproduction
"""Every disarm surface on the Emoji Attack construction: an Emoji_Presentation code point inside a word.
The ladder is the meta-benchmark's (`benchmarks/meta/suites/academic.py`, EmojiDelimiterSegmentation):
rejoined (emoji gone, word restored) > survives (still two runs) > widened (a name or gloss:
extra runs) / substituted (became letters) / destroyed (carrier lost). Reads emoji-data.txt from
$DISARM_META_EMOJI_DATA, else downloads it from unicode.org. Needs disarm importable; run from a
checkout for the profile names. Expected output on disarm 0.15.0 at 43e24b6b:
disarm 0.15.0; Emoji_Presentation code points: 1219
surface rejoined survives widened subst destroyed aa🔥bb ->
@raeq
raeq / tag_block_parcel_rescoring.py
Created September 4, 2026 06:41
raeq/disarm#970: mcp-tag-block-concealment re-scored under three parcel rules (T1 inverted; detection and transform parcelled apart) — reproduction
"""Re-score `mcp-tag-block-concealment` under three parcel rules, with the harness's own code.
Run from a disarm checkout that has `benchmarks/meta` (main at 2b2db6cf or later). The
measurements are quoted from the 2026-09-04 run of `python -m benchmarks.meta --run
--subject all` on disarm 0.15.0 at 2b2db6cf, so no run is needed. Expected output:
A. as scored today (detected_t1_plain: flagged = good)
one parcel disarm rank 8 of 12 (parcel +0.026); top +0.760; confusable-homoglyphs rank 7 (+0.707)
B. detected_t1_plain inverted (not flagged = good)
one parcel disarm rank 7 of 12 (parcel +0.733); top +0.760; confusable-homoglyphs rank 8 (-0.707)
@raeq
raeq / meta_composed_strip_zalgo_repro.py
Created September 3, 2026 19:07
raeq/disarm#958: benchmarks/meta composes three disarm subjects with strip_zalgo=0, which strips every diacritic — reproduction
"""benchmarks/meta composes three disarm pipelines with strip_zalgo=0, which strips every diacritic.
disarm documents max_marks=0 as "strip all combining marks" (docs/api/transforms.md:
strip_zalgo("café", max_marks=0) == "cafe"), and None as off. The composed subjects in
benchmarks/meta/subjects.py pass 0. Run from a disarm checkout that has benchmarks/meta,
with disarm importable. Expected output on disarm 0.15.0 at da174c83:
disarm 0.15.0
TextPipeline(strip_zalgo=None) steps=[] 'Čeština, naïve café'
TextPipeline(strip_zalgo=0 ) steps=[('strip_zalgo', '0')] 'Cestina, naive cafe'
@raeq
raeq / is_confusable_ascii_repro.py
Created September 3, 2026 19:06
raeq/disarm#957: is_confusable fires on printable ASCII (", |, backtick are TR39 sources) — reproduction
"""is_confusable fires on printable ASCII: `"`, `|` and backtick are TR39 sources.
Run from a disarm checkout (the four prose files are read relative to it) with disarm
importable. Expected output on disarm 0.15.0 at da174c83 (tables Unicode 17.0.0,
confusables 17.0.0):
disarm 0.15.0
is_confusable('"') True
is_confusable('|') True
is_confusable('`') True
@raeq
raeq / bad_characters_by_class.py
Created September 3, 2026 11:09
disarm #937: Boucher's four classes scored one at a time (canonicalize recovers 14 of 4,800 reorderings, 0 of 4,820 deletions) and a cell-aware cursor model that recovers 4,820 of 4,820
#!/usr/bin/env python3
"""disarm #937: Boucher et al.'s four classes scored one at a time, and a cursor model
for the deletion class.
The meta-benchmark scores the released corpus as one population. Split by the
experiment key, the aggregate is the sum of two classes near 100%, one near 0%, and
one at 0% — and the two near zero are the two whose *rendering* is the clean string.
Corpus: https://raw.githubusercontent.com/nickboucher/imperceptible/main/results/adversarial-examples.json
(MIT). Nested experiment -> budget -> row -> {"adv_example", "input", ...}. Pass the path
@raeq
raeq / deletion_recovery.py
Last active September 3, 2026 11:09
disarm: #934 shipped detection of the deletion class; no surface applies the erase, so 0 of 57 surface/form pairs recover and the smuggled character reaches the consumer
#!/usr/bin/env python3
"""#739's detection half shipped in #934; the recovery half did not (revised 2026-09-03).
Boucher et al.'s deletion class hides a character from a human reader by
following it with an erasing control. disarm now *detects* every form of it and
*names* the semantics — "erases the preceding 'a'" — and no surface applies
them, so the smuggled character reaches whatever reads the output.
pip install disarm==0.15.0
python deletion_recovery.py
@raeq
raeq / coverage_cost_plot.py
Created September 2, 2026 19:14
Reproduce every number on the disarm meta-benchmark coverage/cost plot (13 benchmarks, 13 subjects, 0.15.0)
#!/usr/bin/env python3
"""Reproduce every number on the disarm meta-benchmark coverage/cost plot.
The plot shows two of thirteen axes. This script recomputes both of them, the
composite, the bootstrap intervals and the Pareto frontier from a run of the
harness, and prints them in the order the page's table lists them.
git clone https://github.com/raeq/disarm && cd disarm
git checkout bench/meta-harness
python -m venv .venv && .venv/bin/pip install -e . maturin