Skip to content

Instantly share code, notes, and snippets.

@pmarreck
Created September 28, 2026 12:03
Show Gist options
  • Select an option

  • Save pmarreck/c9659d1ac71665065fb20f7960208f2c to your computer and use it in GitHub Desktop.

Select an option

Save pmarreck/c9659d1ac71665065fb20f7960208f2c to your computer and use it in GitHub Desktop.
dna-repeats now faster, mor accurate and more memory-efficient than every other known tool in the DNA-hunting space

I have to correct something I told you last night. I said dna-repeats was about 7x faster than MinCED, but that was on all 64 cores. Measured on one core, as MinCED and PILER-CR run, it was 2.4 to 3.7x slower than both. That's now fixed, and dna-repeats wins on every axis even single-threaded. I've emailed you the correction.

Why it was slow: 98% of the time went to the regex step that finds where repeats can start. I replaced that step with a plain Zig scan (src/kmer_scan.zig) that gives exactly the same positions. Output is byte-identical to the old version on all 144 real runs I compared. Held-out #3 went from 59 s to 4.13 s on one core.

Held-out #3 (30 genomes, 73 known arrays), every tool on one core:

Tool Recall Precision Time Peak memory Exact repeat
dna-repeats 68/73 95.9% 4.13 s 14.0 MB 83.3%
MinCED 67/73 94.4% 14.74 s 461.2 MB 76.1%
PILER-CR 63/73 94.7% 24.23 s 24.0 MB 73.2%

With 64 threads dna-repeats takes 0.55 s. One place PILER-CR still leads: on held-out #3 its reported repeats are slightly closer on average (edit distance 0.45 vs 0.51), even though ours are exactly right more often.

The report is published: https://claude.ai/artifact/YDA9CSb9S3wHVikR2ge4Te. It's private, so nobody else can open it until you share it.

Also done tonight:

  • --sensitive: starts seeds at 14 bases instead of 18. On held-out #3 it scores identically to the default, and it finds 16 of 20 arrays at 8% mutation versus 12 (MinCED finds 8).
  • Scoreboard: a crashing tool used to be scored as "found nothing"; now the run fails loudly, and each log row records tool paths, an input checksum and whether the tree was dirty. I reran every set and found no hidden failures.
  • Repeat accuracy: this is now a scored axis. Arrays report the consensus repeat instead of one copy, which raised exact matches with the database from 123 to 143 of 164.
  • Array caller tests: all 22 mutants from a mutation sweep are now caught (18 slipped through before), and I removed a filter that never fired.
  • Output bug: several runs writing to one redirected file used to overwrite each other. Fixed, with regression tests.

Everything is committed and pushed (latest 9471d67). The repo is still private, and I won't publish anything until you've looked. It's 2:35 AM, so this can all wait until you've slept.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment