I’ll run a memory-only version, keeping the benchmark separate from the criteria used to explain it.
One limitation matters: I can propose a benchmark based on remembered usage, but cannot establish overwhelming popular consensus empirically without sources. I’ll distinguish a promising result from a result that passes your full stopping rule.