Freebie for the VERIFY keyword (reel-79B, "It read a page that doesn't exist"). Deliver as a public GitHub Gist. By the end you have a 30-line script and a Claude skill that refuse to pass you a number that isn't actually on the page.
I pointed an AI scraper at a product page and asked for the app name, the developer, the pricing tiers and the rating. It gave me all of it, formatted, confident:
{ "app_name": "WorkflowMail", "developer": "Shopify",
"pricing_plans": [ { "name": "Basic Plan", "price_usd_per_month": 29 },
{ "name": "Pro Plan", "price_usd_per_month": 59 } ],
"rating": 4.7, "review_count": 1200 }The page does not exist. It never did. The response carried the answer in one field and
"statusCode": 404 in another, and I read the first one.
I ran it twice more to be sure it wasn't a fluke. Same URL, same 404, three different answers:
| run | plans it invented | rating | reviews |
|---|---|---|---|
| 1 | $10 / $25 / $50 | 4.7 | 150 |
| 2 | $29 / $59 | 4.7 | 1,200 |
| 3 | $0 / $29 / $79 | 4.5 | 200 |
Three runs, three price lists, one page that isn't there. The model wasn't lying about the page — it was filling in what a page like that usually says.
Why it happens: a 404 from a big site is not an empty response. That one returned 79,877 bytes — navigation, footer, help links, other apps. Plenty of text to pattern-match against, none of it the thing you asked about. So the extraction step does what extraction steps do, and the structure it produces looks exactly like a real answer.
Never read a number out of a scrape without checking that the page returned 200 and that the number appears verbatim in the page's own text.
Two conditions, both cheap. The first catches the dead page. The second catches the live page that simply doesn't say what the model claims — which is the sneakier half.
Save as checksrc somewhere on your PATH, then chmod +x checksrc.
#!/usr/bin/env bash
# checksrc — does this page exist, and does it actually contain the value?
# usage: checksrc <url> [value ...]
set -uo pipefail
url="$1"; shift
body=$(mktemp)
code=$(curl -sS -L -o "$body" -w '%{http_code}' \
-A 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7)' "$url")
bytes=$(wc -c < "$body" | tr -d ' ')
echo "HTTP $code · ${bytes} bytes · $url"
if [ "$code" != "200" ]; then
echo "STOP — page did not return 200. Anything extracted from it is invented."
rm -f "$body"; exit 1
fi
title=$(tr '\n' ' ' < "$body" | grep -o -i '<title>[^<]*' | head -1 | cut -c8-)
echo "title: ${title:-<none>}"
case "$(printf '%s' "$title" | tr 'A-Z' 'a-z')" in
*"not found"*|*"doesn't exist"*|*"does not exist"*|*"404"*|*"oh no"*|*"error"*)
echo "STOP — 200 status but the title reads like an error page." ; rm -f "$body"; exit 1 ;;
esac
text=$(sed -e 's/<[^>]*>/ /g' "$body" | tr -s ' \n' ' ')
fail=0
for v in "$@"; do
if printf '%s' "$text" | grep -qiF -- "$v"; then echo " ok $v"
else echo " MISSING $v"; fail=1; fi
done
rm -f "$body"
[ "$fail" = 0 ] || { echo "STOP — at least one value is not in the page text."; exit 1; }
echo "OK — page exists and every value appears in its text."The title check is there because plenty of sites answer 200 with an error page. A status code alone is not enough on those.
Dead page, claimed price:
$ checksrc "https://apps.shopify.com/workflowmail" "29"
HTTP 404 · 79877 bytes · https://apps.shopify.com/workflowmail
STOP — page did not return 200. Anything extracted from it is invented.
Real page, real values:
$ checksrc "https://github.com/heygen-com/hyperframes" "Apache-2.0" "Write HTML"
HTTP 200 · 479262 bytes
title: GitHub - heygen-com/hyperframes: Write HTML. Render video. Built for agents. · GitHub
ok Apache-2.0
ok Write HTML
OK — page exists and every value appears in its text.
Real page, invented values:
$ checksrc "https://github.com/heygen-com/hyperframes" "4.7 stars" '$49/month'
HTTP 200 · 479263 bytes
MISSING 4.7 stars
MISSING $49/month
STOP — at least one value is not in the page text.
Exit code is 0 only in the middle case, so you can chain it: checksrc "$url" "$price" && ./import.sh.
Save as ~/.claude/skills/verify-scrape/SKILL.md and restart Claude Code.
---
name: verify-scrape
description: Use whenever a scrape, crawl, fetch or extraction produced values (prices, ratings, counts, names, dates) that are about to be reported, stored or acted on. Verifies the source returned 200 and that each value appears verbatim in the page text before the values are used.
---
# Verify a scrape before trusting it
An extraction step returns a well-formed object whether or not the page exists. Structure is
not evidence. Before any scraped value is reported to the user, written to a file or database,
or used in a calculation, run this gate.
## The gate
1. **Status.** Confirm the response carried HTTP 200. If the scraping tool reports a status
code field, read it — it is often in the response you already have, next to the answer.
Anything other than 200 means STOP and say the page did not resolve. Do not report the
extracted values, not even as a guess, and do not retry with a different phrasing of the
same URL.
2. **Error page at 200.** Check the page title. If it contains "not found", "doesn't exist",
"404", "oh no" or "error", treat it as case 1.
3. **Verbatim presence.** For each value you are about to report, confirm the value appears in
the page's own text. Numbers must match as written, including separators. A value you
cannot find is not a value you may report.
4. **Report what you checked.** State the status code and which values you found in the source.
If a value failed, name it and leave it out rather than substituting a plausible one.
## Tooling
`checksrc <url> [value ...]` runs all of this in one call and exits non-zero on any failure.
Fall back to `curl -sS -L -o /tmp/p -w '%{http_code}'` plus `grep -F` if it is not installed.
## What this is not
This does not check whether the page is *correct* — only that it exists and says what you are
about to repeat. Truth about the world still needs a second source.Same failure, different dress:
- A product that was delisted. The URL used to work, so it is in your spreadsheet. It now 404s, and the scraper refills it with the category average.
- A geo-blocked or login-walled page. You get a 200 and a consent wall. The extractor reads the wall and invents the rest.
- A renamed API endpoint. Returns 200 with
{"error": "..."}. Structure parses fine.
The gate above catches all three, which is why it is worth the 30 lines.
Evidence: three live scrapes of the same non-existent URL on 10 and 11 Sep 2026, each response
carrying statusCode: 404 alongside a complete invented answer. Script tested on 22 Sep 2026
against one 404 page, one live page with real values and the same live page with invented
values — stop, pass, stop, exit codes 1 / 0 / 1.