Skip to content

Instantly share code, notes, and snippets.

@conradcaffier03
Created September 24, 2026 22:50
Show Gist options
  • Select an option

  • Save conradcaffier03/d7d6788c99dfb93176b5ad7a7b0f6076 to your computer and use it in GitHub Desktop.

Select an option

Save conradcaffier03/d7d6788c99dfb93176b5ad7a7b0f6076 to your computer and use it in GitHub Desktop.
The check that stops an AI scraper from handing you made-up numbers — 30 lines of bash + a Claude skill (free, by @buildwith.conrad)

✅ VERIFY — the check that stops a scraper from handing you made-up numbers

Freebie for the VERIFY keyword (reel-79B, "It read a page that doesn't exist"). Deliver as a public GitHub Gist. By the end you have a 30-line script and a Claude skill that refuse to pass you a number that isn't actually on the page.


What happened to me

I pointed an AI scraper at a product page and asked for the app name, the developer, the pricing tiers and the rating. It gave me all of it, formatted, confident:

{ "app_name": "WorkflowMail", "developer": "Shopify",
  "pricing_plans": [ { "name": "Basic Plan", "price_usd_per_month": 29 },
                     { "name": "Pro Plan",   "price_usd_per_month": 59 } ],
  "rating": 4.7, "review_count": 1200 }

The page does not exist. It never did. The response carried the answer in one field and "statusCode": 404 in another, and I read the first one.

I ran it twice more to be sure it wasn't a fluke. Same URL, same 404, three different answers:

run plans it invented rating reviews
1 $10 / $25 / $50 4.7 150
2 $29 / $59 4.7 1,200
3 $0 / $29 / $79 4.5 200

Three runs, three price lists, one page that isn't there. The model wasn't lying about the page — it was filling in what a page like that usually says.

Why it happens: a 404 from a big site is not an empty response. That one returned 79,877 bytes — navigation, footer, help links, other apps. Plenty of text to pattern-match against, none of it the thing you asked about. So the extraction step does what extraction steps do, and the structure it produces looks exactly like a real answer.

The rule

Never read a number out of a scrape without checking that the page returned 200 and that the number appears verbatim in the page's own text.

Two conditions, both cheap. The first catches the dead page. The second catches the live page that simply doesn't say what the model claims — which is the sneakier half.

Step 1 — The script

Save as checksrc somewhere on your PATH, then chmod +x checksrc.

#!/usr/bin/env bash
# checksrc — does this page exist, and does it actually contain the value?
# usage: checksrc <url> [value ...]
set -uo pipefail
url="$1"; shift
body=$(mktemp)
code=$(curl -sS -L -o "$body" -w '%{http_code}' \
        -A 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7)' "$url")
bytes=$(wc -c < "$body" | tr -d ' ')
echo "HTTP $code · ${bytes} bytes · $url"
if [ "$code" != "200" ]; then
  echo "STOP — page did not return 200. Anything extracted from it is invented."
  rm -f "$body"; exit 1
fi
title=$(tr '\n' ' ' < "$body" | grep -o -i '<title>[^<]*' | head -1 | cut -c8-)
echo "title: ${title:-<none>}"
case "$(printf '%s' "$title" | tr 'A-Z' 'a-z')" in
  *"not found"*|*"doesn't exist"*|*"does not exist"*|*"404"*|*"oh no"*|*"error"*)
    echo "STOP — 200 status but the title reads like an error page." ; rm -f "$body"; exit 1 ;;
esac
text=$(sed -e 's/<[^>]*>/ /g' "$body" | tr -s ' \n' ' ')
fail=0
for v in "$@"; do
  if printf '%s' "$text" | grep -qiF -- "$v"; then echo "  ok      $v"
  else echo "  MISSING $v"; fail=1; fi
done
rm -f "$body"
[ "$fail" = 0 ] || { echo "STOP — at least one value is not in the page text."; exit 1; }
echo "OK — page exists and every value appears in its text."

The title check is there because plenty of sites answer 200 with an error page. A status code alone is not enough on those.

Step 2 — Use it on the three cases

Dead page, claimed price:

$ checksrc "https://apps.shopify.com/workflowmail" "29"
HTTP 404 · 79877 bytes · https://apps.shopify.com/workflowmail
STOP — page did not return 200. Anything extracted from it is invented.

Real page, real values:

$ checksrc "https://github.com/heygen-com/hyperframes" "Apache-2.0" "Write HTML"
HTTP 200 · 479262 bytes
title: GitHub - heygen-com/hyperframes: Write HTML. Render video. Built for agents. · GitHub
  ok      Apache-2.0
  ok      Write HTML
OK — page exists and every value appears in its text.

Real page, invented values:

$ checksrc "https://github.com/heygen-com/hyperframes" "4.7 stars" '$49/month'
HTTP 200 · 479263 bytes
  MISSING 4.7 stars
  MISSING $49/month
STOP — at least one value is not in the page text.

Exit code is 0 only in the middle case, so you can chain it: checksrc "$url" "$price" && ./import.sh.

Step 3 — Make the agent do it without being asked

Save as ~/.claude/skills/verify-scrape/SKILL.md and restart Claude Code.

---
name: verify-scrape
description: Use whenever a scrape, crawl, fetch or extraction produced values (prices, ratings, counts, names, dates) that are about to be reported, stored or acted on. Verifies the source returned 200 and that each value appears verbatim in the page text before the values are used.
---

# Verify a scrape before trusting it

An extraction step returns a well-formed object whether or not the page exists. Structure is
not evidence. Before any scraped value is reported to the user, written to a file or database,
or used in a calculation, run this gate.

## The gate

1. **Status.** Confirm the response carried HTTP 200. If the scraping tool reports a status
   code field, read it — it is often in the response you already have, next to the answer.
   Anything other than 200 means STOP and say the page did not resolve. Do not report the
   extracted values, not even as a guess, and do not retry with a different phrasing of the
   same URL.
2. **Error page at 200.** Check the page title. If it contains "not found", "doesn't exist",
   "404", "oh no" or "error", treat it as case 1.
3. **Verbatim presence.** For each value you are about to report, confirm the value appears in
   the page's own text. Numbers must match as written, including separators. A value you
   cannot find is not a value you may report.
4. **Report what you checked.** State the status code and which values you found in the source.
   If a value failed, name it and leave it out rather than substituting a plausible one.

## Tooling

`checksrc <url> [value ...]` runs all of this in one call and exits non-zero on any failure.
Fall back to `curl -sS -L -o /tmp/p -w '%{http_code}'` plus `grep -F` if it is not installed.

## What this is not

This does not check whether the page is *correct* — only that it exists and says what you are
about to repeat. Truth about the world still needs a second source.

Where else this bites

Same failure, different dress:

  • A product that was delisted. The URL used to work, so it is in your spreadsheet. It now 404s, and the scraper refills it with the category average.
  • A geo-blocked or login-walled page. You get a 200 and a consent wall. The extractor reads the wall and invents the rest.
  • A renamed API endpoint. Returns 200 with {"error": "..."}. Structure parses fine.

The gate above catches all three, which is why it is worth the 30 lines.


Evidence: three live scrapes of the same non-existent URL on 10 and 11 Sep 2026, each response carrying statusCode: 404 alongside a complete invented answer. Script tested on 22 Sep 2026 against one 404 page, one live page with real values and the same live page with invented values — stop, pass, stop, exit codes 1 / 0 / 1.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment