Skip to content

Instantly share code, notes, and snippets.

@jordotech
Last active August 4, 2026 17:12
Show Gist options
  • Select an option

  • Save jordotech/984745b0caeb400cc85993fe41819ff0 to your computer and use it in GitHub Desktop.

Select an option

Save jordotech/984745b0caeb400cc85993fe41819ff0 to your computer and use it in GitHub Desktop.

Discovery Institute — Data Scope Status Report

Date: 2026-04-03 (updated 2026-06-17, 2026-08-04) Prepared for: Customer Success Collection: org_42caca5c-da0e-4c91-9bf3-d546266fd2e6_discovery-v1 Collection ID: c85d4618-b581-4413-a855-a4739125e705 Total chunks in Qdrant: 131,896


Table of Contents (by site)

Currently ingested / scoped (original report)

  1. Discovery.org (ID section)
  2. Discovery.org (Culture / Intelligent Design)
  3. ScienceAndCulture.com
  4. Bio-Complexity.org
  5. IntelligentDesign.org
  6. IDTheFuture.com
  7. 60+ Affiliate Books (PDFs)

Additional sites requested 2026-06-17 (new inventory)

  1. exploreevolution.com
  2. faithandevolution.org
  3. revolutionarybehe.com
  4. dissentfromdarwin.org
  5. humanzoos.org
  6. discovering.design
  7. aquinas.design
  8. biologicinstitute.orgTumblr, not WordPress
  9. freescience.today
  10. teachingevolution.org
  11. caseyluskin.com
  12. darwinsdoubt.com
  13. signatureinthecell.com
  14. johngwest.com
  15. traipsingintoevolution.org
  16. richardsternberg.com
  17. stephencmeyer.org
  18. returnofthegodhypothesis.com
  19. michaelbehe.com
  20. ncseexposed.org

Additional site requested 2026-08-04

  1. mindmatters.ailargest satellite site by volume

Status Summary

# Source URL Type Status Chunks in Qdrant Notes
1 Discovery.org (ID section) https://discovery.org/id WordPress (subdirectory install) Not ingested 11 Separate WP install with only 2 posts + 44 pages. Minimal unique content — most redirects to main site. Low priority.
2 Discovery.org (Culture) https://www.discovery.org/c/intelligent-design/ WordPress (category on main site) Ingested 33,077 (full site) This is category #71 on the main discovery.org WP site. All discovery.org content is ingested nightly via automated cron, including articles in this category.
3 ScienceAndCulture.com https://scienceandculture.com/ WordPress Ingested 80,769 Largest site by volume (15,070 posts). Automated nightly ingestion active.
4 Bio-Complexity.org https://bio-complexity.org/ Open Journal Systems (OJS) Ingested (via PDF upload) ~12,574 (est.) Not a WordPress site — runs Open Journal Systems. All 54 journal articles + affiliate books ingested as PDFs via Reducto parsing. No additional web content to ingest.
5 IntelligentDesign.org https://intelligentdesign.org/ WordPress Ingested 248 Small site (77 posts, 26 pages). ~90% content overlap with discovery.org. Automated nightly ingestion active.
6 IDTheFuture.com https://idthefuture.com/ WordPress Ingested 4,303 Podcast site with 2,710 episodes stored as posts. Automated nightly ingestion active.
7 60+ Discovery Institute Affiliate Books PDF files provided PDF → Reducto Ingested 13,505 (non-WP chunks) 78 unique PDFs ingested via S3 upload + Reducto.ai parsing. Includes Bio-Complexity journal articles and affiliate books.

Additional Sites — Inventory Summary (requested 2026-06-17)

Customer provided 20 additional sites. Probed each via WordPress REST API (?rest_route=/wp/v2/...). 19 of 20 are WordPress; only biologicinstitute.org is Tumblr (matches customer's own assessment). None require authentication — all REST endpoints are publicly readable. All sites are behind Cloudflare, so the REST API must be reached via the ?rest_route= query form with a browser User-Agent (the pretty-permalink /wp-json/ path returns a themed HTML page to non-browser clients).

# Site Platform Auth posts pages Custom types Ingest-relevant total
A1 exploreevolution.com WordPress None 32 15 ab_portfolio (0) 47
A2 faithandevolution.org WordPress None 0 36 36
A3 revolutionarybehe.com WordPress None 123 4 127
A4 dissentfromdarwin.org WordPress None 15 24 39
A5 humanzoos.org WordPress None 51 14 65
A6 discovering.design WordPress None 2 9 11
A7 aquinas.design WordPress None 12 8 20
A8 biologicinstitute.org Tumblr n/a n/a n/a n/a n/a
A9 freescience.today WordPress None 21 25 story (12) 58
A10 teachingevolution.org WordPress None 0 9 9
A11 caseyluskin.com WordPress None 17 10 gsm_styles (0, UI) 27
A12 darwinsdoubt.com WordPress None 0 16 16
A13 signatureinthecell.com WordPress None 17 11 ab_portfolio (0) 28
A14 johngwest.com WordPress None 475 6 gsm_styles (0, UI) 481
A15 traipsingintoevolution.org WordPress None 0 1 1
A16 richardsternberg.com WordPress None 0 15 15
A17 stephencmeyer.org WordPress None 233 26 guest-author (9, meta) 259
A18 returnofthegodhypothesis.com WordPress None 1 22 gsm_styles (0, UI) 23
A19 michaelbehe.com WordPress None 74 15 89
A20 ncseexposed.org WordPress None 0 1 1
Totals (19 WP sites) 1,073 267 story 12 + guest-author 9 ~1,352 items

Notes on custom post types:

  • ab_portfolio (exploreevolution, signatureinthecell): 0 items — empty, no content to ingest.
  • gsm_styles (caseyluskin, johngwest, returnofthegodhypothesis, stephencmeyer): UI style fragments from the GhostPool/theme builder, not content. Exclude.
  • story (freescience.today): 12 items — real content (scientist profiles, e.g. Scott Minnich, Richard Sternberg). Ingest-worthy.
  • guest-author (stephencmeyer): 9 items — Co-Authors Plus author metadata, not article content. Exclude from chunking, but useful for author attribution.

Sites with zero posts (faithandevolution, teachingevolution, darwinsdoubt, traipsingintoevolution, richardsternberg, ncseexposed) are page-only brochure sites built around a book or topic. They still have ingestible page content, though traipsingintoevolution (1 page) and ncseexposed (1 page) are near-empty landing pages — low value.

High-value targets (most content): johngwest.com (481), stephencmeyer.org (259), revolutionarybehe.com (127), michaelbehe.com (89), humanzoos.org (65), freescience.today (58).


Additional Site — mindmatters.ai (requested 2026-08-04)

Mind Matters ("Natural and Artificial Intelligence News and Analysis") is the Walter Bradley Center's publication. Probed with the same method as the 2026-06-17 batch: WordPress, public REST API, no auth required.

This one site is larger than the entire 2026-06-17 batch combined — 6,276 ingestible items vs ~1,352 across those 19 sites. Among all Discovery properties it is second only to scienceandculture.com (15,070 posts).

# Site Platform Auth posts pages Custom types Ingest-relevant total
A21 mindmatters.ai WordPress None 5,539 13 brief (331), podcast (393) 6,276

Estimated chunk impact: ~30,000 chunks. Basis: scienceandculture.com runs 5.36 chunks/post (15,070 posts → 80,769 chunks) and Mind Matters posts are comparable in length (sampled bodies 4.6k-9.0k characters). This would grow the collection by roughly 23% over its current 131,896 chunks — worth confirming with the customer before configuring, since it materially changes retrieval mix and nightly ingestion time.

Two custom types here carry real content and must be included in the ingest config, unlike the UI/meta types seen elsewhere:

  • brief (331 items) — short commentary posts with full article bodies (~3.4k characters sampled). Ingest-worthy.
  • podcast (393 items) — episode pages with substantive show-note bodies (~1.4-1.7k characters sampled). Ingest-worthy. Same shape as the idthefuture.com podcast content already ingested.
  • guest-author (6 items) — Co-Authors Plus metadata, no title or body in the REST payload. Exclude from chunking (same treatment as stephencmeyer.org), though useful for author attribution.
  • gsm_styles (0 items) — empty theme-builder type. Exclude.
  • media (12,539 items) — attachments, not content. Not counted above.

Coverage: posts run 2018-06-01 through 2026-08-03 and the site is actively publishing (3 posts on the most recent day sampled), so it needs the nightly incremental schedule rather than a one-off backfill.


Detailed Notes

1. Discovery.org (ID section) — https://discovery.org/id

  • Status: Not ingested (not configured as a WordPress site in the collection)
  • Investigation: This is a separate WordPress installation at the /id subdirectory with its own REST API at /id/wp-json/wp/v2/
  • Content: Only 2 posts, 44 pages, 120 gsm_block entries (UI fragments, not content), 0 video-series
  • Recommendation: Low value — most content is navigational pages. The 2 actual posts could be ingested if needed, but the 11 chunks already present (from URL overlap) likely cover it. No action required unless customer specifically requests it.

2. Discovery.org (Culture / Intelligent Design) — https://www.discovery.org/c/intelligent-design/

  • Status: Fully ingested as part of the main discovery.org WordPress site
  • How it works: This URL is a category filter (category #71) on the main discovery.org site. All discovery.org content — including articles, posts, pages, videos, and books — is ingested nightly via the WordPress cron scheduler. Content in the "Intelligent Design" category is included automatically.
  • Chunks: 33,077 total for all discovery.org content
  • Post types ingested: a (articles), posts, pages, v (video), b (books)

3. ScienceAndCulture.com

  • Status: Fully ingested with automated nightly cron
  • Chunks: 80,769
  • Post types: posts, pages
  • Notes: Largest contributor to the collection by far (15,070 posts). Incremental ingestion runs nightly using modified_after to pick up new/updated content.

4. Bio-Complexity.org

  • Status: Fully ingested via PDF upload
  • Platform: Open Journal Systems (OJS), not WordPress — cannot use WP API ingestion
  • Content: 54 peer-reviewed journal articles covering intelligent design biology research
  • Ingestion method: PDFs uploaded to S3 → parsed by Reducto.ai → chunked and embedded
  • Notes: No additional web-only content beyond what's in the PDFs. The site is a thin JS-rendered wrapper around the journal articles.

5. IntelligentDesign.org

  • Status: Fully ingested with automated nightly cron
  • Chunks: 248
  • Post types: posts, pages
  • Notes: Small site (77 posts, 26 pages). ~90% content overlap with discovery.org due to WordPress Distributor syndication plugin. Both copies are ingested.

6. IDTheFuture.com

  • Status: Fully ingested with automated nightly cron
  • Chunks: 4,303
  • Post types: posts, pages
  • Notes: Podcast/media site. Episodes are stored as WordPress posts (2,710 items), not the podcast custom post type. ~80% overlap with discovery.org's podcast type via syndication.

7. 60+ Discovery Institute Affiliate Books (PDFs)

  • Status: Fully ingested
  • Chunks: 13,505 (all non-WordPress content in the collection)
  • Unique PDFs: 78 files
  • Ingestion method: S3 upload → Reducto.ai PDF parsing → chunking → Qdrant
  • Includes: Bio-Complexity journal articles + standalone books/publications

Additional Sites — Detailed Notes (2026-06-17)

Probe method: each site's WordPress REST API was queried via https://<site>/?rest_route=/wp/v2/types (post-type discovery) and https://<site>/?rest_route=/wp/v2/<type>&per_page=1 reading the X-WP-Total response header for volume. Auth checked by requesting /wp/v2/posts anonymously (200 = public, 401 = auth required).

A1. exploreevolution.com

  • Platform: WordPress (block theme) — Auth: none (public REST)
  • Content: 32 posts, 15 pages. Custom type ab_portfolio exists but is empty (0).
  • Total ingestible: ~47 items.

A2. faithandevolution.org

  • Platform: WordPress — Auth: none
  • Content: 0 posts, 36 pages. Page-only site (resource/debate library).
  • Total ingestible: 36 pages.

A3. revolutionarybehe.com

  • Platform: WordPress — Auth: none
  • Content: 123 posts, 4 pages. Blog around Michael Behe's work.
  • Total ingestible: ~127 items.

A4. dissentfromdarwin.org

  • Platform: WordPress — Auth: none
  • Content: 15 posts, 24 pages.
  • Total ingestible: ~39 items.

A5. humanzoos.org

  • Platform: WordPress — Auth: none
  • Content: 51 posts, 14 pages. Documentary/book companion site.
  • Total ingestible: ~65 items.

A6. discovering.design

  • Platform: WordPress — Auth: none
  • Content: 2 posts, 9 pages. Small companion site.
  • Total ingestible: ~11 items.

A7. aquinas.design

  • Platform: WordPress — Auth: none
  • Content: 12 posts, 8 pages.
  • Total ingestible: ~20 items.

A8. biologicinstitute.org

  • Platform: Tumblrnot WordPress. Matches customer's own note.
  • Investigation: Homepage is a Tumblr theme (tumblr.com/assets, "Tumblr Theme by Easton Thomas"). No WordPress REST API.
  • Recommendation: Cannot ingest via the WordPress connector. Options: (a) Tumblr API / RSS feed scrape, (b) sitemap-based crawl, or (c) skip. Needs a different ingestion path than the rest of this batch. Flag for separate handling.

A9. freescience.today

  • Platform: WordPress — Auth: none
  • Content: 21 posts, 25 pages, plus custom type story (12 items) — scientist profile pages (Scott Minnich, Richard Sternberg, Günter Bechly, etc.). Real content, ingest-worthy.
  • Total ingestible: ~58 items (include the story type in the ingest config).

A10. teachingevolution.org

  • Platform: WordPress — Auth: none
  • Content: 0 posts, 9 pages. Page-only.
  • Total ingestible: 9 pages.

A11. caseyluskin.com

  • Platform: WordPress (WP Engine hosted) — Auth: none
  • Content: 17 posts, 10 pages. gsm_styles (0) is a UI fragment type — exclude.
  • Total ingestible: ~27 items.

A12. darwinsdoubt.com

  • Platform: WordPress — Auth: none
  • Content: 0 posts, 16 pages. Book companion site, page-only.
  • Total ingestible: 16 pages.

A13. signatureinthecell.com

  • Platform: WordPress — Auth: none
  • Content: 17 posts, 11 pages. ab_portfolio empty (0).
  • Total ingestible: ~28 items.

A14. johngwest.com

  • Platform: WordPress — Auth: none
  • Content: 475 posts, 6 pages. Largest of the new batch by post count. gsm_styles (0) UI fragments — exclude.
  • Total ingestible: ~481 items.

A15. traipsingintoevolution.org

  • Platform: WordPress — Auth: none
  • Content: 0 posts, 1 page. Near-empty landing page (Kitzmiller v. Dover).
  • Total ingestible: 1 page. Low value.

A16. richardsternberg.com

  • Platform: WordPress — Auth: none
  • Content: 0 posts, 15 pages. Page-only personal site.
  • Total ingestible: 15 pages.

A17. stephencmeyer.org

  • Platform: WordPress (Site Kit by Google) — Auth: none
  • Content: 233 posts, 26 pages, plus guest-author (9 — Co-Authors Plus metadata, not article content; exclude from chunking but useful for attribution).
  • Total ingestible: ~259 items.

A18. returnofthegodhypothesis.com

  • Platform: WordPress — Auth: none
  • Content: 1 post, 22 pages. Book promo site, mostly pages. gsm_styles (0) UI fragments — exclude.
  • Total ingestible: ~23 items.

A19. michaelbehe.com

  • Platform: WordPress (Site Kit by Google) — Auth: none
  • Content: 74 posts, 15 pages.
  • Total ingestible: ~89 items.

A20. ncseexposed.org

  • Platform: WordPress — Auth: none
  • Content: 0 posts, 1 page. Near-empty landing page.
  • Total ingestible: 1 page. Low value.

A21. mindmatters.ai

  • Platform: WordPress (block theme, Co-Authors Plus) — Auth: none (public REST)
  • Site name: Mind Matters — "Natural and Artificial Intelligence News and Analysis" (Walter Bradley Center)
  • Content: 5,539 posts, 13 pages, plus custom types brief (331) and podcast (393) — both carry full bodies and are ingest-worthy. guest-author (6) is CAP metadata; gsm_styles (0) is empty. 12,539 media attachments are not content.
  • Date range: 2018-06-01 to 2026-08-03, actively publishing — needs the nightly incremental schedule.
  • Total ingestible: 6,276 items (~30,000 chunks est.), larger than the whole 2026-06-17 batch combined.
  • Requested: 2026-08-04. Not yet configured.

Automated Ingestion Schedule

All 4 original WordPress sites run on automated nightly cron schedules:

  • discovery.org — nightly incremental ingestion
  • scienceandculture.com — nightly incremental ingestion
  • intelligentdesign.org — nightly incremental ingestion
  • idthefuture.com — nightly incremental ingestion

New/modified posts are automatically detected and ingested each night. A collection-level observability system tracks expected vs actual post counts per site and emits Sentry alerts for any discrepancies (shipped in PR #80, pending merge).


Open Items

Item Priority Status
Add mindmatters.ai to multi-WP collection High Inventoried 2026-08-04 — public, no auth. 6,276 items / ~30,000 chunks est. (+23% collection growth). Include brief and podcast custom types; needs nightly schedule. Confirm scope with customer before configuring.
Add 18 new WordPress sites to multi-WP collection High Inventoried 2026-06-17 — all public, no auth. ~1,352 items total. Ready to configure.
freescience.today story custom type Medium Include story (12 items) in ingest config — it carries real content, not just posts/pages.
biologicinstitute.org (Tumblr) Medium Needs separate ingestion path (Tumblr API / RSS / sitemap crawl). Not WP-connector compatible.
Exclude UI/meta custom types Low Skip gsm_styles (UI) and guest-author (CAP metadata) and empty ab_portfolio during ingest.
Near-empty sites (traipsingintoevolution, ncseexposed) Low 1 page each — confirm with customer whether worth configuring.
Discovery.org /id section ingestion Low Not needed — minimal unique content (2 posts)
Content deduplication across syndicated sites Medium Known ~80-90% overlap between discovery.org and satellite sites. Currently both copies are ingested.
Nightly run observability Done PR #80 adds collection-level expected/actual tracking with Sentry alerts

Generated by engineering team, 2026-04-03. Additional-sites inventory added 2026-06-17. mindmatters.ai inventory added 2026-08-04.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment