Date: 2026-04-03 (updated 2026-06-17, 2026-08-04)
Prepared for: Customer Success
Collection: org_42caca5c-da0e-4c91-9bf3-d546266fd2e6_discovery-v1
Collection ID: c85d4618-b581-4413-a855-a4739125e705
Total chunks in Qdrant: 131,896
- Discovery.org (ID section)
- Discovery.org (Culture / Intelligent Design)
- ScienceAndCulture.com
- Bio-Complexity.org
- IntelligentDesign.org
- IDTheFuture.com
- 60+ Affiliate Books (PDFs)
- exploreevolution.com
- faithandevolution.org
- revolutionarybehe.com
- dissentfromdarwin.org
- humanzoos.org
- discovering.design
- aquinas.design
- biologicinstitute.org — Tumblr, not WordPress
- freescience.today
- teachingevolution.org
- caseyluskin.com
- darwinsdoubt.com
- signatureinthecell.com
- johngwest.com
- traipsingintoevolution.org
- richardsternberg.com
- stephencmeyer.org
- returnofthegodhypothesis.com
- michaelbehe.com
- ncseexposed.org
- mindmatters.ai — largest satellite site by volume
| # | Source | URL | Type | Status | Chunks in Qdrant | Notes |
|---|---|---|---|---|---|---|
| 1 | Discovery.org (ID section) | https://discovery.org/id | WordPress (subdirectory install) | Not ingested | 11 | Separate WP install with only 2 posts + 44 pages. Minimal unique content — most redirects to main site. Low priority. |
| 2 | Discovery.org (Culture) | https://www.discovery.org/c/intelligent-design/ | WordPress (category on main site) | Ingested | 33,077 (full site) | This is category #71 on the main discovery.org WP site. All discovery.org content is ingested nightly via automated cron, including articles in this category. |
| 3 | ScienceAndCulture.com | https://scienceandculture.com/ | WordPress | Ingested | 80,769 | Largest site by volume (15,070 posts). Automated nightly ingestion active. |
| 4 | Bio-Complexity.org | https://bio-complexity.org/ | Open Journal Systems (OJS) | Ingested (via PDF upload) | ~12,574 (est.) | Not a WordPress site — runs Open Journal Systems. All 54 journal articles + affiliate books ingested as PDFs via Reducto parsing. No additional web content to ingest. |
| 5 | IntelligentDesign.org | https://intelligentdesign.org/ | WordPress | Ingested | 248 | Small site (77 posts, 26 pages). ~90% content overlap with discovery.org. Automated nightly ingestion active. |
| 6 | IDTheFuture.com | https://idthefuture.com/ | WordPress | Ingested | 4,303 | Podcast site with 2,710 episodes stored as posts. Automated nightly ingestion active. |
| 7 | 60+ Discovery Institute Affiliate Books | PDF files provided | PDF → Reducto | Ingested | 13,505 (non-WP chunks) | 78 unique PDFs ingested via S3 upload + Reducto.ai parsing. Includes Bio-Complexity journal articles and affiliate books. |
Customer provided 20 additional sites. Probed each via WordPress REST API
(?rest_route=/wp/v2/...). 19 of 20 are WordPress; only
biologicinstitute.org is Tumblr (matches customer's own assessment). None
require authentication — all REST endpoints are publicly readable. All sites
are behind Cloudflare, so the REST API must be reached via the
?rest_route= query form with a browser User-Agent (the pretty-permalink
/wp-json/ path returns a themed HTML page to non-browser clients).
| # | Site | Platform | Auth | posts | pages | Custom types | Ingest-relevant total |
|---|---|---|---|---|---|---|---|
| A1 | exploreevolution.com | WordPress | None | 32 | 15 | ab_portfolio (0) |
47 |
| A2 | faithandevolution.org | WordPress | None | 0 | 36 | — | 36 |
| A3 | revolutionarybehe.com | WordPress | None | 123 | 4 | — | 127 |
| A4 | dissentfromdarwin.org | WordPress | None | 15 | 24 | — | 39 |
| A5 | humanzoos.org | WordPress | None | 51 | 14 | — | 65 |
| A6 | discovering.design | WordPress | None | 2 | 9 | — | 11 |
| A7 | aquinas.design | WordPress | None | 12 | 8 | — | 20 |
| A8 | biologicinstitute.org | Tumblr | n/a | n/a | n/a | n/a | n/a |
| A9 | freescience.today | WordPress | None | 21 | 25 | story (12) |
58 |
| A10 | teachingevolution.org | WordPress | None | 0 | 9 | — | 9 |
| A11 | caseyluskin.com | WordPress | None | 17 | 10 | gsm_styles (0, UI) |
27 |
| A12 | darwinsdoubt.com | WordPress | None | 0 | 16 | — | 16 |
| A13 | signatureinthecell.com | WordPress | None | 17 | 11 | ab_portfolio (0) |
28 |
| A14 | johngwest.com | WordPress | None | 475 | 6 | gsm_styles (0, UI) |
481 |
| A15 | traipsingintoevolution.org | WordPress | None | 0 | 1 | — | 1 |
| A16 | richardsternberg.com | WordPress | None | 0 | 15 | — | 15 |
| A17 | stephencmeyer.org | WordPress | None | 233 | 26 | guest-author (9, meta) |
259 |
| A18 | returnofthegodhypothesis.com | WordPress | None | 1 | 22 | gsm_styles (0, UI) |
23 |
| A19 | michaelbehe.com | WordPress | None | 74 | 15 | — | 89 |
| A20 | ncseexposed.org | WordPress | None | 0 | 1 | — | 1 |
| Totals (19 WP sites) | 1,073 | 267 | story 12 + guest-author 9 | ~1,352 items |
Notes on custom post types:
ab_portfolio(exploreevolution, signatureinthecell): 0 items — empty, no content to ingest.gsm_styles(caseyluskin, johngwest, returnofthegodhypothesis, stephencmeyer): UI style fragments from the GhostPool/theme builder, not content. Exclude.story(freescience.today): 12 items — real content (scientist profiles, e.g. Scott Minnich, Richard Sternberg). Ingest-worthy.guest-author(stephencmeyer): 9 items — Co-Authors Plus author metadata, not article content. Exclude from chunking, but useful for author attribution.
Sites with zero posts (faithandevolution, teachingevolution, darwinsdoubt,
traipsingintoevolution, richardsternberg, ncseexposed) are page-only brochure
sites built around a book or topic. They still have ingestible page content,
though traipsingintoevolution (1 page) and ncseexposed (1 page) are near-empty
landing pages — low value.
High-value targets (most content): johngwest.com (481), stephencmeyer.org (259), revolutionarybehe.com (127), michaelbehe.com (89), humanzoos.org (65), freescience.today (58).
Mind Matters ("Natural and Artificial Intelligence News and Analysis") is the Walter Bradley Center's publication. Probed with the same method as the 2026-06-17 batch: WordPress, public REST API, no auth required.
This one site is larger than the entire 2026-06-17 batch combined — 6,276 ingestible items vs ~1,352 across those 19 sites. Among all Discovery properties it is second only to scienceandculture.com (15,070 posts).
| # | Site | Platform | Auth | posts | pages | Custom types | Ingest-relevant total |
|---|---|---|---|---|---|---|---|
| A21 | mindmatters.ai | WordPress | None | 5,539 | 13 | brief (331), podcast (393) |
6,276 |
Estimated chunk impact: ~30,000 chunks. Basis: scienceandculture.com runs 5.36 chunks/post (15,070 posts → 80,769 chunks) and Mind Matters posts are comparable in length (sampled bodies 4.6k-9.0k characters). This would grow the collection by roughly 23% over its current 131,896 chunks — worth confirming with the customer before configuring, since it materially changes retrieval mix and nightly ingestion time.
Two custom types here carry real content and must be included in the ingest config, unlike the UI/meta types seen elsewhere:
brief(331 items) — short commentary posts with full article bodies (~3.4k characters sampled). Ingest-worthy.podcast(393 items) — episode pages with substantive show-note bodies (~1.4-1.7k characters sampled). Ingest-worthy. Same shape as the idthefuture.com podcast content already ingested.guest-author(6 items) — Co-Authors Plus metadata, no title or body in the REST payload. Exclude from chunking (same treatment as stephencmeyer.org), though useful for author attribution.gsm_styles(0 items) — empty theme-builder type. Exclude.media(12,539 items) — attachments, not content. Not counted above.
Coverage: posts run 2018-06-01 through 2026-08-03 and the site is actively publishing (3 posts on the most recent day sampled), so it needs the nightly incremental schedule rather than a one-off backfill.
- Status: Not ingested (not configured as a WordPress site in the collection)
- Investigation: This is a separate WordPress installation at the
/idsubdirectory with its own REST API at/id/wp-json/wp/v2/ - Content: Only 2 posts, 44 pages, 120
gsm_blockentries (UI fragments, not content), 0 video-series - Recommendation: Low value — most content is navigational pages. The 2 actual posts could be ingested if needed, but the 11 chunks already present (from URL overlap) likely cover it. No action required unless customer specifically requests it.
- Status: Fully ingested as part of the main discovery.org WordPress site
- How it works: This URL is a category filter (category #71) on the main discovery.org site. All discovery.org content — including articles, posts, pages, videos, and books — is ingested nightly via the WordPress cron scheduler. Content in the "Intelligent Design" category is included automatically.
- Chunks: 33,077 total for all discovery.org content
- Post types ingested:
a(articles),posts,pages,v(video),b(books)
- Status: Fully ingested with automated nightly cron
- Chunks: 80,769
- Post types:
posts,pages - Notes: Largest contributor to the collection by far (15,070 posts). Incremental ingestion runs nightly using
modified_afterto pick up new/updated content.
- Status: Fully ingested via PDF upload
- Platform: Open Journal Systems (OJS), not WordPress — cannot use WP API ingestion
- Content: 54 peer-reviewed journal articles covering intelligent design biology research
- Ingestion method: PDFs uploaded to S3 → parsed by Reducto.ai → chunked and embedded
- Notes: No additional web-only content beyond what's in the PDFs. The site is a thin JS-rendered wrapper around the journal articles.
- Status: Fully ingested with automated nightly cron
- Chunks: 248
- Post types:
posts,pages - Notes: Small site (77 posts, 26 pages). ~90% content overlap with discovery.org due to WordPress Distributor syndication plugin. Both copies are ingested.
- Status: Fully ingested with automated nightly cron
- Chunks: 4,303
- Post types:
posts,pages - Notes: Podcast/media site. Episodes are stored as WordPress
posts(2,710 items), not thepodcastcustom post type. ~80% overlap with discovery.org'spodcasttype via syndication.
- Status: Fully ingested
- Chunks: 13,505 (all non-WordPress content in the collection)
- Unique PDFs: 78 files
- Ingestion method: S3 upload → Reducto.ai PDF parsing → chunking → Qdrant
- Includes: Bio-Complexity journal articles + standalone books/publications
Probe method: each site's WordPress REST API was queried via
https://<site>/?rest_route=/wp/v2/types (post-type discovery) and
https://<site>/?rest_route=/wp/v2/<type>&per_page=1 reading the
X-WP-Total response header for volume. Auth checked by requesting
/wp/v2/posts anonymously (200 = public, 401 = auth required).
- Platform: WordPress (block theme) — Auth: none (public REST)
- Content: 32 posts, 15 pages. Custom type
ab_portfolioexists but is empty (0). - Total ingestible: ~47 items.
- Platform: WordPress — Auth: none
- Content: 0 posts, 36 pages. Page-only site (resource/debate library).
- Total ingestible: 36 pages.
- Platform: WordPress — Auth: none
- Content: 123 posts, 4 pages. Blog around Michael Behe's work.
- Total ingestible: ~127 items.
- Platform: WordPress — Auth: none
- Content: 15 posts, 24 pages.
- Total ingestible: ~39 items.
- Platform: WordPress — Auth: none
- Content: 51 posts, 14 pages. Documentary/book companion site.
- Total ingestible: ~65 items.
- Platform: WordPress — Auth: none
- Content: 2 posts, 9 pages. Small companion site.
- Total ingestible: ~11 items.
- Platform: WordPress — Auth: none
- Content: 12 posts, 8 pages.
- Total ingestible: ~20 items.
- Platform: Tumblr — not WordPress. Matches customer's own note.
- Investigation: Homepage is a Tumblr theme (
tumblr.com/assets, "Tumblr Theme by Easton Thomas"). No WordPress REST API. - Recommendation: Cannot ingest via the WordPress connector. Options: (a) Tumblr API / RSS feed scrape, (b) sitemap-based crawl, or (c) skip. Needs a different ingestion path than the rest of this batch. Flag for separate handling.
- Platform: WordPress — Auth: none
- Content: 21 posts, 25 pages, plus custom type
story(12 items) — scientist profile pages (Scott Minnich, Richard Sternberg, Günter Bechly, etc.). Real content, ingest-worthy. - Total ingestible: ~58 items (include the
storytype in the ingest config).
- Platform: WordPress — Auth: none
- Content: 0 posts, 9 pages. Page-only.
- Total ingestible: 9 pages.
- Platform: WordPress (WP Engine hosted) — Auth: none
- Content: 17 posts, 10 pages.
gsm_styles(0) is a UI fragment type — exclude. - Total ingestible: ~27 items.
- Platform: WordPress — Auth: none
- Content: 0 posts, 16 pages. Book companion site, page-only.
- Total ingestible: 16 pages.
- Platform: WordPress — Auth: none
- Content: 17 posts, 11 pages.
ab_portfolioempty (0). - Total ingestible: ~28 items.
- Platform: WordPress — Auth: none
- Content: 475 posts, 6 pages. Largest of the new batch by post count.
gsm_styles(0) UI fragments — exclude. - Total ingestible: ~481 items.
- Platform: WordPress — Auth: none
- Content: 0 posts, 1 page. Near-empty landing page (Kitzmiller v. Dover).
- Total ingestible: 1 page. Low value.
- Platform: WordPress — Auth: none
- Content: 0 posts, 15 pages. Page-only personal site.
- Total ingestible: 15 pages.
- Platform: WordPress (Site Kit by Google) — Auth: none
- Content: 233 posts, 26 pages, plus
guest-author(9 — Co-Authors Plus metadata, not article content; exclude from chunking but useful for attribution). - Total ingestible: ~259 items.
- Platform: WordPress — Auth: none
- Content: 1 post, 22 pages. Book promo site, mostly pages.
gsm_styles(0) UI fragments — exclude. - Total ingestible: ~23 items.
- Platform: WordPress (Site Kit by Google) — Auth: none
- Content: 74 posts, 15 pages.
- Total ingestible: ~89 items.
- Platform: WordPress — Auth: none
- Content: 0 posts, 1 page. Near-empty landing page.
- Total ingestible: 1 page. Low value.
- Platform: WordPress (block theme, Co-Authors Plus) — Auth: none (public REST)
- Site name: Mind Matters — "Natural and Artificial Intelligence News and Analysis" (Walter Bradley Center)
- Content: 5,539 posts, 13 pages, plus custom types
brief(331) andpodcast(393) — both carry full bodies and are ingest-worthy.guest-author(6) is CAP metadata;gsm_styles(0) is empty. 12,539 media attachments are not content. - Date range: 2018-06-01 to 2026-08-03, actively publishing — needs the nightly incremental schedule.
- Total ingestible: 6,276 items (~30,000 chunks est.), larger than the whole 2026-06-17 batch combined.
- Requested: 2026-08-04. Not yet configured.
All 4 original WordPress sites run on automated nightly cron schedules:
- discovery.org — nightly incremental ingestion
- scienceandculture.com — nightly incremental ingestion
- intelligentdesign.org — nightly incremental ingestion
- idthefuture.com — nightly incremental ingestion
New/modified posts are automatically detected and ingested each night. A collection-level observability system tracks expected vs actual post counts per site and emits Sentry alerts for any discrepancies (shipped in PR #80, pending merge).
| Item | Priority | Status |
|---|---|---|
| Add mindmatters.ai to multi-WP collection | High | Inventoried 2026-08-04 — public, no auth. 6,276 items / ~30,000 chunks est. (+23% collection growth). Include brief and podcast custom types; needs nightly schedule. Confirm scope with customer before configuring. |
| Add 18 new WordPress sites to multi-WP collection | High | Inventoried 2026-06-17 — all public, no auth. ~1,352 items total. Ready to configure. |
freescience.today story custom type |
Medium | Include story (12 items) in ingest config — it carries real content, not just posts/pages. |
| biologicinstitute.org (Tumblr) | Medium | Needs separate ingestion path (Tumblr API / RSS / sitemap crawl). Not WP-connector compatible. |
| Exclude UI/meta custom types | Low | Skip gsm_styles (UI) and guest-author (CAP metadata) and empty ab_portfolio during ingest. |
| Near-empty sites (traipsingintoevolution, ncseexposed) | Low | 1 page each — confirm with customer whether worth configuring. |
| Discovery.org /id section ingestion | Low | Not needed — minimal unique content (2 posts) |
| Content deduplication across syndicated sites | Medium | Known ~80-90% overlap between discovery.org and satellite sites. Currently both copies are ingested. |
| Nightly run observability | Done | PR #80 adds collection-level expected/actual tracking with Sentry alerts |
Generated by engineering team, 2026-04-03. Additional-sites inventory added 2026-06-17. mindmatters.ai inventory added 2026-08-04.