You are an assistant with access to my workplace tools (calendar, online-meeting / transcript service, and local file + zip capabilities). Build and MAINTAIN a growing archive of meeting transcripts: each transcript saved as WebVTT with a matching YAML sidecar, tracked in a CSV ledger so repeat runs only fetch what's new or missing, delivered as a zip snapshot with a README and manifest.
────────────────────────────────────────────────────────
──────────────────────────────────────────────────────── Do not assume calendar access implies transcript access. PROBE, don't interrogate: a. list calendar events for a small recent window, b. extract one online-meeting identity from an event, c. list that meeting's transcript metadata, d. download one transcript's content, e. create a local file, and f. create a zip. If any of (a)–(f) fails, STOP and return an execution plan only, naming the exact blocked capability. Never fabricate around a missing capability.
Connector note: this runs end-to-end only where a Microsoft 365 / Teams connector is available (native in Microsoft Copilot & Copilot Cowork; ChatGPT/Claude need an M365 connector or MCP integration). Expect retention to dominate results — most historical transcript CONTENT is purged even when metadata still lists it, so a "5 year" request typically yields only recent months. Set that expectation explicitly.
────────────────────────────────────────────────────────
────────────────────────────────────────────────────────
- LOOKBACK: {{e.g. "the last 5 years"}}
- SCOPE: {{"own tenant" | "all"}}
- EXPORT_DIR: {{persistent folder for the archive, e.g. "MeetingTranscripts/"}}
- LEDGER_FILE: "{{EXPORT_DIR}}/transcript-ledger.csv" # persists across runs
- RECHECK_GAPS: {{"transient-only" (default) | "all" | "none"}}
- TIME ZONE / LANGUAGE: {{optional}}
SCOPE definition: for "own tenant", include a meeting when the organiser's TENANT ID (preferred) or email domain matches mine. Do NOT exclude meetings merely because external attendees were present — instead tag external attendees in metadata.
────────────────────────────────────────────────────────
──────────────────────────────────────────────────────── LEDGER_FILE is a CSV that is the source of truth for "what I already have". Columns:
transcript_id,meeting_id,event_id,series_id,subject,occurrence_start_utc, occurrence_end_utc,organiser,source_format,retrieval_status,gap_reason, vtt_file,yaml_file,word_count,retrieved_utc,last_checked_utc
On every run, decide per discovered transcript_id:
- In ledger as
includedAND its vtt_file exists on disk → SKIP (already have it). - In ledger as
includedbut file MISSING → RE-FETCH (self-heal). - In ledger as gap
past_retention→ SKIP (won't return), unless RECHECK_GAPS = "all". - In ledger as gap
policy_denied/error/not_found→ RE-ATTEMPT if RECHECK_GAPS is "transient-only" or "all"; else skip. - Not in ledger → FETCH (new).
Only fetch content for the resulting MISSING set. Never re-download an included,
file-present transcript. Update
last_checked_utcfor every discovered transcript.
────────────────────────────────────────────────────────
────────────────────────────────────────────────────────
-
RESOLVE "NOW": establish today's date and time zone; convert LOOKBACK to explicit start/end dates.
-
LOAD LEDGER: read LEDGER_FILE if it exists; else start an empty ledger.
-
ENUMERATE FROM THE CALENDAR (source of truth — a "recent transcripts" feed alone is unreliable). Page the whole range. Keep online meetings (with a join link), apply SCOPE, skip cancelled events. Capture per event: subject, event_id, online-meeting id, series/master id (if recurring), occurrence start, organiser. NOTE these blind spots in the README: ad-hoc "Meet now" calls, Teams CHANNEL meetings, webinars/town halls, meetings I wasn't invited to, and deleted events — these may not appear on the calendar.
-
IDENTIFY TRANSCRIPTS (don't lose occurrences). De-duplicate MEETINGS by meeting/thread id only to avoid querying the same meeting twice — a recurring series is usually ONE meeting object whose transcript list already contains EVERY occurrence. Query each unique meeting's transcript list and treat each TRANSCRIPT as unique by (transcript_id + occurrence_start). Confirm the transcript count is consistent with the number of occurrences you expected; never collapse occurrences.
-
DIFF against the ledger (per the Incremental model) to produce the MISSING set. Report how many are already archived vs. how many will be fetched this run.
-
FETCH CONTENT — MISSING set only, honestly.
- Retention purges content over time; expect "not found" on older items. Retry once.
- NEVER invent, paraphrase, sample, or placeholder transcript text.
- Record a gap REASON for anything not fetched: past_retention | policy_denied | error | not_found. (Split these in the manifest — don't lump them.)
- Note each transcript's SOURCE FORMAT: native_webvtt or derived_text.
-
PRODUCE WEBVTT.
- If the source is ALREADY valid WebVTT → pass it through UNCHANGED (set source_format: native_webvtt, reconstructed: false). Do not rebuild it.
- If the source is derived text (e.g. "[HH:MM:SS] Speaker: text") → reconstruct
valid WebVTT (source_format: derived_text, reconstructed: true):
- First line WEBVTT; one numbered cue per utterance: 1 00:01:22.000 --> 00:01:29.000 What they said.
- Escape cue text: replace & < > with & < >; the literal "-->" must never appear in payload.
- Missing speaker → omit the tag; empty payload → "(inaudible)".
- Cue end = next cue's start, and end MUST be > start. If timestamps are coarse (e.g. per-second), spread same-second utterances so cues are strictly sequential and NEVER overlap; mark sub-second/end times as interpolated.
- FINAL cue: use the source end time if present; else start + 5s (or +1s if no later boundary is known). Flag its end time as interpolated.
- Validate every timestamp matches HH:MM:SS.mmm and every cue parses.
-
WRITE A YAML SIDECAR per transcript (basename shared: file.vtt + file.vtt.yaml). Quote STRING scalars; use native numbers, booleans, and arrays; no folded blocks. Include provenance so any file traces back to its source:
title: "Weekly Ops Sync" date: "2026-06-24" start_time_utc: "2026-06-24T03:59:44Z" end_time_utc: "2026-06-24T04:40:25Z" duration_minutes: 40.7 source: "Microsoft Teams meeting transcript (Microsoft 365)" content_type: "meeting_transcript" format: "WebVTT" media_type: "text/vtt" source_format: "derived_text" # or "native_webvtt" reconstructed: true # false when native_webvtt timing_precision_seconds: 1 # omit when native_webvtt timing_note: "Sub-second positions and end times are interpolated; whole-second starts reflect the source." language: "en" participants: ["Alex Kim", "Sam Rivera"] external_attendees: [] # tag any out-of-tenant attendees here speaker_count: 2 cue_count: 128 word_count: 4210 retrieved_utc: "2026-07-01T09:00:00Z" retrieval_status: "included" gap_reason: "" transcript_file: "2026-06-24_1359_weekly-ops-sync_a1b2c3d4.vtt" source_event_id: "<calendar event id>" source_meeting_id: "<online meeting id>" source_transcript_id: "<transcript id>" series_id: "<recurring series master id, or empty>" organiser: "alex.kim@example.com" -
FILE NAMING (deterministic, cross-OS, collision-proof): YYYY-MM-DD_HHMM__.vtt Slug: lowercase; spaces→hyphens; strip characters illegal on Windows/macOS/Linux ( \ / : * ? " < > | and control chars ); transliterate or drop emoji/CJK; collapse repeats; truncate slug to 80 chars. Short id = first 8 chars of the transcript id. The sidecar is the same name + ".yaml".
-
UPDATE THE LEDGER: upsert one row per discovered transcript (included and gaps), with provenance IDs, status, gap_reason, filenames, word_count, retrieved_utc, last_checked_utc. Write LEDGER_FILE back to EXPORT_DIR. Gap rows carry the same source IDs so gaps are auditable too.
-
REGENERATE README.md AND MANIFEST.md from the FULL ledger (cumulative, not just this run):
- MANIFEST.md: table of every meeting — transcripts included, date range, and gap counts split by reason (past_retention / policy_denied / error / not_found).
- README.md: concrete counts, retrievable date range, file-naming convention, how to use the sidecars, the reconstruction caveat, retention + calendar-coverage blind spots, how the incremental ledger works, and a "nothing was fabricated" statement.
-
VALIDATE-BEFORE-ZIP GATE: parse EVERY .vtt (header + cue timings, zero illegal overlaps, end>start) and EVERY .yaml (loads, required keys, points to an existing .vtt). If anything fails, fix or record it — do not ship a broken file silently.
-
PACKAGE: zip the full current set of .vtt + .yaml files, the ledger CSV, README.md, and MANIFEST.md into a dated snapshot in EXPORT_DIR. If no zip utility exists, create the archive programmatically.
────────────────────────────────────────────────────────
────────────────────────────────────────────────────────
- Accuracy over completeness: report exactly what was and wasn't retrieved, and why.
- Never fabricate names, dates, quotes, or transcript text.
- Privacy (high-sensitivity data): do NOT email, share, upload to third-party storage, summarise, or quote transcript content in the chat unless I explicitly ask. Only include transcripts the connector reports as accessible to me and within SCOPE.
- Rate limits: parallelise lookups/downloads only where the connector supports it safely; on throttling, back off and retry; record still-unresolved items as gaps.
────────────────────────────────────────────────────────
──────────────────────────────────────────────────────── State: archive/zip location and ledger path; total transcripts in the archive; how many were NEW this run vs. already present; gaps by reason; and the meetings covered. Counts and locations only — no transcript excerpts.