Skip to content

Instantly share code, notes, and snippets.

@contactbrenton
Last active July 1, 2026 09:12
Show Gist options
  • Select an option

  • Save contactbrenton/8005e7456888c800af4f3e2b1bdfd936 to your computer and use it in GitHub Desktop.

Select an option

Save contactbrenton/8005e7456888c800af4f3e2b1bdfd936 to your computer and use it in GitHub Desktop.
copilot-get-transcripts.md

TASK: Incrementally Archive My Meeting Transcripts as WebVTT + YAML Sidecars

You are an assistant with access to my workplace tools (calendar, online-meeting / transcript service, and local file + zip capabilities). Build and MAINTAIN a growing archive of meeting transcripts: each transcript saved as WebVTT with a matching YAML sidecar, tracked in a CSV ledger so repeat runs only fetch what's new or missing, delivered as a zip snapshot with a README and manifest.

────────────────────────────────────────────────────────

0. CAPABILITY PROBE (do this first — do not skip)

──────────────────────────────────────────────────────── Do not assume calendar access implies transcript access. PROBE, don't interrogate: a. list calendar events for a small recent window, b. extract one online-meeting identity from an event, c. list that meeting's transcript metadata, d. download one transcript's content, e. create a local file, and f. create a zip. If any of (a)–(f) fails, STOP and return an execution plan only, naming the exact blocked capability. Never fabricate around a missing capability.

Connector note: this runs end-to-end only where a Microsoft 365 / Teams connector is available (native in Microsoft Copilot & Copilot Cowork; ChatGPT/Claude need an M365 connector or MCP integration). Expect retention to dominate results — most historical transcript CONTENT is purged even when metadata still lists it, so a "5 year" request typically yields only recent months. Set that expectation explicitly.

────────────────────────────────────────────────────────

Parameters (edit these)

────────────────────────────────────────────────────────

  • LOOKBACK: {{e.g. "the last 5 years"}}
  • SCOPE: {{"own tenant" | "all"}}
  • EXPORT_DIR: {{persistent folder for the archive, e.g. "MeetingTranscripts/"}}
  • LEDGER_FILE: "{{EXPORT_DIR}}/transcript-ledger.csv" # persists across runs
  • RECHECK_GAPS: {{"transient-only" (default) | "all" | "none"}}
  • TIME ZONE / LANGUAGE: {{optional}}

SCOPE definition: for "own tenant", include a meeting when the organiser's TENANT ID (preferred) or email domain matches mine. Do NOT exclude meetings merely because external attendees were present — instead tag external attendees in metadata.

────────────────────────────────────────────────────────

Incremental model (the ledger) — read before the steps

──────────────────────────────────────────────────────── LEDGER_FILE is a CSV that is the source of truth for "what I already have". Columns:

transcript_id,meeting_id,event_id,series_id,subject,occurrence_start_utc, occurrence_end_utc,organiser,source_format,retrieval_status,gap_reason, vtt_file,yaml_file,word_count,retrieved_utc,last_checked_utc

On every run, decide per discovered transcript_id:

  • In ledger as included AND its vtt_file exists on disk → SKIP (already have it).
  • In ledger as included but file MISSING → RE-FETCH (self-heal).
  • In ledger as gap past_retention → SKIP (won't return), unless RECHECK_GAPS = "all".
  • In ledger as gap policy_denied / error / not_found → RE-ATTEMPT if RECHECK_GAPS is "transient-only" or "all"; else skip.
  • Not in ledger → FETCH (new). Only fetch content for the resulting MISSING set. Never re-download an included, file-present transcript. Update last_checked_utc for every discovered transcript.

────────────────────────────────────────────────────────

Method

────────────────────────────────────────────────────────

  1. RESOLVE "NOW": establish today's date and time zone; convert LOOKBACK to explicit start/end dates.

  2. LOAD LEDGER: read LEDGER_FILE if it exists; else start an empty ledger.

  3. ENUMERATE FROM THE CALENDAR (source of truth — a "recent transcripts" feed alone is unreliable). Page the whole range. Keep online meetings (with a join link), apply SCOPE, skip cancelled events. Capture per event: subject, event_id, online-meeting id, series/master id (if recurring), occurrence start, organiser. NOTE these blind spots in the README: ad-hoc "Meet now" calls, Teams CHANNEL meetings, webinars/town halls, meetings I wasn't invited to, and deleted events — these may not appear on the calendar.

  4. IDENTIFY TRANSCRIPTS (don't lose occurrences). De-duplicate MEETINGS by meeting/thread id only to avoid querying the same meeting twice — a recurring series is usually ONE meeting object whose transcript list already contains EVERY occurrence. Query each unique meeting's transcript list and treat each TRANSCRIPT as unique by (transcript_id + occurrence_start). Confirm the transcript count is consistent with the number of occurrences you expected; never collapse occurrences.

  5. DIFF against the ledger (per the Incremental model) to produce the MISSING set. Report how many are already archived vs. how many will be fetched this run.

  6. FETCH CONTENT — MISSING set only, honestly.

    • Retention purges content over time; expect "not found" on older items. Retry once.
    • NEVER invent, paraphrase, sample, or placeholder transcript text.
    • Record a gap REASON for anything not fetched: past_retention | policy_denied | error | not_found. (Split these in the manifest — don't lump them.)
    • Note each transcript's SOURCE FORMAT: native_webvtt or derived_text.
  7. PRODUCE WEBVTT.

    • If the source is ALREADY valid WebVTT → pass it through UNCHANGED (set source_format: native_webvtt, reconstructed: false). Do not rebuild it.
    • If the source is derived text (e.g. "[HH:MM:SS] Speaker: text") → reconstruct valid WebVTT (source_format: derived_text, reconstructed: true):
      • First line WEBVTT; one numbered cue per utterance: 1 00:01:22.000 --> 00:01:29.000 What they said.
      • Escape cue text: replace & < > with & < >; the literal "-->" must never appear in payload.
      • Missing speaker → omit the tag; empty payload → "(inaudible)".
      • Cue end = next cue's start, and end MUST be > start. If timestamps are coarse (e.g. per-second), spread same-second utterances so cues are strictly sequential and NEVER overlap; mark sub-second/end times as interpolated.
      • FINAL cue: use the source end time if present; else start + 5s (or +1s if no later boundary is known). Flag its end time as interpolated.
      • Validate every timestamp matches HH:MM:SS.mmm and every cue parses.
  8. WRITE A YAML SIDECAR per transcript (basename shared: file.vtt + file.vtt.yaml). Quote STRING scalars; use native numbers, booleans, and arrays; no folded blocks. Include provenance so any file traces back to its source:

    title: "Weekly Ops Sync"
    date: "2026-06-24"
    start_time_utc: "2026-06-24T03:59:44Z"
    end_time_utc: "2026-06-24T04:40:25Z"
    duration_minutes: 40.7
    source: "Microsoft Teams meeting transcript (Microsoft 365)"
    content_type: "meeting_transcript"
    format: "WebVTT"
    media_type: "text/vtt"
    source_format: "derived_text"        # or "native_webvtt"
    reconstructed: true                   # false when native_webvtt
    timing_precision_seconds: 1           # omit when native_webvtt
    timing_note: "Sub-second positions and end times are interpolated; whole-second starts reflect the source."
    language: "en"
    participants: ["Alex Kim", "Sam Rivera"]
    external_attendees: []                # tag any out-of-tenant attendees here
    speaker_count: 2
    cue_count: 128
    word_count: 4210
    retrieved_utc: "2026-07-01T09:00:00Z"
    retrieval_status: "included"
    gap_reason: ""
    transcript_file: "2026-06-24_1359_weekly-ops-sync_a1b2c3d4.vtt"
    source_event_id: "<calendar event id>"
    source_meeting_id: "<online meeting id>"
    source_transcript_id: "<transcript id>"
    series_id: "<recurring series master id, or empty>"
    organiser: "alex.kim@example.com"
    
  9. FILE NAMING (deterministic, cross-OS, collision-proof): YYYY-MM-DD_HHMM__.vtt Slug: lowercase; spaces→hyphens; strip characters illegal on Windows/macOS/Linux ( \ / : * ? " < > | and control chars ); transliterate or drop emoji/CJK; collapse repeats; truncate slug to 80 chars. Short id = first 8 chars of the transcript id. The sidecar is the same name + ".yaml".

  10. UPDATE THE LEDGER: upsert one row per discovered transcript (included and gaps), with provenance IDs, status, gap_reason, filenames, word_count, retrieved_utc, last_checked_utc. Write LEDGER_FILE back to EXPORT_DIR. Gap rows carry the same source IDs so gaps are auditable too.

  11. REGENERATE README.md AND MANIFEST.md from the FULL ledger (cumulative, not just this run):

    • MANIFEST.md: table of every meeting — transcripts included, date range, and gap counts split by reason (past_retention / policy_denied / error / not_found).
    • README.md: concrete counts, retrievable date range, file-naming convention, how to use the sidecars, the reconstruction caveat, retention + calendar-coverage blind spots, how the incremental ledger works, and a "nothing was fabricated" statement.
  12. VALIDATE-BEFORE-ZIP GATE: parse EVERY .vtt (header + cue timings, zero illegal overlaps, end>start) and EVERY .yaml (loads, required keys, points to an existing .vtt). If anything fails, fix or record it — do not ship a broken file silently.

  13. PACKAGE: zip the full current set of .vtt + .yaml files, the ledger CSV, README.md, and MANIFEST.md into a dated snapshot in EXPORT_DIR. If no zip utility exists, create the archive programmatically.

────────────────────────────────────────────────────────

Guardrails

────────────────────────────────────────────────────────

  • Accuracy over completeness: report exactly what was and wasn't retrieved, and why.
  • Never fabricate names, dates, quotes, or transcript text.
  • Privacy (high-sensitivity data): do NOT email, share, upload to third-party storage, summarise, or quote transcript content in the chat unless I explicitly ask. Only include transcripts the connector reports as accessible to me and within SCOPE.
  • Rate limits: parallelise lookups/downloads only where the connector supports it safely; on throttling, back off and retry; record still-unresolved items as gaps.

────────────────────────────────────────────────────────

Final report

──────────────────────────────────────────────────────── State: archive/zip location and ledger path; total transcripts in the archive; how many were NEW this run vs. already present; gaps by reason; and the meetings covered. Counts and locations only — no transcript excerpts.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment