--- name: "get-meeting-transcript" description: "Grab a Teams/Stream meeting transcript by driving a logged-in browser when the Graph download is blocked. Captures the underlying .vtt from network traffic, or scrapes the recap transcript pane. Accepts a Teams recap, Stream/SharePoint, or meet.microsoft.com URL, or a calendar reference it resolves to one. Saves .vtt + .md + the meeting chat into the Obsidian vault. Triggers: get transcript, scrape transcript, grab the transcript, meeting chat, transcript for my meeting." --- ## get-meeting-transcript When the user wants a transcript for a meeting they can view in the UI but can't pull via Graph (`/me/onlineMeetings/.../transcripts` returns nothing or 403s), use this skill instead of giving up. **Before scraping, try the official download.** The scraper is the third choice, not the first — see "Choosing a capture path" below. A downloaded `.vtt` is strictly better than anything this script synthesizes, and getting one costs the user about ten seconds. ### Choosing a capture path Try these in order. Stop at the first that works. **1. Native Graph.** (`m365_*` does not currently expose transcripts; the EOD automation uses an internal Stream call.) If it succeeds, you don't need this skill at all. **2. Official `.vtt` download from the UI — prefer this over scraping.** In the Teams recap (or the Stream/SharePoint player), the transcript pane has a **download** action that yields the real `.vtt`: every cue at full millisecond precision, proper `` tags, and cue identifiers. That is strictly better than the synthesized file this script produces from the DOM, which is lossy in timing (see "Fidelity of the DOM scrape"). If the meeting is one the user can open right now, **ask them to download it** before reaching for the scraper — it is faster than a manual-mode walk (which already requires them to navigate to the same pane) and higher fidelity. Have them save it into the meeting's output dir, then: - adopt it as `transcript.vtt` (replacing any scraped copy), and - render `transcript.md` from it using the documented format — `### HH:MM:SS — Speaker` headings, consecutive same-speaker cues collapsed into paragraphs — **skipping the GUID cue-identifier lines**, which are not transcript text. Record `source: "official Teams transcript download (.vtt)"` in the markdown front matter so later readers know the timings are exact. **3. This scraper.** Use it when the download action is unavailable, blocked by tenant policy, or the user isn't around to click it. Common cases: meetings the user attended but didn't organize, restrictive transcript-export policy, externally-organized meetings. Content and speaker attribution are reliable; timings are not. If a scrape has already run and a download later becomes available, prefer the download and regenerate `transcript.md` from it — the scraped `.vtt` should be replaced, not kept alongside. ### How it works A bundled Playwright script (`scrape-transcript.mjs`) drives Edge against a persistent profile (`.pw-profile/`) so the user's Teams/Stream/SharePoint sign-in carries over. **Two capture paths** (auto-detected, network first, DOM fallback): 1. **Network interception** (rare — Teams encrypts transcript blobs in transit, so this usually fails for Teams recap). Listens on every network response for either a WEBVTT signature or a JSON transcript shape with recognizable cue arrays (`{ text, startTime/startOffset, speaker }`). 2. **DOM scrape of the recap iframe** (the usual path for Teams meetings the user doesn't organize). The transcript pane lives inside the `RecapxPlatIframe` SharePoint iframe at `*_layouts/15/xplatplugins.aspx?...hv=Recap*` inside a region with `aria-label="Transcript"`. Cue entries are `[id^="entry-"]` with `aria-label="@N M minutes S seconds"`. Speaker headers (`.itemDisplayName-*`) only appear when the speaker changes — those get forward-filled across following cues. **Virtualization handling.** The recap iframe uses a virtual list (only ~10 cues rendered at once). To materialize every cue, the script focuses the first entry then drives **PageDown + ArrowDown keypresses** to walk through the entire list (`aria-setsize` tells us the total; normal-zone pacing is PageDown+5×ArrowDown per pass with a 150ms settle wait). Cues are accumulated by their stable `entry-N` id, so re-renders don't cause duplicates. **Tail resilience.** Big `PageDown` jumps are more likely to overshoot the render boundary near the *end* of the list than in the middle — observed failure mode: a run plateaus short of `aria-setsize` (e.g. 307/329, 369/402, 355/366) because the walk's stuck-counter trips right at the tail. Once collection reaches 85% of the reported total, the script automatically switches to gentler single `ArrowDown` steps with a 275ms settle wait, and tolerates 3x more "stuck" passes before giving up. Keyboard nav only moves list *focus*, though — some virtualized-list implementations render more rows off a real, trusted scroll/wheel input rather than focus changes alone, so a stall that survives repeated keyboard nudges can still respond to an actual scroll. **Note:** a JS-dispatched synthetic `WheelEvent` + manual `scrollTop` bump was tried first and confirmed *not* to work here — diagnostics showed the transcript region has no CSS overflow-scroll container at all (`scrollHeight === clientHeight` wherever probed), so those synthetic events landed on nothing. The script instead uses Playwright's `page.mouse.wheel()`: a trusted, OS-level input event dispatched at the transcript pane's real screen coordinates, which can reach handlers that reject untrusted/synthetic events or rely on the browser's native wheel-to-scroll pipeline. Every 4th stuck pass in the tail zone, and throughout the dedicated tail-recovery phase below, the script fires a mouse wheel nudge as a supplement to keyboard nudges. If it still falls short after the main loop, a dedicated tail-recovery phase spends an extra bounded budget (`--tail-wait`, default 45s) alternating keyboard (`End`+`ArrowDown`) and mouse-wheel nudges. The result JSON reports `cueCount`, `targetCueCount`, and `truncated` so you can tell at a glance whether anything was missed — no need to hand-write a gap analysis script. When truncated, the missing cues are virtually always a contiguous block at the point the walk stalled (usually the tail), not scattered through the transcript. If `--headed`, the user can click into the transcript pane manually; the DOM scrape still works because it's keyed off the iframe structure, not on the navigation path used to reach it. ### Input modes **1. Direct URL.** Anything that resolves to a page showing the transcript: - Teams recap URL: `https://teams.microsoft.com/.../recap/...` - Stream / SharePoint video page: `*.sharepoint.com/.../stream.aspx?id=...` or `*.sharepoint.com/personal/.../_layouts/15/stream.aspx?...` **Live-join links auto-fallback to manual.** `onlineMeeting.joinUrl` / `onlineMeetingUrl` values — `teams.microsoft.com/meet/...`, `teams.microsoft.com/l/meetup-join/...`, `meet.microsoft.com/...` — always open the pre-join/lobby screen, **never** the recap, even for meetings that already ended. (Earlier revisions of this skill assumed Teams would redirect these into the recap; it doesn't.) The script detects this pattern itself (`isLiveJoinUrl()`) and immediately switches into the manual walk instead of burning the `--wait` timeout on a dead-end screen — you don't need to pre-filter these URLs or retry with `--manual` yourself; just pass whatever URL you have and let the script decide. Don't waste a turn calling it with the join URL expecting a direct hit — assume it'll need the manual walk and tell the user up front so they're ready to navigate to Recap → Transcript. **2. Calendar reference.** User says "transcript from yesterday's 1:1 with Alex" or names a meeting. The agent (you) must resolve this to a URL before invoking the scraper: 1. Call `m365_list_events` with an appropriate `startDate`/`endDate` window. 2. Match by subject / attendee. Confirm with the user via `m_ask_user` if multiple events match. 3. Pull the transcript URL from the event: - any SharePoint stream.aspx URL in the event body (often present after recording is auto-saved to OneDrive) — try this first, it can go straight to a direct-URL capture. - `onlineMeeting.joinUrl` / a `teams.microsoft.com/meet/...` link in the body — this will auto-fallback to the manual walk (see above), so treat it as "go straight to manual" rather than a URL worth waiting on. 4. Hand the URL to `scrape-transcript.mjs`. If neither lookup succeeds, ask the user to paste the URL directly. If you already know the only URL available is a join link, you can skip straight to `--manual` yourself and tell the user to expect the manual walk — no need to round-trip through a direct-URL attempt first. ### Invocation ```pwsh node "C:\Users\dkucinski\.scout\m-skills\get-meeting-transcript\scrape-transcript.mjs" ` --url "" ` [--out ""] ` [--subject ""] ` [--headed] ` [--manual] ` [--wait 60] ` [--tail-wait 45] ` [--debug] ``` - `--url` is the page to land on. Skip when using `--manual`. If `--url` is a live-join link (`teams.microsoft.com/meet/...`, `/l/meetup-join/...`, `meet.microsoft.com/...`), the script auto-detects it and behaves as if `--manual` were passed (overrides the URL to `teams.microsoft.com/v2/`, extends the wait to 300s) — printing a note to stderr explaining why. - `--manual` opens `teams.microsoft.com` and waits while the user navigates to the recap themselves (use this directly when you already know the only URL you have is a join link — no need to let the script rediscover that). Implies `--headed`. Default wait jumps to 300s in manual mode. Extraction starts **automatically** as soon as the transcript pane is detected — you just navigate to Recap → Transcript and the scraper proceeds on its own. After reaching the pane, clicking one cue is recommended (gives the list keyboard focus for the PageDown/ArrowDown walk), but the scraper also self-focuses the first cue. Pressing Enter in the terminal is an optional override that forces extraction to start immediately. The old behavior — blocking until you typed something — has been removed, so manual mode works whether or not stdin is interactive. - `--out` defaults to `D:\Repos\Obsidian\02 - Meetings\YYYY-MM-DD - \` using today's date and `--subject` (falls back to `meeting`). Pass an explicit path with the meeting's own date when scraping past meetings. - `--headed` opens a visible browser window. Required on first run so the user can sign in. - `--wait N` overrides the seconds to wait for transcript readiness/capture after the page loads (default 45; 300 in `--manual`). Bump higher for very long meetings (the keyboard walk takes ~1 ArrowDown per cue). - `--tail-wait N` extra seconds (default 45) spent on gentle End+ArrowDown recovery nudges if the DOM-scrape walk plateaus short of the transcript's reported total (`aria-setsize`) — see "Tail resilience" above. Set to `0` to disable and fail fast instead. - `--debug` logs every response URL + content-type the listener sees plus scroll-progress dumps, for diagnosing capture failures. stdout is a single JSON object: ```json { "ok": true, "vtt": "C:\\...\\transcript.vtt", "md": "C:\\...\\transcript.md", "cueCount": 785, "targetCueCount": 785, "truncated": false, "bytes": 100976, "source": "dom" } ``` `source` is `"network"` when captured from a WEBVTT/JSON response, or `"dom"` when extracted from the recap iframe DOM. For Teams meetings, expect `"dom"` — the transcript blob is encrypted in transit. `targetCueCount` is only present for DOM-scrape captures (the iframe's `aria-setsize`); if `truncated` is `true`, `cueCount` is short of it — check stderr for the `[warn] transcript is likely INCOMPLETE` line, and consider re-running with a larger `--tail-wait` if it keeps happening. Exit codes: `0` ok, `2` bad args, `3` no transcript captured before timeout, `4` page never loaded. ### One-time setup First run, always pass `--headed`. The script opens Edge against `.pw-profile/`. Sign in to teams.microsoft.com and/or the SharePoint tenant when prompted. The profile persists; subsequent runs reuse it silently. ### Output Files written into the output dir: - `transcript.vtt` — WEBVTT. When `source` is `"network"`, this is byte-for-byte what the UI player consumes. When `source` is `"dom"` (the usual case for Teams), it is a **synthesized** WEBVTT and is **lossy in timing** — see "Fidelity of the DOM scrape" below. - `transcript.md` — same content as readable markdown with `### HH:MM:SS — Speaker` headings and consecutive same-speaker cues collapsed into single paragraphs. Easy to read in Obsidian. - `chat.md` — the meeting chat, when one exists. Written by the agent, not the script. See "Meeting chat capture" below. Default output dir matches the existing EOD meeting-archive layout under `D:\Repos\Obsidian\02 - Meetings\` so transcripts grabbed this way slot into the same vault structure as the automated ones. ### Fidelity of the DOM scrape Measured 2026-08-14 against an official `.vtt` download of the same 57-minute meeting (Astra Analytics walkthrough), the DOM scrape was: - **Faithful on text** — 10,507 words vs 10,503 in the official file. The 4 extra are the "X started/stopped transcription" system lines, which the recap DOM exposes and the official export omits. - **Faithful on speaker attribution** — per-speaker word counts matched within ±4 words across all five participants. - **Lossy on timing.** 433 captured entries vs 1,091 real cues. The recap DOM exposes *grouped* entries, and the script rounds each to whole seconds, so a synthesized cue spans several real ones (e.g. `00:00:03.000 --> 00:00:24.000` covers five cues running `00:00:03.288 --> 00:00:24.168`). So: **the DOM scrape is trustworthy for content and attribution, and should not be trusted for precise timings.** Don't cite DOM-scraped timestamps as exact, and don't use them to align against recordings or other time-series. Note that `cueCount` counts *recap entries*, not VTT cues — `433/433` with `truncated: false` is a complete capture even though the real transcript has 1,091 cues. The two numbers are not comparable, and a mismatch is not a bug. **If you verify a capture against a reference `.vtt`, strip the cue-identifier lines first.** Official Teams exports put a GUID id line (e.g. `63fe10ca-…-152225ba9c70/7-1`) before each timestamp; a naive parser counts those as transcript text and will report a large phantom word deficit (~7 tokens × cue count). This has already produced one false "41% of content is missing" alarm. ### Meeting chat capture (`chat.md`) After a successful transcript capture, also grab the **meeting chat** — the Teams conversation attached to the same meeting. It routinely carries what the transcript can't: links pasted mid-call, files shared, and the follow-up traffic that lands in the hours or days afterwards. This is an **agent** step. `scrape-transcript.mjs` does not do it — the script only handles the transcript pane. Do it after the script returns `ok: true`. **1. Resolve the meeting chat id.** - `workiq_search_chats` with the meeting subject is usually enough. Meeting chats come back with `chatType: "meeting"` and an id shaped `19:meeting_@thread.v2`. - Sanity-check the returned `members` against the calendar event's attendees before using it. If the subject is generic and several chats match, confirm via `m_ask_user` — never guess. - If no chat exists (common for meetings nobody typed in), say so and move on. A missing chat is not a failure of the run. **2. Export it to `/chat.md`.** Follow the `/teams-export` rendering conventions so it reads identically to everything under `00 - Chats\`: `##` for the date (YYYY-MM-DD), `###` for each message (`HH:mm — Author`), replies as nested blockquotes, attachments collapsed to a single `**Attachments:**` line (never download binaries), message HTML converted to markdown, timestamps in America/Los_Angeles 24-hour, and `---` between root messages. Front matter: ```yaml --- type: reference subtype: meeting-chat date: chat_id: "19:meeting_...@thread.v2" topic: captured: --- ``` Export the **whole** chat rather than a date window — meeting chats are small, and the valuable part is often the post-meeting tail. **3. Emit a `sources.yaml` suggestion — do NOT write it.** Most meeting chats are dead within a day, and `chat.md` is the right home for those: self-contained, sitting next to the transcript it belongs to. But some keep producing follow-up traffic for days, and those are worth promoting to a standing source so `/teams-export` keeps pulling them. That promotion is the user's call, not this skill's. `sources.yaml` is hand-maintained and `/teams-export` is explicitly forbidden from writing to it; this skill must not write to it either. Instead, print a ready-to-paste block and let the user decide: ```yaml # Suggested addition to D:\Repos\Obsidian\00 - Chats\sources.yaml (groups:) - alias: type: group chat_id: "19:meeting_...@thread.v2" topic: notes: >- Added . Meeting chat for the session. Transcript captured under "02 - Meetings/ - ". Retire with enabled:false once follow-up traffic stops. ``` Recommend promotion when the meeting has a live follow-up thread — handoffs, access grants, artifacts still to be shared, or a recurring series. Recommend against it for one-off sessions that concluded in the room; `chat.md` already covers those, and every dead entry left in `sources.yaml` is re-polled on every future export run. ### Guardrails - `.pw-profile/` and any captured cookies are credentials. Never echo them, never copy them outside the skill dir. - Transcripts may contain sensitive content (MIP labels are not surfaced via the UI capture). Treat any transcript saved this way as at least as sensitive as the calendar event itself. Warn the user before any outbound share. - Only write inside the `--out` dir. Do not modify anything else in the Obsidian vault. This includes `00 - Chats\sources.yaml` — that file is hand-maintained by the user; suggest additions, never apply them. - `chat.md` inherits the same sensitivity treatment as the transcript, and can be worse — chats carry pasted links, file names, and customer detail that never got said out loud. Same rule: warn before any outbound share. - If the user supplies a calendar reference that resolves to multiple events, always confirm via `m_ask_user` — never guess. Same for a meeting subject that matches multiple Teams chats.