19 KiB
| name | description |
|---|---|
| get-meeting-transcript | Grab a Teams/Stream meeting transcript by driving a logged-in browser when the Graph download is blocked. Captures the underlying .vtt from network traffic, or scrapes the recap transcript pane. Accepts a Teams recap, Stream/SharePoint, or meet.microsoft.com URL, or a calendar reference it resolves to one. Saves .vtt + .md + the meeting chat into the Obsidian vault. Triggers: get transcript, scrape transcript, grab the transcript, meeting chat, transcript for my meeting. |
get-meeting-transcript
When the user wants a transcript for a meeting they can view in the UI but
can't pull via Graph (/me/onlineMeetings/.../transcripts returns nothing or
403s), use this skill instead of giving up.
Before scraping, try the official download. The scraper is the third
choice, not the first — see "Choosing a capture path" below. A downloaded
.vtt is strictly better than anything this script synthesizes, and getting
one costs the user about ten seconds.
Choosing a capture path
Try these in order. Stop at the first that works.
1. Native Graph. (m365_* does not currently expose transcripts; the EOD
automation uses an internal Stream call.) If it succeeds, you don't need this
skill at all.
2. Official .vtt download from the UI — prefer this over scraping.
In the Teams recap (or the Stream/SharePoint player), the transcript pane has
a download action that yields the real .vtt: every cue at full
millisecond precision, proper <v Speaker> tags, and cue identifiers. That is
strictly better than the synthesized file this script produces from the DOM,
which is lossy in timing (see "Fidelity of the DOM scrape").
If the meeting is one the user can open right now, ask them to download it before reaching for the scraper — it is faster than a manual-mode walk (which already requires them to navigate to the same pane) and higher fidelity. Have them save it into the meeting's output dir, then:
- adopt it as
transcript.vtt(replacing any scraped copy), and - render
transcript.mdfrom it using the documented format —### HH:MM:SS — Speakerheadings, consecutive same-speaker cues collapsed into paragraphs — skipping the GUID cue-identifier lines, which are not transcript text.
Record source: "official Teams transcript download (.vtt)" in the markdown
front matter so later readers know the timings are exact.
3. This scraper. Use it when the download action is unavailable, blocked by tenant policy, or the user isn't around to click it. Common cases: meetings the user attended but didn't organize, restrictive transcript-export policy, externally-organized meetings. Content and speaker attribution are reliable; timings are not.
If a scrape has already run and a download later becomes available, prefer the
download and regenerate transcript.md from it — the scraped .vtt should be
replaced, not kept alongside.
How it works
A bundled Playwright script (scrape-transcript.mjs) drives Edge against a
persistent profile (.pw-profile/) so the user's Teams/Stream/SharePoint
sign-in carries over.
Two capture paths (auto-detected, network first, DOM fallback):
- Network interception (rare — Teams encrypts transcript blobs in
transit, so this usually fails for Teams recap). Listens on every network
response for either a WEBVTT signature or a JSON transcript shape with
recognizable cue arrays (
{ text, startTime/startOffset, speaker }). - DOM scrape of the recap iframe (the usual path for Teams meetings the
user doesn't organize). The transcript pane lives inside the
RecapxPlatIframeSharePoint iframe at*_layouts/15/xplatplugins.aspx?...hv=Recap*inside a region witharia-label="Transcript". Cue entries are[id^="entry-"]witharia-label="@N M minutes S seconds". Speaker headers (.itemDisplayName-*) only appear when the speaker changes — those get forward-filled across following cues.
Virtualization handling. The recap iframe uses a virtual list (only ~10
cues rendered at once). To materialize every cue, the script focuses the
first entry then drives PageDown + ArrowDown keypresses to walk through
the entire list (aria-setsize tells us the total; normal-zone pacing is
PageDown+5×ArrowDown per pass with a 150ms settle wait). Cues are accumulated
by their stable entry-N id, so re-renders don't cause duplicates.
Tail resilience. Big PageDown jumps are more likely to overshoot the
render boundary near the end of the list than in the middle — observed
failure mode: a run plateaus short of aria-setsize (e.g. 307/329, 369/402,
355/366) because the walk's stuck-counter trips right at the tail. Once
collection reaches 85% of the reported total, the script automatically
switches to gentler single ArrowDown steps with a 275ms settle wait, and
tolerates 3x more "stuck" passes before giving up. Keyboard nav only moves
list focus, though — some virtualized-list implementations render more
rows off a real, trusted scroll/wheel input rather than focus changes alone,
so a stall that survives repeated keyboard nudges can still respond to an
actual scroll. Note: a JS-dispatched synthetic WheelEvent +
manual scrollTop bump was tried first and confirmed not to work here —
diagnostics showed the transcript region has no CSS overflow-scroll
container at all (scrollHeight === clientHeight wherever probed), so those
synthetic events landed on nothing. The script instead uses Playwright's
page.mouse.wheel(): a trusted, OS-level input event dispatched at the
transcript pane's real screen coordinates, which can reach handlers that
reject untrusted/synthetic events or rely on the browser's native
wheel-to-scroll pipeline. Every 4th stuck pass in the tail zone, and
throughout the dedicated tail-recovery phase below, the script fires a mouse
wheel nudge as a supplement to keyboard nudges. If it still falls short
after the main loop, a dedicated tail-recovery phase spends an extra bounded
budget (--tail-wait, default 45s) alternating keyboard (End+ArrowDown)
and mouse-wheel nudges. The result JSON reports cueCount, targetCueCount,
and truncated so you can tell at a glance whether anything was missed — no
need to hand-write a gap analysis script. When truncated, the missing cues
are virtually always a contiguous block at the point the walk stalled
(usually the tail), not scattered through the transcript.
If --headed, the user can click into the transcript pane manually; the
DOM scrape still works because it's keyed off the iframe structure, not on
the navigation path used to reach it.
Input modes
1. Direct URL. Anything that resolves to a page showing the transcript:
- Teams recap URL:
https://teams.microsoft.com/.../recap/... - Stream / SharePoint video page:
*.sharepoint.com/.../stream.aspx?id=...or*.sharepoint.com/personal/.../_layouts/15/stream.aspx?...
Live-join links auto-fallback to manual. onlineMeeting.joinUrl /
onlineMeetingUrl values — teams.microsoft.com/meet/...,
teams.microsoft.com/l/meetup-join/..., meet.microsoft.com/... — always
open the pre-join/lobby screen, never the recap, even for meetings that
already ended. (Earlier revisions of this skill assumed Teams would redirect
these into the recap; it doesn't.) The script detects this pattern itself
(isLiveJoinUrl()) and immediately switches into the manual walk instead of
burning the --wait timeout on a dead-end screen — you don't need to
pre-filter these URLs or retry with --manual yourself; just pass whatever
URL you have and let the script decide. Don't waste a turn calling it with
the join URL expecting a direct hit — assume it'll need the manual walk and
tell the user up front so they're ready to navigate to Recap → Transcript.
2. Calendar reference. User says "transcript from yesterday's 1:1 with Alex" or names a meeting. The agent (you) must resolve this to a URL before invoking the scraper:
- Call
m365_list_eventswith an appropriatestartDate/endDatewindow. - Match by subject / attendee. Confirm with the user via
m_ask_userif multiple events match. - Pull the transcript URL from the event:
- any SharePoint stream.aspx URL in the event body (often present after recording is auto-saved to OneDrive) — try this first, it can go straight to a direct-URL capture.
onlineMeeting.joinUrl/ ateams.microsoft.com/meet/...link in the body — this will auto-fallback to the manual walk (see above), so treat it as "go straight to manual" rather than a URL worth waiting on.
- Hand the URL to
scrape-transcript.mjs.
If neither lookup succeeds, ask the user to paste the URL directly. If you
already know the only URL available is a join link, you can skip straight to
--manual yourself and tell the user to expect the manual walk — no need to
round-trip through a direct-URL attempt first.
Invocation
node "C:\Users\dkucinski\.scout\m-skills\get-meeting-transcript\scrape-transcript.mjs" `
--url "<meeting or stream URL>" `
[--out "<output dir>"] `
[--subject "<filename slug>"] `
[--headed] `
[--manual] `
[--wait 60] `
[--tail-wait 45] `
[--debug]
--urlis the page to land on. Skip when using--manual. If--urlis a live-join link (teams.microsoft.com/meet/...,/l/meetup-join/...,meet.microsoft.com/...), the script auto-detects it and behaves as if--manualwere passed (overrides the URL toteams.microsoft.com/v2/, extends the wait to 300s) — printing a note to stderr explaining why.--manualopensteams.microsoft.comand waits while the user navigates to the recap themselves (use this directly when you already know the only URL you have is a join link — no need to let the script rediscover that). Implies--headed. Default wait jumps to 300s in manual mode. Extraction starts automatically as soon as the transcript pane is detected — you just navigate to Recap → Transcript and the scraper proceeds on its own. After reaching the pane, clicking one cue is recommended (gives the list keyboard focus for the PageDown/ArrowDown walk), but the scraper also self-focuses the first cue. Pressing Enter in the terminal is an optional override that forces extraction to start immediately. The old behavior — blocking until you typed something — has been removed, so manual mode works whether or not stdin is interactive.--outdefaults toD:\Repos\Obsidian\02 - Meetings\YYYY-MM-DD - <subject>\using today's date and--subject(falls back tomeeting). Pass an explicit path with the meeting's own date when scraping past meetings.--headedopens a visible browser window. Required on first run so the user can sign in.--wait Noverrides the seconds to wait for transcript readiness/capture after the page loads (default 45; 300 in--manual). Bump higher for very long meetings (the keyboard walk takes ~1 ArrowDown per cue).--tail-wait Nextra seconds (default 45) spent on gentle End+ArrowDown recovery nudges if the DOM-scrape walk plateaus short of the transcript's reported total (aria-setsize) — see "Tail resilience" above. Set to0to disable and fail fast instead.--debuglogs every response URL + content-type the listener sees plus scroll-progress dumps, for diagnosing capture failures.
stdout is a single JSON object:
{ "ok": true, "vtt": "C:\\...\\transcript.vtt", "md": "C:\\...\\transcript.md", "cueCount": 785, "targetCueCount": 785, "truncated": false, "bytes": 100976, "source": "dom" }
source is "network" when captured from a WEBVTT/JSON response, or
"dom" when extracted from the recap iframe DOM. For Teams meetings, expect
"dom" — the transcript blob is encrypted in transit. targetCueCount is
only present for DOM-scrape captures (the iframe's aria-setsize); if
truncated is true, cueCount is short of it — check stderr for the
[warn] transcript is likely INCOMPLETE line, and consider re-running with
a larger --tail-wait if it keeps happening.
Exit codes: 0 ok, 2 bad args, 3 no transcript captured before
timeout, 4 page never loaded.
One-time setup
First run, always pass --headed. The script opens Edge against
.pw-profile/. Sign in to teams.microsoft.com and/or the SharePoint tenant
when prompted. The profile persists; subsequent runs reuse it silently.
Output
Files written into the output dir:
transcript.vtt— WEBVTT. Whensourceis"network", this is byte-for-byte what the UI player consumes. Whensourceis"dom"(the usual case for Teams), it is a synthesized WEBVTT and is lossy in timing — see "Fidelity of the DOM scrape" below.transcript.md— same content as readable markdown with### HH:MM:SS — Speakerheadings and consecutive same-speaker cues collapsed into single paragraphs. Easy to read in Obsidian.chat.md— the meeting chat, when one exists. Written by the agent, not the script. See "Meeting chat capture" below.
Default output dir matches the existing EOD meeting-archive layout under
D:\Repos\Obsidian\02 - Meetings\ so transcripts grabbed this way slot
into the same vault structure as the automated ones.
Fidelity of the DOM scrape
Measured 2026-08-14 against an official .vtt download of the same 57-minute
meeting (Astra Analytics walkthrough), the DOM scrape was:
- Faithful on text — 10,507 words vs 10,503 in the official file. The 4 extra are the "X started/stopped transcription" system lines, which the recap DOM exposes and the official export omits.
- Faithful on speaker attribution — per-speaker word counts matched within ±4 words across all five participants.
- Lossy on timing. 433 captured entries vs 1,091 real cues. The recap DOM
exposes grouped entries, and the script rounds each to whole seconds, so
a synthesized cue spans several real ones (e.g.
00:00:03.000 --> 00:00:24.000covers five cues running00:00:03.288 --> 00:00:24.168).
So: the DOM scrape is trustworthy for content and attribution, and should not be trusted for precise timings. Don't cite DOM-scraped timestamps as exact, and don't use them to align against recordings or other time-series.
Note that cueCount counts recap entries, not VTT cues — 433/433 with
truncated: false is a complete capture even though the real transcript has
1,091 cues. The two numbers are not comparable, and a mismatch is not a bug.
If you verify a capture against a reference .vtt, strip the cue-identifier
lines first. Official Teams exports put a GUID id line (e.g.
63fe10ca-…-152225ba9c70/7-1) before each timestamp; a naive parser counts
those as transcript text and will report a large phantom word deficit
(~7 tokens × cue count). This has already produced one false "41% of content
is missing" alarm.
Meeting chat capture (chat.md)
After a successful transcript capture, also grab the meeting chat — the Teams conversation attached to the same meeting. It routinely carries what the transcript can't: links pasted mid-call, files shared, and the follow-up traffic that lands in the hours or days afterwards.
This is an agent step. scrape-transcript.mjs does not do it — the script
only handles the transcript pane. Do it after the script returns ok: true.
1. Resolve the meeting chat id.
workiq_search_chatswith the meeting subject is usually enough. Meeting chats come back withchatType: "meeting"and an id shaped19:meeting_<base64>@thread.v2.- Sanity-check the returned
membersagainst the calendar event's attendees before using it. If the subject is generic and several chats match, confirm viam_ask_user— never guess. - If no chat exists (common for meetings nobody typed in), say so and move on. A missing chat is not a failure of the run.
2. Export it to <out>/chat.md.
Follow the /teams-export rendering conventions so it reads identically to
everything under 00 - Chats\: ## for the date (YYYY-MM-DD), ### for each
message (HH:mm — Author), replies as nested blockquotes, attachments
collapsed to a single **Attachments:** line (never download binaries),
message HTML converted to markdown, timestamps in America/Los_Angeles 24-hour,
and --- between root messages.
Front matter:
---
type: reference
subtype: meeting-chat
date: <meeting date, YYYY-MM-DD>
chat_id: "19:meeting_...@thread.v2"
topic: <meeting subject>
captured: <YYYY-MM-DD>
---
Export the whole chat rather than a date window — meeting chats are small, and the valuable part is often the post-meeting tail.
3. Emit a sources.yaml suggestion — do NOT write it.
Most meeting chats are dead within a day, and chat.md is the right home for
those: self-contained, sitting next to the transcript it belongs to. But some
keep producing follow-up traffic for days, and those are worth promoting to a
standing source so /teams-export keeps pulling them.
That promotion is the user's call, not this skill's. sources.yaml is
hand-maintained and /teams-export is explicitly forbidden from writing to it;
this skill must not write to it either. Instead, print a ready-to-paste block
and let the user decide:
# Suggested addition to D:\Repos\Obsidian\00 - Chats\sources.yaml (groups:)
- alias: <kebab-case-slug>
type: group
chat_id: "19:meeting_...@thread.v2"
topic: <meeting subject>
notes: >-
Added <YYYY-MM-DD>. Meeting chat for the <date> <subject> session.
Transcript captured under "02 - Meetings/<date> - <subject>".
Retire with enabled:false once follow-up traffic stops.
Recommend promotion when the meeting has a live follow-up thread — handoffs,
access grants, artifacts still to be shared, or a recurring series. Recommend
against it for one-off sessions that concluded in the room; chat.md already
covers those, and every dead entry left in sources.yaml is re-polled on every
future export run.
Guardrails
.pw-profile/and any captured cookies are credentials. Never echo them, never copy them outside the skill dir.- Transcripts may contain sensitive content (MIP labels are not surfaced via the UI capture). Treat any transcript saved this way as at least as sensitive as the calendar event itself. Warn the user before any outbound share.
- Only write inside the
--outdir. Do not modify anything else in the Obsidian vault. This includes00 - Chats\sources.yaml— that file is hand-maintained by the user; suggest additions, never apply them. chat.mdinherits the same sensitivity treatment as the transcript, and can be worse — chats carry pasted links, file names, and customer detail that never got said out loud. Same rule: warn before any outbound share.- If the user supplies a calendar reference that resolves to multiple
events, always confirm via
m_ask_user— never guess. Same for a meeting subject that matches multiple Teams chats.