Working notes — Plains Curriculum atlas

Freeform observations, surprises and open questions, newest sections at the bottom. Numbers are as of 2026-09-25 unless noted; regenerate with the pipeline (inventory → extract → paths → curriculum → link → similarity → build_db).

The library

  • SharePoint site PlainsCurriculum, synced via OneDrive shortcut “Plains Curriculum - Documents”. 6 courses (FP, PCV, RU, EM, MSI, RLC), 4,788 files, ~74 GB. OneDrive kept discovering files for hours after the first sync (4,240 → 4,788): early counts undercounted, inventory must be re-run.
  • 74 GB is mostly pptx (~51 GB; embedded media) and mp4 (~18 GB, 95 files). Text is ~107M chars across 137,847 pages/slides. The text corpus is tiny relative to the bytes.
  • Layout: course / Course Build | Resource Collection / AY folder / stream / week / event folder / file. Naming drifts by course and year (AY23-24, AY2024-2025-FP, AY25-26PCVLearningResources, 06-Renal … Rolled from AY23-24 for AY24-25 — the rolled folder is AY23-24 content, confirmed by dated file names).
  • Streams: AMC (Anschutz), FCB (Fort Collins branch), H&S (Health & Society pillar), DOCS (clinical skills pillar), plus “Additional Exploration” and “Pharmacology Drug Lists”.
  • Years present are uneven: AY23-24 only for EM, RLC, RU; AY26-27 materials only for FP (335) and PCV (117) — built ahead. Course Build planning sheets go out to AY27-28.
  • User believes FCB material lives (also? mainly?) in a separate drive location. Not synced on this Mac. This library does hold 1,289 FCB files. Open: where, and is ours partial?

Course Build workbooks are the backbone (big finding)

  • 54 usable workbooks: <course> AY#### Planning.xlsx, … LOs.xlsx, … Granular Build.xlsx. They look like exports/imports for the curriculum management system (and AAMC Curriculum Inventory):
    • Instructional/Assessment Methods vocabularies = AAMC CI standard lists.
    • System vocabulary = USMLE Step 1 content-outline systems.
    • MEPO = 28 Medical Education Program Objectives in 6 domains: MK (knowledge), PC (patient care), CRTY (curiosity/inquiry), IPCS (interpersonal & communication), CMT (commitment/professionalism), LDSP (leadership/systems). Crosswalk to AAMC PCRS / ACGME should be a small hand-built table — ask the curriculum office if one exists.
    • Chief Concern: 53 clinical presentations in the vocabulary, 24 used in these six courses. The curriculum seems organized around them.
  • Granular Build sheets: per-campus, MEA Event ID, material status (Final / Draft / Waiting on Final / Missing / Uploaded / uploaded / Uploaded-A — vocabulary drift; 1,333 blank), attendance, recording. They also carry instructor names + emails — deliberately NOT copied.
  • Parsing gotchas: column LO Description. (trailing period); sheets under revision have LO Description 2026 + LO Description 2027 (take newest); “Course Director Review” / “Content Director Review” columns show a review workflow; files named “(Do Not Use)”, “OLD Version”, “forDemo”, “StateSave” skipped.
  • Objectives are authored ONCE against AMC event titles. FCB events share them (by title twin). No FCB-titled objective exists.

What the objectives say (AY25-26: 3,073 LOs, 560 events)

  • MK1.1 (knowledge) tags 2,317 LOs (~75%). Everything else is thin: PC2.2 physical exam 116, CRTY3.x ~250, LDSP6.2 53, IPCS4.1 32 … Some MEPOs are barely touched in these six courses (PC2.4 4, PC2.7 2, CMT5.1 1, CMT5.5 1). Expected for a preclinical block, but worth showing.
  • Bloom’s: “Understanding” dominates every course; Evaluating/Creating rare (PCV 11+18, RLC 4+7). RU and FP have relatively more Evaluating.
  • Systems: GI (2) and Nervous (50) barely appear — presumably taught in courses not in this library. Health Systems Science is large (349) via H&S.
  • MS disciplines: Physiology, Pathophysiology, Gross Anatomy lead; Behavioral Science nearly absent (8).
  • H&S keywords: Service Learning, Culturally Effective Medicine, LGBT+, Advocacy, Structural Competency, Disability … (note: may be politically sensitive framing for some audiences — present neutrally).

Linking materials to the schedule

  • Event folders mostly mirror event titles (00030-Tubular Transport Na+, Cl- & H2O - AMC ≈ Tubular Transport: Na+, Cl- & H2O - AMC); newer years use YYYYMMDD_HHMM-Title (exact date/time join).
  • H&S files carry session codes: PCV 4-2 EBM… = block 4, week 2; A/B suffixes = separate sessions. H&S events are generic (“Week 2: H&S Small Group”); the topic is in the event description. Large/small group share LOs.
  • DOCS events are generic per week/student group → week-level links only.
  • 89% of resource files link to an event (AMC/FCB 94%, H&S 99%, DOCS 53%). Unlinked: mostly AY23-24 (no planning data), Additional Exploration, drug lists.
  • Bug I made and fixed: stripping campus from the key merged FCB into AMC.

Coverage & reuse

  • AY25-26 events with ≥1 file: Lecture 84%, Case-based 85%, Lab 86%, TBL 100%; Small Group 19%, Review/Feedback 0%, Panel 0%, Exam 0% (good — exams aren’t here). Small-group materials probably live in Canvas or reuse lecture decks.
  • ~190 events marked Uploaded/Final in build sheets have no linked file here (linking miss, or stored elsewhere — e.g. the FCB location).
  • Near-duplicates (MinHash, Jaccard ≥ 0.8): 4,584 texted files → 2,861 distinct materials. Only 33–45% of AY25-26 files are near-copies of an AY24-25 file → most material is revised year to year (or ≥20% edited).
  • FCB vs AMC: only 7–21% of FCB files are near-identical to an AMC file. FCB materials are largely distinct despite shared objectives. Question for leadership: intentional adaptation or duplicated effort?
  • Largest “families” are admin templates (Prework-Event-Description-Template ×65, peer-feedback forms) — tag as forms, not content.
  • “NO ANSWERS” and “WITH ANSWERS” versions of a deck land in the same family (answers are a small fraction of the text). Tutor must select by role tags, not content similarity.

Readiness for AI uses

  • Text extraction: MarkItDown for pptx/xlsx (slide numbers + speaker notes), LiteParse for the rest (real page numbers, legacy .doc/.ppt via LibreOffice). Speaker notes on 1,866 decks / 30,888 slides — a lot of narrative that plain slide text would miss.
  • ~9% of slides and ~8% of PDF pages are image-only (histology, anatomy, radiology): a text-only tutor is weakest exactly there. OCR / multimodal later.
  • Role tags from names: answer_key 403, facilitator 458, restricted 175 (“DO NOT POST”, “Release at 5pm”, “to be projected in room”, “course directors only”), assessment 110 (make-up quizzes). ~3,561 texted files are candidates for student-facing use after excluding those — needs faculty sign-off, filenames aren’t reliable enough alone.
  • Some DOCS files contain individual names (e.g. “DOCS Student Case_…_“) — possibly students → FERPA. Exclude from tutoring; never show names in reports.
  • Access model agreed with user: links go to SharePoint (permissions stay there); extracted text in a DB/search index is fine.

Tooling notes

  • Bakeoff (32 files): MarkItDown only tool with speaker notes; LiteParse 20/20 page counts; Kreuzberg pptx pages broken (1–2 “pages” per deck, 0-based); Docling slow, drops empty PDF pages. LiteParse needs_ocr fires on “sparse-text” pages → use own rule.
  • MarkItDown crashes on unrecognized pptx shape types (2 decks) → LiteParse fallback. LiteParse parse_timeout requires pool_size.
  • OneDrive: reading a cloud-only placeholder triggers download; check SF_DATALESS via lstat. Disk filled once (74 GB library on a nearly full disk) and OneDrive died.

Open questions

  1. Where is the FCB drive, and is it the authoritative FCB copy?
  2. Does the curriculum office have a MEPO → PCRS/EPA crosswalk and the AAMC CI submission?
  3. Which vendors/APIs are approved for sending course text to an LLM?
  4. What are the DOCS files with personal names?
  5. Is “Uploaded-A” a real status or a typo?
  6. Where do small-group, review and panel materials live (Canvas?)

Phase 1–2 findings (2026-09-26)

Details are in docs/research/ (phase1, validity, workbench-landscape) and docs/workbench.md.

  • Keyword search is enough for “where is X taught”, because slides and speaker notes name things. It fails on conceptual queries. Naive MeSH expansion hurts: Frank-Starling pulls in starlings (the birds). MeSH has no “HFrEF”.
  • Objective-list slides fool naive alignment. 444 of 641 “best matches” were the objectives slide itself. Exclude restatements before measuring anything.
  • FCB’s AY25-26 “CONTENT_GCA” decks tag slides with deck-local objective numbers, e.g. “(LO1.1, P4)”: about 1,400 pages across courses. These are author-made alignment evidence. Is the GCA format school-wide? Who makes it?
  • Bloom labels follow wording (87% agree with the leading verb; FP is lowest at 69%). Higher-order labels are rarely backed by in-material activities. 28% of AY25-26 “Evaluating” objectives are one weekly template sentence.
  • Repeated objective text means two different things:
    • weekly format templates, like chief concern (“advance organizer”) and VISTA;
    • content threaded through prework → cases → small group in the same week (RU especially), sometimes under two title variants of the same session.
  • AMC vs FCB (PCV): 41 sessions are separate builds, 16 partly or largely shared, and 33 have no FCB material in this library.
  • Linking bugs fixed:
    • week-range folders (Wk1-5);
    • sessions swapped between time slots after their folders were named (24 date links);
    • H&S A/B sessions (95 files).
  • “Autonomic Regulation of the CV System” (PCV AMC) had 39 slides in AY24-25 and has no AMC material in AY25-26. Missing upload, or stored elsewhere?