Working notes — Plains Curriculum atlas
Freeform observations, surprises and open questions, newest sections at the bottom. Numbers are as of 2026-09-25 unless noted; regenerate with the pipeline (inventory → extract → paths → curriculum → link → similarity → build_db).
The library
- SharePoint site
PlainsCurriculum, synced via OneDrive shortcut “Plains Curriculum - Documents”. 6 courses (FP, PCV, RU, EM, MSI, RLC), 4,788 files, ~74 GB. OneDrive kept discovering files for hours after the first sync (4,240 → 4,788): early counts undercounted, inventory must be re-run. - 74 GB is mostly pptx (~51 GB; embedded media) and mp4 (~18 GB, 95 files). Text is ~107M chars across 137,847 pages/slides. The text corpus is tiny relative to the bytes.
- Layout:
course / Course Build | Resource Collection / AY folder / stream / week / event folder / file. Naming drifts by course and year (AY23-24,AY2024-2025-FP,AY25-26PCVLearningResources,06-Renal … Rolled from AY23-24 for AY24-25— the rolled folder is AY23-24 content, confirmed by dated file names). - Streams: AMC (Anschutz), FCB (Fort Collins branch), H&S (Health & Society pillar), DOCS (clinical skills pillar), plus “Additional Exploration” and “Pharmacology Drug Lists”.
- Years present are uneven: AY23-24 only for EM, RLC, RU; AY26-27 materials only for FP (335) and PCV (117) — built ahead. Course Build planning sheets go out to AY27-28.
- User believes FCB material lives (also? mainly?) in a separate drive location. Not synced on this Mac. This library does hold 1,289 FCB files. Open: where, and is ours partial?
Course Build workbooks are the backbone (big finding)
- 54 usable workbooks:
<course> AY#### Planning.xlsx,… LOs.xlsx,… Granular Build.xlsx. They look like exports/imports for the curriculum management system (and AAMC Curriculum Inventory):- Instructional/Assessment Methods vocabularies = AAMC CI standard lists.
Systemvocabulary = USMLE Step 1 content-outline systems.MEPO= 28 Medical Education Program Objectives in 6 domains: MK (knowledge), PC (patient care), CRTY (curiosity/inquiry), IPCS (interpersonal & communication), CMT (commitment/professionalism), LDSP (leadership/systems). Crosswalk to AAMC PCRS / ACGME should be a small hand-built table — ask the curriculum office if one exists.Chief Concern: 53 clinical presentations in the vocabulary, 24 used in these six courses. The curriculum seems organized around them.
- Granular Build sheets: per-campus, MEA Event ID, material status (Final / Draft / Waiting on Final / Missing / Uploaded / uploaded / Uploaded-A — vocabulary drift; 1,333 blank), attendance, recording. They also carry instructor names + emails — deliberately NOT copied.
- Parsing gotchas: column
LO Description.(trailing period); sheets under revision haveLO Description 2026+LO Description 2027(take newest); “Course Director Review” / “Content Director Review” columns show a review workflow; files named “(Do Not Use)”, “OLD Version”, “forDemo”, “StateSave” skipped. - Objectives are authored ONCE against AMC event titles. FCB events share them (by title twin). No FCB-titled objective exists.
What the objectives say (AY25-26: 3,073 LOs, 560 events)
- MK1.1 (knowledge) tags 2,317 LOs (~75%). Everything else is thin: PC2.2 physical exam 116, CRTY3.x ~250, LDSP6.2 53, IPCS4.1 32 … Some MEPOs are barely touched in these six courses (PC2.4 4, PC2.7 2, CMT5.1 1, CMT5.5 1). Expected for a preclinical block, but worth showing.
- Bloom’s: “Understanding” dominates every course; Evaluating/Creating rare (PCV 11+18, RLC 4+7). RU and FP have relatively more Evaluating.
- Systems: GI (2) and Nervous (50) barely appear — presumably taught in courses not in this library. Health Systems Science is large (349) via H&S.
- MS disciplines: Physiology, Pathophysiology, Gross Anatomy lead; Behavioral Science nearly absent (8).
- H&S keywords: Service Learning, Culturally Effective Medicine, LGBT+, Advocacy, Structural Competency, Disability … (note: may be politically sensitive framing for some audiences — present neutrally).
Linking materials to the schedule
- Event folders mostly mirror event titles (
00030-Tubular Transport Na+, Cl- & H2O - AMC≈Tubular Transport: Na+, Cl- & H2O - AMC); newer years useYYYYMMDD_HHMM-Title(exact date/time join). - H&S files carry session codes:
PCV 4-2 EBM…= block 4, week 2; A/B suffixes = separate sessions. H&S events are generic (“Week 2: H&S Small Group”); the topic is in the event description. Large/small group share LOs. - DOCS events are generic per week/student group → week-level links only.
- 89% of resource files link to an event (AMC/FCB 94%, H&S 99%, DOCS 53%). Unlinked: mostly AY23-24 (no planning data), Additional Exploration, drug lists.
- Bug I made and fixed: stripping campus from the key merged FCB into AMC.
Coverage & reuse
- AY25-26 events with ≥1 file: Lecture 84%, Case-based 85%, Lab 86%, TBL 100%; Small Group 19%, Review/Feedback 0%, Panel 0%, Exam 0% (good — exams aren’t here). Small-group materials probably live in Canvas or reuse lecture decks.
- ~190 events marked Uploaded/Final in build sheets have no linked file here (linking miss, or stored elsewhere — e.g. the FCB location).
- Near-duplicates (MinHash, Jaccard ≥ 0.8): 4,584 texted files → 2,861 distinct materials. Only 33–45% of AY25-26 files are near-copies of an AY24-25 file → most material is revised year to year (or ≥20% edited).
- FCB vs AMC: only 7–21% of FCB files are near-identical to an AMC file. FCB materials are largely distinct despite shared objectives. Question for leadership: intentional adaptation or duplicated effort?
- Largest “families” are admin templates (Prework-Event-Description-Template ×65, peer-feedback forms) — tag as forms, not content.
- “NO ANSWERS” and “WITH ANSWERS” versions of a deck land in the same family (answers are a small fraction of the text). Tutor must select by role tags, not content similarity.
Readiness for AI uses
- Text extraction: MarkItDown for pptx/xlsx (slide numbers + speaker notes), LiteParse for the rest (real page numbers, legacy .doc/.ppt via LibreOffice). Speaker notes on 1,866 decks / 30,888 slides — a lot of narrative that plain slide text would miss.
- ~9% of slides and ~8% of PDF pages are image-only (histology, anatomy, radiology): a text-only tutor is weakest exactly there. OCR / multimodal later.
- Role tags from names: answer_key 403, facilitator 458, restricted 175 (“DO NOT POST”, “Release at 5pm”, “to be projected in room”, “course directors only”), assessment 110 (make-up quizzes). ~3,561 texted files are candidates for student-facing use after excluding those — needs faculty sign-off, filenames aren’t reliable enough alone.
- Some DOCS files contain individual names (e.g. “DOCS Student Case_…_
“) — possibly students → FERPA. Exclude from tutoring; never show names in reports. - Access model agreed with user: links go to SharePoint (permissions stay there); extracted text in a DB/search index is fine.
Tooling notes
- Bakeoff (32 files): MarkItDown only tool with speaker notes; LiteParse 20/20 page counts; Kreuzberg pptx pages broken (1–2 “pages” per deck, 0-based); Docling slow, drops empty PDF pages. LiteParse
needs_ocrfires on “sparse-text” pages → use own rule. - MarkItDown crashes on unrecognized pptx shape types (2 decks) → LiteParse fallback. LiteParse
parse_timeoutrequirespool_size. - OneDrive: reading a cloud-only placeholder triggers download; check
SF_DATALESSvia lstat. Disk filled once (74 GB library on a nearly full disk) and OneDrive died.
Open questions
- Where is the FCB drive, and is it the authoritative FCB copy?
- Does the curriculum office have a MEPO → PCRS/EPA crosswalk and the AAMC CI submission?
- Which vendors/APIs are approved for sending course text to an LLM?
- What are the DOCS files with personal names?
- Is “Uploaded-A” a real status or a typo?
- Where do small-group, review and panel materials live (Canvas?)
Phase 1–2 findings (2026-09-26)
Details are in docs/research/ (phase1, validity, workbench-landscape) and docs/workbench.md.
- Keyword search is enough for “where is X taught”, because slides and speaker notes name things. It fails on conceptual queries. Naive MeSH expansion hurts: Frank-Starling pulls in starlings (the birds). MeSH has no “HFrEF”.
- Objective-list slides fool naive alignment. 444 of 641 “best matches” were the objectives slide itself. Exclude restatements before measuring anything.
- FCB’s AY25-26 “CONTENT_GCA” decks tag slides with deck-local objective numbers, e.g. “(LO1.1, P4)”: about 1,400 pages across courses. These are author-made alignment evidence. Is the GCA format school-wide? Who makes it?
- Bloom labels follow wording (87% agree with the leading verb; FP is lowest at 69%). Higher-order labels are rarely backed by in-material activities. 28% of AY25-26 “Evaluating” objectives are one weekly template sentence.
- Repeated objective text means two different things:
- weekly format templates, like chief concern (“advance organizer”) and VISTA;
- content threaded through prework → cases → small group in the same week (RU especially), sometimes under two title variants of the same session.
- AMC vs FCB (PCV): 41 sessions are separate builds, 16 partly or largely shared, and 33 have no FCB material in this library.
- Linking bugs fixed:
- week-range folders (
Wk1-5); - sessions swapped between time slots after their folders were named (24 date links);
- H&S A/B sessions (95 files).
- week-range folders (
- “Autonomic Regulation of the CV System” (PCV AMC) had 39 slides in AY24-25 and has no AMC material in AY25-26. Missing upload, or stored elsewhere?