This program is designed to read content from a Moodle system, sanitize it, and save it as text or publish it to a LLM vector store.
It works by searching for courses, and then extracting the content of the course either into a .csv file or (todo) a vector store or search engine.
To install the Moodle Content API, follow these steps:
- Clone the repository:
git clone https://github.com/brianlmerritt/learning_tools-content-api.git - Install the required dependencies:
pip install -r requirements.txt# Note you are better to install a virtual env
You also need to setup Web Services (REST) and generate a user token
If that user doesn't have full view all courses & categories, restrict requests to course by course or search of courses by pattern instead of find all courses.
Minimum access (verified on live, 2026-08-25): a COURSE VIEWER role for the
token's user is enough for this app - it needs to enter course contexts
without enrolment (moodle/course:view, plus view-hidden capabilities if
hidden courses matter), not to edit anything. A system-level editing teacher
alone is NOT enough (that archetype lacks moodle/course:view), which
presents as "Course or activity not accessible" on every unenrolled course
and an almost-empty full course list.
Configuration is split in two (since 2026-08-25):
.env— credentials and server URLs only (MOODLE_*,UAT_MOODLE_*); use.env_exampleto help.config.yaml(committed) — run settings:use_uat,course_data_dir, and the course filters (idnumber_search,idnumber_search_list,idnumber_list).
Once the configuration is set up, you can run the program using the following command:
python3 get_moodle_courses_data.py
Extracted content lands in a sqlite store: <course_data_dir>/content.sqlite3
(since 2026-08-25, replacing the per-course CSV tree). Tables keep the old CSV
column names (sections, books, pages, labels, urls, files,
folders, blocks, forums), plus:
course— one manifest row per course withcontents_hashandfirst/last_extracted_at. A course whosecore_course_get_contentspayload hashes the same as last time is skipped (no content refetch); pass--forceto re-extract everything.course_modules— the true all-modules union fromget_contents, subsection cms included (the CSV era discarded this).modules— keeps its historical (misleading) name: it is themod_resource_get_resources_by_coursespayload.sections.parent_section_id— owning parent for Moodle 4.5 delegated subsections, empty for plain sections.
build_staging_content.py reads the store and still writes
course_content.csv / activity_content.csv for the curriculum_mapping
loader; cleaned assets still land in <course_data_dir>/<idnumber>/.
Delegated-section names are staged parent-prefixed ("Week 1 - 6th Oct › Monday").
It is transitional: the plan (2026-08-25) is for the staging loader to read
content.sqlite3 directly, after which this script retires.
Still CSV-based (read the pre-sqlite course_data* trees, migrate when next
needed): analyse_course.py, extract_urls.py, analyze_csv.py, and the
sofia experiment's 4b_extract_import.py. NOTE: 4b reads per-course
{idnumber}_modules_with_strands.csv files, which are NO LONGER produced -
migrate 4b to read the store's modules_with_strands table before its next
use.
A helper utility can extract all urls from the activity content.
python3 extract_urls.py
- Pages
- Books
- Files (Course level files)
- Folders (To be tested)
- Labels
- Blocks
- URLs
- Forums
Strand mapping runs at the END of every get_moodle_courses_data.py run, so
the sqlite store is complete at fetch time. Two passes over the store: pass 1
collects the strand-map legends across every stored course, pass 2 rebuilds
each course's rows (pure local computation, no API calls - a full rebuild per
run keeps everything consistent with the freshest legends). The logic lives
in lib/strand_mapping.py (carried verbatim from the superseded
build_modules_with_strands.py, which remains only as the CSV-corpus-era
tool and is not maintained). Note: on learn-uat copies of live courses the
strand-map week table resolves to 0 entries by design - the page HTML links
to LIVE cmids that do not exist in the copy.
Outputs are STORE TABLES ONLY (sqlite-only output, ruled later 2026-08-25 -
no per-course folders or CSVs from the fetch beyond log_events.csv):
modules_with_strands- one row per content item (books, pages, labels, URLs, resources, folders, forums) with section, strand, sub-strand, week andin_strand_map. Replaced per course, so single-course reprocessing never clobbers the estate.courses_with_strands- one row per course: course id/fullname/ shortname/category,course_strand(see below),has_strand_map, item counts. The course-level companion to the per-item table (together they prefigure the planned two-table moodle_content design).
Strand assignment, in priority order:
- Strand Map page links - courses with a "Strand Map" page (the
BVetMed hub estate) resolve items via the map,
in_strand_map = True. - Section inheritance - items sharing a section with strand-mapped content inherit the section's strand.
- Strand-course any-term rule - a course whose idnumber matches the
SRS strand-module dialect (
1VETS01_A_Y_202627-style, or the UBVETMD PVP course) and whose fullname resolves to a known strand IS a strand course (course_strandis set). Inside it, any full medical strand term in an item name or label text flags the item - no adjacency needed. Short abbreviations (ah, cs, pos...) and the word "skin" still require rule 4's adjacency; both restrictions are deliberate false-positive guards. - Term+strand adjacency rule (all courses) - a strand term immediately preceded or followed by the word "strand"/"strands" ("Alimentary Strand", "the POS strand", "Strand: Alimentary").
- Otherwise the
strandcolumn falls back to the content type.
The vocabulary (STRAND_TERMS) is vendored from moodle-local_curricmap
matcher.php::default_rules() synonyms - keep the two in sync. Same
deliberate exclusions as the plugin: end, nma.
Note on file resources: mod_resource/mod_folder contents are files
with no html sibling; for those, is_used means "module visible" (the
file IS the resource), whereas book/page files require their filename to
appear in the prose. Earlier extracts silently dropped files-only groups
entirely - re-extract any course_data captured before 2026-07-27.
- Extract data from Moodle for quizzes (to start)
- Extract study map function if applicable (at RVC it is strand map)
- Output content to RAG api for indexing, vector database, retrieval
- Build text import routines to save sanitised data (plain text & .md format?) with meta data from course, section, module, and study map if applicable
- Set up contributing possibility
- Add LTI & other content via Selenium?
- Add lecture capture
Coming soon