▗ ▗ ▖ ▖ Claude Code v2.0.36 Sonnet 4.5 · Claude Pro ▘▘ ▝▝ $HOME/Documents/ws/claude-skills
run the chapter-content-generator skill on chapters 6 and 7
⏺ I'll run the chapter-content-generator skill to generate content for chapters 6 and 7.
The "chapter-content-generator" skill is running
The "chapter-content-generator" skill is running
⏺ I'll generate content for chapters 6 and 7. Let me start by finding and verifying these chapter files.
⏺ Bash(ls -la $HOME/Documents/ws/claude-skills/docs/chapters/ | grep -E "^d.*0[67]-") ⎿ drwxr-xr-x 3 dan staff 96 Nov 8 12:07 06-learning-graph-quality-validation drwxr-xr-x 3 dan staff 96 Nov 8 12:07 07-taxonomy-data-formats
⏺ Read(docs/chapters/06-learning-graph-quality-validation/index.md) ⎿ Read 40 lines
⏺ Read(docs/chapters/07-taxonomy-data-formats/index.md) ⎿ Read 46 lines
⏺ Read(docs/course-description.md) ⎿ Read 121 lines
⏺ Read(~/.claude/skills/chapter-content-generator/references/reading-levels.md) ⎿ Read 245 lines
⏺ Read(~/.claude/skills/chapter-content-generator/references/content-element-types.md) ⎿ Read 595 lines
⏺ Now I'll generate content for both chapters. Based on the course description, the target audience is "Professional development" with programming prerequisites, indicating a college-level reading level.
⏺ Update(docs/chapters/06-learning-graph-quality-validation/index.md)
⎿ Updated docs/chapters/06-learning-graph-quality-validation/index.md with 621
additions and 1 removal
36
37 ---
38
39 - TODO: Generate Chapter Content
39 + ## Introduction to Learning Graph Quality Validation
40 +
41 + Creating a learning graph is a significant achievement, but ensuring its
+ quality is equally important for effective educational outcomes. A
+ well-constructed learning graph serves as the foundation for your
+ intelligent textbook, guiding students through concepts in a logical,
+ dependency-aware sequence. Poor quality graphs—those with circular
+ dependencies, orphaned concepts, or imbalanced taxonomy distributions—can
+ confuse learners and undermine the pedagogical value of your materials.
42 +
43 + This chapter introduces systematic approaches for validating and
+ assessing the quality of your learning graph. You'll learn both
+ structural validation techniques that ensure your graph is mathematically
+ sound as a Directed Acyclic Graph (DAG), and quality metrics that
+ measure pedagogical effectiveness. These validation techniques are
+ essential for identifying and correcting issues before generating chapter
+ content, as structural problems in your graph will propagate throughout
+ your entire textbook.
44 +
45 + The validation process combines automated analysis through Python
+ scripts with manual review of quality reports. By the end of this
+ chapter, you'll be able to generate comprehensive quality assessments for
+ your learning graphs and make data-driven improvements to enhance their
+ educational value.
46 +
47 + ## Directed Acyclic Graphs and Educational Dependencies
48 +
49 + Learning graphs must be structured as Directed Acyclic Graphs (DAGs)
+ to represent prerequisite relationships correctly. In a DAG, directed
+ edges point from prerequisite concepts to dependent concepts, and the
+ graph contains no cycles—you cannot follow the dependency arrows and
+ return to your starting concept.
50 +
51 + This DAG structure ensures that students can learn concepts in a valid
+ sequence. If your graph contains a cycle (Concept A depends on B, B
+ depends on C, and C depends on A), there is no valid starting point for
+ learning—a logical impossibility that must be detected and corrected.
52 +
53 + ### DAG Validation
54 +
55 + Validating that your learning graph is a proper DAG involves checking
+ two critical properties:
56 +
57 + 1. Acyclicity: No circular dependency chains exist in the graph
58 + 2. Connectivity: All concepts are reachable from foundational nodes
59 +
60 + The analyze-graph.py Python script performs DAG validation
+ automatically by implementing a depth-first search (DFS) algorithm with
+ cycle detection. During traversal, the algorithm maintains three node
+ states:
61 +
62 + - White (unvisited): Node has not been explored
63 + - Gray (in progress): Node is being explored, currently on the
+ recursion stack
64 + - Black (completed): Node and all its descendants have been fully
+ explored
65 +
66 + If the algorithm encounters a gray node during traversal, it has
+ detected a back edge indicating a cycle. This validation runs in O(V + E)
+ time complexity, where V is the number of vertices (concepts) and E is
+ the number of edges (dependencies).
67 +
68 +
69 + DAG Validation Algorithm Visualization
70 + Type: diagram
71 +
72 + Purpose: Illustrate the three-color DFS algorithm used for cycle
+ detection in learning graphs
73 +
74 + Components to show:
75 + - A sample learning graph with 8 nodes arranged in a network
76 + - Color-coded nodes showing White (gray), Gray (yellow), Black
+ (green)
77 + - Directed edges showing dependencies
78 + - One back edge highlighted in red creating a cycle
79 + - DFS traversal stack shown on the right side
80 + - Traversal order numbered 1-8
81 +
82 + Layout: Network graph on left (70%), DFS stack visualization on
+ right (30%)
83 +
84 + Example nodes:
85 + - Node 1: "Variables" (Black - completed)
86 + - Node 2: "Functions" (Black - completed)
87 + - Node 3: "Loops" (Gray - in progress)
88 + - Node 4: "Recursion" (Gray - in progress)
89 + - Node 5: "Data Structures" (White - unvisited)
90 + - Node 6: "Algorithms" (White - unvisited)
91 +
92 + Edges:
93 + - Black arrows: Valid forward edges
94 + - Red arrow: Back edge from "Recursion" to "Loops" (cycle detected!)
95 +
96 + Annotations:
97 + - Arrow pointing to red edge: "Cycle detected: Loops ← Recursion ←
+ Loops"
98 + - Stack showing: [Loops, Recursion]
99 +
100 + Style: Network diagram with color-coded nodes and directional arrows
101 +
102 + Implementation: SVG diagram with color-coded circles and arrows
103 +
104 +
105 + ### Circular Dependency Detection
106 +
107 + Circular dependencies represent the most critical structural flaw in a
+ learning graph. They create logical impossibilities in the learning
+ sequence and must be identified and eliminated before proceeding with
+ content generation.
108 +
109 + Common sources of circular dependencies include:
110 +
111 + - Bidirectional prerequisites: Concept A requires B, and B requires
+ A
112 + - Multi-hop cycles: A requires B, B requires C, C requires A
113 + - Self-dependencies: A concept incorrectly lists itself as a
+ prerequisite
114 +
115 + The analyze-graph.py script reports all cycles found, displaying the
+ complete dependency chain for each cycle. This detailed output allows you
+ to identify which dependency link to remove to break the cycle.
116 +
117 + Here's an example of cycle detection output:
118 +
119 + 120 + CYCLE DETECTED: 121 + Graph Databases (ID: 45) 122 + → Query Performance (ID: 52) 123 + → Index Selection (ID: 48) 124 + → Database Design (ID: 44) 125 + → Graph Databases (ID: 45) 126 + 127 + Recommendation: Remove dependency "Database Design → Graph Databases" 128 +
129 +
130 + ### Self-Dependency Checking
131 +
132 + Self-dependencies occur when a concept incorrectly lists its own
+ ConceptID in its dependencies column. While technically a special case of
+ circular dependencies, self-dependencies are so common—often resulting
+ from copy-paste errors in CSV editing—that the validation script checks
+ for them explicitly before running the general cycle detection algorithm.
133 +
134 + The self-dependency check is trivial but essential:
135 +
136 + python 137 + for concept in learning_graph: 138 + if concept.id in concept.dependencies: 139 + report_error(f"Concept {concept.id} depends on itself") 140 +
141 +
142 + Any self-dependencies detected indicate data entry errors that should be
+ corrected immediately in your learning-graph.csv file.
143 +
144 + ## Quality Metrics for Learning Graphs
145 +
146 + Beyond structural validation, effective learning graphs exhibit certain
+ quality characteristics that enhance their pedagogical value. Quality
+ metrics quantify these characteristics, providing objective measures for
+ assessing and comparing learning graphs.
147 +
148 + The following metrics help identify potential issues that, while not
+ structurally invalid, may indicate pedagogical problems or opportunities
+ for improvement.
149 +
150 + ### Orphaned Nodes
151 +
152 + An orphaned node is a concept that no other concept depends upon—it
+ has an outdegree of zero. While terminal concepts (endpoints in the
+ learning journey) naturally have no dependents, excessive orphaned nodes
+ suggest concepts that may be:
153 +
154 + - Too specialized or advanced for the course scope
155 + - Improperly isolated from the main learning progression
156 + - Missing their dependent concepts due to incomplete graph construction
157 +
158 + A well-designed learning graph typically has 5-10% orphaned nodes,
+ representing culminating concepts and specialized topics. If more than
+ 20% of your concepts are orphaned, review them to determine whether they
+ should be connected to later material or removed from the graph entirely.
159 +
160 +
161 + Orphaned Nodes Identification Chart
162 + Type: chart
163 +
164 + Chart type: Scatter plot
165 +
166 + Purpose: Visualize concept connectivity by showing indegree vs
+ outdegree for all concepts, highlighting orphaned nodes
167 +
168 + X-axis: Indegree (number of prerequisites, 0-8)
169 + Y-axis: Outdegree (number of dependents, 0-12)
170 +
171 + Data series:
172 + 1. Foundational concepts (green dots, indegree = 0, outdegree > 0)
173 + - Example: "Introduction to Learning Graphs" (0, 8)
174 + - Example: "What is a Concept?" (0, 6)
175 +
176 + 2. Intermediate concepts (blue dots, indegree > 0, outdegree > 0)
177 + - Scatter of 150+ points representing well-connected concepts
178 + - Example: "DAG Validation" (2, 4)
179 +
180 + 3. Orphaned concepts (red dots, indegree > 0, outdegree = 0)
181 + - Example: "Advanced Quality Metrics" (5, 0)
182 + - Example: "Future of Learning Graphs" (3, 0)
183 + - Show approximately 15-20 red dots
184 +
185 + Title: "Concept Connectivity Analysis: Indegree vs Outdegree"
186 +
187 + Annotations:
188 + - Vertical line at outdegree=0 labeled "Orphaned Zone"
189 + - Horizontal line at indegree=0 labeled "Foundation Zone"
190 + - Callout: "12% orphaned (healthy range: 5-15%)"
191 +
192 + Legend: Position top-right with color coding explanation
193 +
194 + Implementation: Chart.js scatter plot with color-coded point
+ categories
195 +
196 +
197 + ### Disconnected Subgraphs
198 +
199 + A disconnected subgraph is a cluster of concepts isolated from the
+ main learning graph—they have no dependency paths connecting them to
+ foundational concepts. This indicates a serious structural problem:
+ students cannot reach these concepts through the normal learning
+ progression.
200 +
201 + Disconnected subgraphs typically result from:
202 +
203 + - Copy-pasting concept blocks without establishing connections
204 + - Incomplete dependency mapping during graph construction
205 + - Accidental deletion of bridging concepts
206 +
207 + The analyze-graph.py script uses a connectivity analysis algorithm to
+ identify all disconnected components. In a valid learning graph, there
+ should be exactly one connected component containing all concepts. Any
+ additional components indicate isolated concept clusters that need to be
+ integrated into the main graph.
208 +
209 + ### Linear Chain Detection
210 +
211 + A linear chain is a sequence of concepts where each concept depends
+ on exactly one predecessor and is depended upon by exactly one successor,
+ forming a single-file progression. While some linear sequences are
+ natural (basic → intermediate → advanced), excessive linear chains
+ indicate missed opportunities for:
212 +
213 + - Parallel learning paths that students could explore in different
+ orders
214 + - Cross-concept connections that reinforce understanding
215 + - Flexible curriculum that accommodates different learning styles
216 +
217 + Linear chains are identified by checking each concept's indegree and
+ outdegree:
218 +
219 + python 220 + def is_linear_chain_node(concept): 221 + return concept.indegree == 1 and concept.outdegree == 1 222 +
223 +
224 + Quality learning graphs typically have 20-40% of concepts in linear
+ chains, with the remainder providing branching paths and concept
+ integration points. If more than 60% of concepts form linear chains,
+ consider adding cross-dependencies to create a richer learning network.
225 +
226 +
227 + Linear Chain vs Network Structure Comparison
228 + Type: diagram
229 +
230 + Purpose: Compare linear chain structure (poor) with network
+ structure (good) for learning graphs
231 +
232 + Layout: Two side-by-side network diagrams
233 +
234 + Left diagram - "Linear Chain Structure (Poor)":
235 + - 10 concepts arranged vertically
236 + - Single path: Concept 1 → 2 → 3 → 4 → 5 → 6 → 7 → 8 → 9 → 10
237 + - All nodes colored orange
238 + - Title: "Linear Chain: 100% of concepts in single path"
239 + - Caption: "No flexibility, single learning route"
240 +
241 + Right diagram - "Network Structure (Good)":
242 + - Same 10 concepts arranged in a network
243 + - Multiple paths and connections:
244 + - Concept 1 (foundation) connects to 2, 3, 4
245 + - Concepts 2, 3, 4 are parallel (same level)
246 + - Concept 5 depends on 2 and 3
247 + - Concept 6 depends on 3 and 4
248 + - Concepts 7, 8 depend on various combinations
249 + - Concepts 9, 10 are terminal (culminating concepts)
250 + - Nodes colored by depth: green (foundation), blue (intermediate),
+ purple (advanced)
251 + - Title: "Network Structure: 40% linear, 60% networked"
252 + - Caption: "Multiple paths, cross-concept integration"
253 +
254 + Visual style: Network diagrams with nodes as circles, directed
+ arrows showing dependencies
255 +
256 + Annotations:
257 + - Left: Red "X" indicating poor structure
258 + - Right: Green checkmark indicating good structure
259 + - Arrow between diagrams showing "Refactor to add
+ cross-dependencies"
260 +
261 + Color scheme: Orange for linear, green/blue/purple gradient for
+ network depth
262 +
263 + Implementation: SVG network diagram with positioned nodes and edges
264 +
265 +
266 + ## Graph Analysis Metrics
267 +
268 + Quantitative metrics provide objective measures of graph structure and
+ complexity. These metrics help you understand your learning graph's
+ characteristics and compare it to best practices for educational graph
+ design.
269 +
270 + ### Indegree and Outdegree Analysis
271 +
272 + Indegree (number of prerequisites) and outdegree (number of
+ dependents) are fundamental graph metrics that reveal concept roles
+ within the learning progression:
273 +
274 + - High indegree: Advanced concepts requiring substantial prior
+ knowledge
275 + - Low indegree (0): Foundational concepts accessible without
+ prerequisites
276 + - High outdegree: Core concepts that enable many subsequent topics
277 + - Low outdegree (0): Specialized or terminal concepts
278 +
279 + Distribution of indegree values across your learning graph indicates its
+ prerequisite structure:
280 +
281 + | Indegree | Interpretation | Typical % of Concepts |
282 + |----------|----------------|----------------------|
283 + | 0 | Foundational concepts | 5-10% |
284 + | 1-2 | Early concepts with minimal prerequisites | 30-40% |
285 + | 3-5 | Intermediate concepts requiring solid foundation | 40-50% |
286 + | 6+ | Advanced concepts requiring extensive background | 5-15% |
287 +
288 + If your graph has too many high-indegree concepts (>20% with indegree ≥
+ 6), consider whether some prerequisites are redundant or if the course
+ scope is too advanced. Conversely, if most concepts have indegree 0-1,
+ you may be missing important prerequisite relationships.
289 +
290 + ### Average Dependencies Per Concept
291 +
292 + The average dependencies per concept metric indicates overall graph
+ connectivity and curriculum density:
293 +
294 + 295 + Average Dependencies = Total Edges / Total Nodes 296 +
297 +
298 + For educational learning graphs, empirical research suggests optimal
+ ranges:
299 +
300 + - 2.0-3.0: Appropriate for introductory courses with linear
+ progressions
301 + - 3.0-4.0: Ideal for intermediate courses with moderate integration
302 + - 4.0-5.0: Suitable for advanced courses with high concept
+ integration
303 + - >5.0: May indicate over-specification of prerequisites
304 +
305 + The analyze-graph.py script calculates this metric and flags values
+ outside the recommended 2.0-4.5 range. Graphs with average dependencies
+ below 2.0 may be too linear, while those above 5.0 may impose unrealistic
+ prerequisite burdens on learners.
306 +
307 +
308 + Average Dependencies Distribution Bar Chart
309 + Type: chart
310 +
311 + Chart type: Histogram (bar chart)
312 +
313 + Purpose: Show distribution of prerequisite counts across all
+ concepts in the learning graph
314 +
315 + X-axis: Number of prerequisites (0, 1, 2, 3, 4, 5, 6, 7, 8+)
316 + Y-axis: Number of concepts
317 +
318 + Data (example for 200-concept graph):
319 + - 0 prerequisites: 12 concepts (foundational)
320 + - 1 prerequisite: 45 concepts
321 + - 2 prerequisites: 58 concepts
322 + - 3 prerequisites: 42 concepts
323 + - 4 prerequisites: 25 concepts
324 + - 5 prerequisites: 12 concepts
325 + - 6 prerequisites: 4 concepts
326 + - 7 prerequisites: 2 concepts
327 + - 8+ prerequisites: 0 concepts
328 +
329 + Title: "Prerequisite Distribution Across Learning Graph"
330 +
331 + Calculated metrics displayed below chart:
332 + - Total concepts: 200
333 + - Total dependencies: 620
334 + - Average dependencies: 3.1 per concept
335 + - Median: 2
336 + - Mode: 2
337 +
338 + Annotations:
339 + - Shaded region (2-4 prerequisites) in light green labeled "Optimal
+ Range"
340 + - Average line (vertical) at 3.1 in blue
341 + - Callout: "84% of concepts in optimal range (1-5 prerequisites)"
342 +
343 + Color scheme: Gold bars with green shading for optimal range
344 +
345 + Implementation: Chart.js bar chart with annotations
346 +
347 +
348 + ### Maximum Dependency Chain Length
349 +
350 + The maximum dependency chain length represents the longest sequence
+ of prerequisite concepts from any foundational node to any terminal node.
+ This metric indicates the depth of your curriculum and affects course
+ duration planning.
351 +
352 + For a 200-concept learning graph, typical maximum chain lengths are:
353 +
354 + - 8-12 concepts: Short course (4-6 weeks)
355 + - 12-18 concepts: Standard semester course (12-15 weeks)
356 + - 18-25 concepts: Extended course or multi-semester sequence
357 + - >25 concepts: May indicate overly linear structure
358 +
359 + The chain length affects student progress velocity. If your maximum
+ chain is 20 concepts deep, students must complete at least 20 learning
+ steps to reach the most advanced material—establishing a minimum time
+ investment regardless of study intensity.
360 +
361 + Critical path analysis identifies these longest chains, helping you
+ understand pacing requirements and potential bottlenecks in the learning
+ progression. Concepts on the critical path deserve extra attention in
+ content development, as delays in mastering these concepts cascade
+ through all dependent material.
362 +
363 + ## Learning Graph Quality Score
364 +
365 + The overall learning graph quality score provides a single metric
+ (0-100) that aggregates multiple quality dimensions into an interpretable
+ assessment. While individual metrics reveal specific issues, the quality
+ score enables quick comparison and tracking of improvements over time.
366 +
367 + The quality scoring algorithm used by analyze-graph.py weights various
+ factors:
368 +
369 + Structural Validity (40 points):
370 +
371 + - DAG validation passes (20 points)
372 + - No self-dependencies (10 points)
373 + - All concepts in single connected component (10 points)
374 +
375 + Connectivity Quality (30 points):
376 +
377 + - Orphaned nodes 5-15% of total (10 points, scaled for deviation)
378 + - Average dependencies 2.5-4.0 per concept (10 points, scaled)
379 + - Maximum chain length appropriate for scope (10 points)
380 +
381 + Distribution Quality (20 points):
382 +
383 + - No linear chains exceeding 20% of graph (10 points)
384 + - Indegree distribution follows expected pattern (10 points)
385 +
386 + Taxonomy Balance (10 points):
387 +
388 + - No single taxonomy category exceeds 30% (5 points)
389 + - At least 5 taxonomy categories represented (5 points)
390 +
391 + Interpretation of quality scores:
392 +
393 + | Score Range | Quality Level | Interpretation |
394 + |-------------|--------------|----------------|
395 + | 90-100 | Excellent | Publication-ready, well-structured graph |
396 + | 75-89 | Good | Minor improvements recommended |
397 + | 60-74 | Acceptable | Several issues to address before content
+ generation |
398 + | 40-59 | Poor | Significant structural or quality problems |
399 + | 0-39 | Critical | Major revision required |
400 +
401 + The quality score should be calculated after every significant graph
+ revision. Track scores over time to ensure your changes improve rather
+ than degrade graph quality.
402 +
403 +
404 + Learning Graph Quality Score Calculator MicroSim
405 + Type: microsim
406 +
407 + Learning objective: Allow students to experiment with how different
+ graph characteristics affect overall quality score
408 +
409 + Canvas layout (900x600px):
410 + - Left side (600x600): Quality score visualization
411 + - Right side (300x600): Interactive controls
412 +
413 + Visual elements (left panel):
414 + - Large circular gauge showing overall score (0-100)
415 + - Color-coded segments: Red (0-39), Orange (40-59), Yellow (60-74),
+ Light Green (75-89), Dark Green (90-100)
416 + - Current score displayed in center in large font
417 + - Four horizontal bars below gauge showing component scores:
418 + * Structural Validity: 0-40 points (blue bar)
419 + * Connectivity Quality: 0-30 points (green bar)
420 + * Distribution Quality: 0-20 points (orange bar)
421 + * Taxonomy Balance: 0-10 points (purple bar)
422 + - Each bar shows points earned out of maximum
423 +
424 + Interactive controls (right panel):
425 + - Slider: "Number of Concepts" (50-300, default 200)
426 + - Slider: "Orphaned Nodes %" (0-40%, default 10%)
427 + - Slider: "Avg Dependencies" (1.0-6.0, default 3.2)
428 + - Slider: "Max Chain Length" (5-35, default 16)
429 + - Slider: "Linear Chain %" (10-80%, default 35%)
430 + - Slider: "Largest Taxonomy %" (10-60%, default 22%)
431 + - Checkbox: "Has Cycles" (default unchecked)
432 + - Checkbox: "Has Disconnected Subgraphs" (default unchecked)
433 + - Button: "Reset to Defaults"
434 + - Button: "Load Example: Poor Graph"
435 + - Button: "Load Example: Excellent Graph"
436 +
437 + Default parameters (Good Graph):
438 + - Concepts: 200
439 + - Orphaned: 10%
440 + - Avg Dependencies: 3.2
441 + - Max Chain: 16
442 + - Linear Chain %: 35%
443 + - Largest Taxonomy: 22%
444 + - No cycles, no disconnected subgraphs
445 + - Expected Score: 82 (Good)
446 +
447 + Behavior:
448 + - Real-time recalculation as sliders move
449 + - Score gauge animates to new value
450 + - Component bars update proportionally
451 + - Color of gauge changes based on score range
452 + - Tooltip on hover shows calculation details for each component
453 + - "Poor Graph" example: cycles=true, orphaned=35%, score28
454 + - "Excellent Graph" example: optimal all parameters, score96
455 +
456 + Implementation notes:
457 + - Use p5.js for rendering gauge and bars
458 + - Implement scoring algorithm matching analyze-graph.py logic
459 + - Use DOM elements for sliders and checkboxes
460 + - Map() function to scale slider values to score components
461 + - Lerp() for smooth score animations
462 +
463 + Implementation: p5.js MicroSim with interactive controls
464 +
465 +
466 + ## Taxonomy Distribution and Balance
467 +
468 + Beyond graph structure, the distribution of concepts across taxonomy
+ categories affects curriculum balance and learning progression. A
+ well-balanced taxonomy distribution ensures students encounter
+ appropriate variety across knowledge domains without over-concentration
+ in any single area.
469 +
470 + ### Taxonomy Categories
471 +
472 + Learning graphs typically categorize concepts using a TaxonomyID
+ field that groups related concepts into domains. Common taxonomy
+ categories for technical courses include:
473 +
474 + - FOUND - Foundational concepts and definitions
475 + - BASIC - Basic principles and core ideas
476 + - ARCH - Architecture and system design
477 + - IMPL - Implementation and practical skills
478 + - TOOL - Tools and technologies
479 + - SKILL - Professional skills and practices
480 + - ADV - Advanced topics and specializations
481 +
482 + The number and specificity of taxonomy categories varies by subject
+ matter. Introductory courses might use 5-8 broad categories, while
+ specialized courses might employ 10-15 granular categories.
483 +
484 + ### TaxonomyID Abbreviations
485 +
486 + TaxonomyIDs use 3-5 letter abbreviations for compactness in CSV files
+ and visualization color-coding. When designing your taxonomy, choose
+ abbreviations that are:
487 +
488 + - Distinctive: No two categories should share the same first 3
+ letters
489 + - Mnemonic: Abbreviation should suggest the full category name
490 + - Consistent: Use similar grammatical forms (nouns vs. adjectives)
491 +
492 + Example taxonomy abbreviations:
493 +
494 + | TaxonomyID | Full Category Name | Color Code (visualization) |
495 + |------------|-------------------|---------------------------|
496 + | FOUND | Foundational Concepts | Red |
497 + | BASIC | Basic Principles | Orange |
498 + | ARCH | Architecture & Design | Yellow |
499 + | IMPL | Implementation | Light Green |
500 + | DATA | Data Management | Green |
501 + | TOOL | Tools & Technologies | Light Blue |
502 + | QUAL | Quality Assurance | Blue |
503 + | ADV | Advanced Topics | Purple |
504 +
505 + ### Category Distribution Analysis
506 +
507 + The category distribution metric shows what percentage of your total
+ concepts fall into each taxonomy category. This distribution should
+ reflect the emphasis and scope of your course.
508 +
509 + Healthy category distributions typically exhibit:
510 +
511 + - No single category exceeds 30%: Avoid over-concentration
512 + - Top 3 categories contain 50-70% of concepts: Natural emphasis
+ areas
513 + - At least 5 categories represented: Adequate coverage breadth
514 + - Foundational category: 5-10% of concepts: Appropriate base layer
515 +
516 + The taxonomy-distribution.py script generates a detailed report
+ showing both absolute counts and percentages for each category, enabling
+ quick identification of imbalanced distributions.
517 +
518 +
519 + Taxonomy Distribution Pie Chart
520 + Type: chart
521 +
522 + Chart type: Pie chart with percentage labels
523 +
524 + Purpose: Visualize the distribution of 200 concepts across taxonomy
+ categories
525 +
526 + Data:
527 + - FOUND (Foundational): 18 concepts (9%) - Red
528 + - BASIC (Basic Principles): 42 concepts (21%) - Orange
529 + - ARCH (Architecture): 38 concepts (19%) - Yellow
530 + - IMPL (Implementation): 35 concepts (17.5%) - Light Green
531 + - DATA (Data Management): 28 concepts (14%) - Green
532 + - TOOL (Tools): 22 concepts (11%) - Light Blue
533 + - QUAL (Quality): 12 concepts (6%) - Blue
534 + - ADV (Advanced): 5 concepts (2.5%) - Purple
535 +
536 + Title: "Learning Graph Taxonomy Distribution (200 Concepts)"
537 +
538 + Label format: "CATEGORY: N concepts (P%)"
539 +
540 + Annotations:
541 + - Callout for BASIC slice: "Largest category: 21% (healthy)"
542 + - Callout for ADV slice: "Smallest category: 2.5% (may need
+ expansion)"
543 + - Legend positioned to right side
544 +
545 + Quality indicators:
546 + - Green checkmark: "No category exceeds 30% ✓"
547 + - Green checkmark: "8 categories represented ✓"
548 + - Green checkmark: "Top 3 categories = 59% ✓"
549 +
550 + Color scheme: Rainbow gradient (red → orange → yellow → green → blue
+ → purple)
551 +
552 + Implementation: Chart.js pie chart with custom colors and labels
553 +
554 +
555 + ### Avoiding Over-Representation
556 +
557 + Over-representation occurs when a single taxonomy category dominates
+ the learning graph, consuming more than 30% of total concepts. This
+ imbalance can result from:
558 +
559 + - Scope creep: Course expanded in one area without proportional
+ breadth
560 + - Expert bias: Instructor's specialization over-emphasized
561 + - Incomplete mapping: Other categories insufficiently developed
562 +
563 + Over-representation in foundational or basic categories suggests the
+ course may be too introductory, while over-representation in advanced or
+ specialized categories indicates potential accessibility issues for
+ learners.
564 +
565 + To correct over-representation:
566 +
567 + 1. Review over-represented category: Identify concepts that could be
+ consolidated or removed
568 + 2. Expand under-represented categories: Add concepts to balance
+ distribution
569 + 3. Reclassify borderline concepts: Move concepts to more appropriate
+ categories
570 + 4. Validate against learning outcomes: Ensure distribution aligns
+ with stated course objectives
571 +
572 + The taxonomy distribution report generated by taxonomy-distribution.py
+ flags any categories exceeding the 30% threshold, enabling quick
+ identification of balance issues.
573 +
574 + ## Generating Quality Reports with Python Scripts
575 +
576 + The learning graph quality validation process relies on three Python
+ scripts located in the docs/learning-graph/ directory. These scripts
+ analyze your learning-graph.csv file and generate comprehensive quality
+ reports in Markdown format.
577 +
578 + ### analyze-graph.py Script
579 +
580 + The analyze-graph.py script performs comprehensive graph validation
+ and quality analysis:
581 +
582 + Usage:
583 + bash 584 + cd docs/learning-graph 585 + python analyze-graph.py learning-graph.csv quality-metrics.md 586 +
587 +
588 + Checks performed:
589 +
590 + 1. CSV format validation
591 + 2. Self-dependency detection
592 + 3. Cycle detection (DAG validation)
593 + 4. Connectivity analysis
594 + 5. Orphaned node identification
595 + 6. Linear chain detection
596 + 7. Indegree/outdegree statistics
597 + 8. Maximum dependency chain calculation
598 + 9. Overall quality score computation
599 +
600 + Output: Generates quality-metrics.md report file containing all
+ findings, metrics, and a final quality score. Any critical issues
+ (cycles, disconnected subgraphs) are highlighted at the top of the
+ report.
601 +
602 + ### csv-to-json.py Script
603 +
604 + The csv-to-json.py script converts your learning graph CSV to
+ vis-network JSON format for visualization:
605 +
606 + Usage:
607 + bash 608 + cd docs/learning-graph 609 + python csv-to-json.py learning-graph.csv learning-graph.json 610 +
611 +
612 + Functionality:
613 +
614 + - Parses CSV with ConceptID, ConceptLabel, Dependencies, TaxonomyID
+ columns
615 + - Generates nodes array with id, label, and group (taxonomy) fields
616 + - Generates edges array with from and to fields (dependency arrows)
617 + - Adds metadata section with graph statistics
618 + - Validates JSON output format
619 +
620 + Output: Creates learning-graph.json file that can be l
…(truncated)