dot-skills Graph Database Schema Design Best Practices
Comprehensive graph database data modeling guide for property graphs (Neo4j, Memgraph, Amazon Neptune, etc.). Contains 46 rules across 8 categories, prioritized by modeling impact from critical (entity classification, relationship design) to incremental (scale and evolution). Each rule includes detailed explanations, real-world Cypher examples comparing incorrect vs. correct models, and specific impact descriptions.
Philosophy: Data modeling correctness first, performance second. Always ask "what is the user trying to achieve?" before choosing structure.
When to Apply
Reference these guidelines when:
Designing a new graph database schema from domain requirements
Translating a relational schema to a graph model
Deciding whether something should be a node, relationship, or property
Reviewing an existing graph schema for modeling errors
Refactoring a graph that produces awkward or slow queries
Planning for schema evolution and data growth
Rule Categories by Priority
Priority
Category
Impact
Prefix
1
Entity Classification
CRITICAL
entity-
2
Relationship Design
CRITICAL
rel-
3
Property Placement
HIGH
prop-
4
Query-Driven Refinement
HIGH
query-
5
Structural Patterns
HIGH
pattern-
6
Anti-Patterns
MEDIUM
anti-
7
Constraints & Integrity
MEDIUM
constraint-
8
Scale & Evolution
LOW-MEDIUM
scale-
Quick Reference
1. Entity Classification (CRITICAL)
entity-events - Model multi-participant events as first-class nodes
entity-shared-values - Promote shared property values to nodes
entity-specific-labels - Use specific labels over generic ones
entity-multi-label - Qualify entities with multiple labels
entity-identity-state - Separate identity from mutable state
entity-reify-actions - Reify lifecycle actions into nodes
rel-properties-scope - Put data on relationships only when it describes the connection
rel-single-semantic - One relationship type per semantic meaning
rel-typed-over-filtered - Prefer typed relationships over generic + property filter
3. Property Placement (HIGH)
prop-no-foreign-keys - Don't embed foreign keys as properties
prop-promote-to-node - Promote frequently-queried values to nodes
prop-correct-data-types - Use appropriate data types for properties
prop-no-arrays-for-connections - Don't use property arrays when you need relationships
prop-relationship-vs-node-data - Know when data belongs on relationship vs. node
4. Query-Driven Refinement (HIGH)
query-critical-traversals - Design for your most critical traversals first
query-shortcut-relationships - Add shortcut relationships for frequent multi-hop queries
query-denormalize-reads - Denormalize for read-heavy paths
query-filter-by-rel-props - Use relationship properties to filter traversals
query-test-before-deploy - Test model against real queries before deploying
5. Structural Patterns (HIGH)
pattern-intermediary-nodes - Use intermediary nodes for multi-entity relationships
pattern-hierarchy - Model hierarchies with category nodes and depth relationships
pattern-linked-list - Use linked lists for ordered sequences
pattern-timeline-tree - Apply timeline trees for temporal data
pattern-fan-out - Fan-out pattern for event streams and activity feeds
pattern-bipartite - Use bipartite structure for many-to-many with context
6. Anti-Patterns (MEDIUM)
anti-join-table-nodes - Don't model relational join tables as nodes
anti-generic-relationships - Don't use generic RELATED_TO or CONNECTED relationships
anti-relational-porting - Don't port relational schemas directly to graph
anti-over-modeling - Don't make everything a node
anti-duplicate-data - Don't duplicate data instead of creating relationships
anti-string-encoded-structure - Don't encode structured data as delimited strings
7. Constraints & Integrity (MEDIUM)
constraint-unique-identifiers - Define uniqueness constraints on natural identifiers
constraint-existence - Use existence constraints for required properties
constraint-index-traversals - Create indexes on traversal entry point properties
constraint-no-over-index - Don't over-index — each index has a write cost
constraint-node-key - Use composite node keys for natural multi-part identifiers
8. Scale & Evolution (LOW-MEDIUM)
scale-supernode-mitigation - Mitigate supernodes with fan-out or partitioning
scale-temporal-versioning - Separate current state from historical state
scale-schema-migration - Plan for label and relationship type evolution
scale-batch-refactoring - Use APOC or batched queries for schema refactoring
scale-dense-node-detection - Monitor and detect emerging supernodes
How to Use
Read individual reference files for detailed explanations and code examples:
Section definitions - Category structure and impact levels
Rule template - Template for adding new rules
Reference Files
File
Description
references/_sections.md
Category definitions and ordering
assets/templates/_template.md
Template for new rules
metadata.json
Version and reference information
1---2name: graph-schema3description: dot-skills Graph Database Schema Design Best Practices4---5# dot-skills Graph Database Schema Design Best Practices67Comprehensive graph database data modeling guide for property graphs (Neo4j, Memgraph, Amazon Neptune, etc.). Contains 46 rules across 8 categories, prioritized by modeling impact from critical (entity classification, relationship design) to incremental (scale and evolution). Each rule includes detailed explanations, real-world Cypher examples comparing incorrect vs. correct models, and specific impact descriptions.89**Philosophy:** Data modeling correctness first, performance second. Always ask "what is the user trying to achieve?" before choosing structure.1011## When to Apply1213Reference these guidelines when:14- Designing a new graph database schema from domain requirements15- Translating a relational schema to a graph model16- Deciding whether something should be a node, relationship, or property17- Reviewing an existing graph schema for modeling errors18- Refactoring a graph that produces awkward or slow queries19- Planning for schema evolution and data growth2021## Rule Categories by Priority2223| Priority | Category | Impact | Prefix |24|----------|----------|--------|--------|25| 1 | Entity Classification | CRITICAL | `entity-` |26| 2 | Relationship Design | CRITICAL | `rel-` |27| 3 | Property Placement | HIGH | `prop-` |28| 4 | Query-Driven Refinement | HIGH | `query-` |29| 5 | Structural Patterns | HIGH | `pattern-` |30| 6 | Anti-Patterns | MEDIUM | `anti-` |31| 7 | Constraints & Integrity | MEDIUM | `constraint-` |32| 8 | Scale & Evolution | LOW-MEDIUM | `scale-` |3334## Quick Reference3536### 1. Entity Classification (CRITICAL)3738- [`entity-events`](references/entity-events.md) - Model multi-participant events as first-class nodes39- [`entity-shared-values`](references/entity-shared-values.md) - Promote shared property values to nodes40- [`entity-specific-labels`](references/entity-specific-labels.md) - Use specific labels over generic ones41- [`entity-multi-label`](references/entity-multi-label.md) - Qualify entities with multiple labels42- [`entity-identity-state`](references/entity-identity-state.md) - Separate identity from mutable state43- [`entity-reify-actions`](references/entity-reify-actions.md) - Reify lifecycle actions into nodes44- [`entity-avoid-god-nodes`](references/entity-avoid-god-nodes.md) - Avoid kitchen-sink entity nodes4546### 2. Relationship Design (CRITICAL)4748- [`rel-specific-types`](references/rel-specific-types.md) - Use specific relationship types over generic ones49- [`rel-meaningful-direction`](references/rel-meaningful-direction.md) - Choose semantically meaningful direction50- [`rel-naming-conventions`](references/rel-naming-conventions.md) - Follow UPPER_SNAKE_CASE for relationship types51- [`rel-no-redundant-reverse`](references/rel-no-redundant-reverse.md) - Don't create redundant reverse relationships52- [`rel-properties-scope`](references/rel-properties-scope.md) - Put data on relationships only when it describes the connection53- [`rel-single-semantic`](references/rel-single-semantic.md) - One relationship type per semantic meaning54- [`rel-typed-over-filtered`](references/rel-typed-over-filtered.md) - Prefer typed relationships over generic + property filter5556### 3. Property Placement (HIGH)5758- [`prop-no-foreign-keys`](references/prop-no-foreign-keys.md) - Don't embed foreign keys as properties59- [`prop-promote-to-node`](references/prop-promote-to-node.md) - Promote frequently-queried values to nodes60- [`prop-correct-data-types`](references/prop-correct-data-types.md) - Use appropriate data types for properties61- [`prop-no-arrays-for-connections`](references/prop-no-arrays-for-connections.md) - Don't use property arrays when you need relationships62- [`prop-relationship-vs-node-data`](references/prop-relationship-vs-node-data.md) - Know when data belongs on relationship vs. node6364### 4. Query-Driven Refinement (HIGH)6566- [`query-critical-traversals`](references/query-critical-traversals.md) - Design for your most critical traversals first67- [`query-shortcut-relationships`](references/query-shortcut-relationships.md) - Add shortcut relationships for frequent multi-hop queries68- [`query-denormalize-reads`](references/query-denormalize-reads.md) - Denormalize for read-heavy paths69- [`query-filter-by-rel-props`](references/query-filter-by-rel-props.md) - Use relationship properties to filter traversals70- [`query-test-before-deploy`](references/query-test-before-deploy.md) - Test model against real queries before deploying7172### 5. Structural Patterns (HIGH)7374- [`pattern-intermediary-nodes`](references/pattern-intermediary-nodes.md) - Use intermediary nodes for multi-entity relationships75- [`pattern-hierarchy`](references/pattern-hierarchy.md) - Model hierarchies with category nodes and depth relationships76- [`pattern-linked-list`](references/pattern-linked-list.md) - Use linked lists for ordered sequences77- [`pattern-timeline-tree`](references/pattern-timeline-tree.md) - Apply timeline trees for temporal data78- [`pattern-fan-out`](references/pattern-fan-out.md) - Fan-out pattern for event streams and activity feeds79- [`pattern-bipartite`](references/pattern-bipartite.md) - Use bipartite structure for many-to-many with context8081### 6. Anti-Patterns (MEDIUM)8283- [`anti-join-table-nodes`](references/anti-join-table-nodes.md) - Don't model relational join tables as nodes84- [`anti-generic-relationships`](references/anti-generic-relationships.md) - Don't use generic RELATED_TO or CONNECTED relationships85- [`anti-relational-porting`](references/anti-relational-porting.md) - Don't port relational schemas directly to graph86- [`anti-over-modeling`](references/anti-over-modeling.md) - Don't make everything a node87- [`anti-duplicate-data`](references/anti-duplicate-data.md) - Don't duplicate data instead of creating relationships88- [`anti-string-encoded-structure`](references/anti-string-encoded-structure.md) - Don't encode structured data as delimited strings8990### 7. Constraints & Integrity (MEDIUM)9192- [`constraint-unique-identifiers`](references/constraint-unique-identifiers.md) - Define uniqueness constraints on natural identifiers93- [`constraint-existence`](references/constraint-existence.md) - Use existence constraints for required properties94- [`constraint-index-traversals`](references/constraint-index-traversals.md) - Create indexes on traversal entry point properties95- [`constraint-no-over-index`](references/constraint-no-over-index.md) - Don't over-index — each index has a write cost96- [`constraint-node-key`](references/constraint-node-key.md) - Use composite node keys for natural multi-part identifiers9798### 8. Scale & Evolution (LOW-MEDIUM)99100- [`scale-supernode-mitigation`](references/scale-supernode-mitigation.md) - Mitigate supernodes with fan-out or partitioning101- [`scale-temporal-versioning`](references/scale-temporal-versioning.md) - Separate current state from historical state102- [`scale-schema-migration`](references/scale-schema-migration.md) - Plan for label and relationship type evolution103- [`scale-batch-refactoring`](references/scale-batch-refactoring.md) - Use APOC or batched queries for schema refactoring104- [`scale-dense-node-detection`](references/scale-dense-node-detection.md) - Monitor and detect emerging supernodes105106## How to Use107108Read individual reference files for detailed explanations and code examples:109110- [Section definitions](references/_sections.md) - Category structure and impact levels111- [Rule template](assets/templates/_template.md) - Template for adding new rules112113## Reference Files114115| File | Description |116|------|-------------|117| [references/_sections.md](references/_sections.md) | Category definitions and ordering |118| [assets/templates/_template.md](assets/templates/_template.md) | Template for new rules |119| [metadata.json](metadata.json) | Version and reference information |
Run npx skillmds@latest add comeonoliver/graph-schema in your terminal (requires Node.js), paste this page's agent-chat prompt into Claude, Cursor, or any MCP-connected agent, or download the SKILL.md file and copy it into your agent's skills directory.
dot-skills Graph Database Schema Design Best Practices It is listed under Coding & Dev Tools on SkillMD.
This skill has not completed SkillMD's automated safety review yet. Independent scanners report: SkillSpector: PASS, Skill Scanner: PASS. SkillMD never runs a skill's scripts for you; review the SKILL.md before installing.
This skill is tagged as working with Claude Code, Claude.ai, OpenAI Codex. SKILL.md is an open format, so most agents that read a skills directory can load it too.
Yes. Installing skills from SkillMD is free, and the skill stays under its author's original license.
ComeOnOliver (@comeonoliver) published this skill. Their other Agent Skills are listed on their SkillMD profile.