Skip to content

NoSQL Data Modeling Explained

Modeling patterns for document, key-value, wide-column, graph, and vector stores — when to embed, reference, denormalize, or bucket.

NoSQL stores trade fixed schemas and joins for horizontal scale. Each family models relationships its own way, so the access pattern — not normal form — dictates the structure.

Reference table · 28 entries
28 of 28 rows
Document
Nested sub-documents live inside one parent document instead of separate tables.One-to-few relationships read together. An order with its line items. A user with three addresses.
Store a foreign id, not the data. Resolve it with a second read or a $lookup.Large, shared, or independently updated sub-entities. Many-to-many links across collections.
The parent embeds a short array of children; the whole set fits and loads as one document.A handful of phone numbers or preferences on a profile. One read returns parent and children.
Children live in their own collection; each child carries the parent id, or the parent keeps a bounded id array.Sets too large or too independent to embed. Query children by parent id.
The child stores the parent id; the parent never lists children because the set is unbounded.Events, logs, or votes tied to a host or user. Filter children by parent id and paginate.
Documents of different entity types share one collection, told apart by a type field.Comments or notifications attached to many kinds of parent, all queried one way.
A few extreme entities are split out of the shared shape so their size or traffic does not distort the rest.Celebrity accounts whose follower lists or counters would outgrow the normal document.
Each revision is a new document with the same id and an increasing version number; the query picks the latest.Audit trails and edit history where old revisions stay queryable.
Derived values are calculated at write time and stored beside the data they summarize.Totals, ratings, and top-N that would otherwise run an aggregation on every read.
Key-value & wide-column
DynamoDB design term: collapse many entity types into one table using overloaded partition and sort keys.Fetch heterogeneous items in one query without a join.
One key attribute serves many entity types, marked by prefixes such as USER# or ORDER#.Single-table designs. Cluster related items under one partition and fetch them in one query.
A partition key spreads rows across nodes; a sort key orders each partition.Cassandra and DynamoDB. Range scans, ordered logs, hierarchical composite keys.
Stack several attributes into one delimited sort key — country#region#city — so a prefix match walks down the hierarchy.Hierarchies and multi-facet filters. One ordered range query replaces many indexed lookups.
Group events into fixed windows — hour, day, month — under dedicated partitions or tables.Metrics, logs & IoT telemetry. Bounds partition size and ages out cold data.
Duplicate read-heavy fields across records so a query never chases a second lookup.Read-heavy workloads where eventual consistency is tolerable and storage is cheap.
A secondary index built on an attribute that only some items carry; items without it never enter the index.Alternate access paths. Find the sparse subset — open orders, unshipped items — without a full scan.
Write down the exact queries first; the table layout is derived from them, not from entities or normal form.Cassandra and DynamoDB schema work. One table per query family; the data model copies the question.
A query must hit many partitions or nodes in parallel because no key matches the question.A smell to design out, not a goal. Re-key the table so the common query lands on one partition.
Copies are pushed to each reader's view at write time, or gathered and merged at read time.Feeds and timelines. Write-time fan-out for small audiences, read-time for huge ones.
A pre-built query result stored as its own table, kept current as the base changes.Alternate read shapes over one write path. Serves queries the base table cannot answer.
Graph
Each edge is a row or tuple pointing from a source node to a target node.Graph and property stores. Friend-of-friend lookups, routing, recommendations.
Each relationship is a document in its own collection, carrying from and to ids plus link attributes.Graphs in document stores. Attributes on the link itself — rating, role, timestamp.
Every node stores left and right counters from a depth-first walk; a subtree lives inside its parent's range.Read-heavy trees. One BETWEEN query returns a subtree; inserts renumber the walk.
A dedicated table records every ancestor-to-descendant pair, including each node at depth zero.Whole-branch queries at any depth. Simple joins, paid for in storage and maintenance.
Property graphs attach key-value attributes to nodes and edges; RDF states everything as subject-predicate-object triples.Property graphs for rich traversal (Neo4j); RDF for linked open data and ontology exchange.
Vector
A high-dimensional float array sits beside metadata. An ANN index finds nearest neighbors.Semantic search, RAG retrieval, recommendations, and similarity over unstructured content.
Source documents are split — by size, overlap, or structure — before each piece is embedded on its own.RAG retrieval quality. Chunk size trades precision against recall; overlap protects split sentences.
Dense vector similarity runs beside keyword matching; the two ranked lists fuse into one result.Queries needing exact terms and meaning at once, with metadata filters on top.