Designs MongoDB schemas, indexes, and aggregation pipelines that perform - embed-vs-reference decision rules, the 16MB document limit, ESR compound-index ordering, explain-plan verification, and operational settings for replica sets and sharding. Use when someone asks "should I embed or reference this", "why is my Mongo query slow", "design a MongoDB schema for X", "how do I structure this aggregation pipeline", or "what should my shard key be". Do NOT use for relational or SQL schema design - use database-schema instead; for tuning SQL queries use sql-query-optimizer; for Postgres-on-Supabase work use supabase-expert.
Click to play with sound.
---
name: MongoDB Expert
description: Designs MongoDB schemas, indexes, and aggregation pipelines that perform - embed-vs-reference decision rules, the 16MB document limit, ESR compound-index ordering, explain-plan verification, and operational settings for replica sets and sharding. Use when someone asks "should I embed or reference this", "why is my Mongo query slow", "design a MongoDB schema for X", "how do I structure this aggregation pipeline", or "what should my shard key be". Do NOT use for relational or SQL schema design - use database-schema instead; for tuning SQL queries use sql-query-optimizer; for Postgres-on-Supabase work use supabase-expert.
---
# MongoDB Expert
Model data, index it, and query it the MongoDB way: around access patterns, not normalization. The costly mistake this skill prevents is a relationally-normalized schema ported into Mongo - a design that forces application-side joins on every read and, in the opposite failure, unbounded embedded arrays that grow until documents hit the 16MB limit and writes start failing in production.
## Operating procedure
Schema first, indexes second, pipelines third - indexes serve queries, and queries are determined by the schema, so working backwards wastes every downstream decision.
### Step 1: Gather inputs
1. The top read and write patterns, ranked by frequency: which entities are fetched together, filtered by what, sorted by what.
2. Cardinality and growth of every one-to-many relationship: bounded ("an order has 5-50 line items") or unbounded ("a user accrues events forever"). Label estimates as guesses.
3. Read/write ratio per collection and expected document counts.
4. Deployment shape: replica set or sharded cluster, and durability requirements.
### Step 2: Decide embed vs reference per relationship
Apply these rules in order; the first that fires decides.
1. Accessed together and the child side is bounded (a known, modest maximum) → embed. One read serves the whole query.… load the full skill through Skill Me