You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Eager-loading for reference fields via depth query option (with future populate extension)
#410
A note from the author: I'm a Japanese developer and not a native English speaker, so I've leaned heavily on AI assistance to write this proposal clearly. The technical content, code references, and design decisions are mine — the AI helped me express them in fluent English. Apologies in advance if any phrasing feels off; please ask and I'll clarify.
Today, reference fields returned by getEmDashCollection() / getEmDashEntry() are plain ID strings (zod-generator.ts:120, fields/reference.ts). Resolving them requires a follow-up getEmDashEntry() call per row, which becomes the textbook N+1 problem when rendering list views (e.g. a blog index that wants the author name, a product list that needs the brand name).
I'd like to propose a small, additive query option — depth — that eagerly hydrates reference fields in batch, with no impact on existing call sites.
Why
Concrete case: I'm migrating a 9-collection / ~30k-entry sakepo product catalog from Strapi to EmDash. The catalog renders an /liquors index page (100 items per page) where each card shows the product name + producer name. With Strapi's populate=producer this is one query. With the current EmDash query API the choices are:
Flatten the producer name onto the liquor row (denormalize) — works but spreads producer state across two tables and requires manual sync on updates.
Fetch all producers then in-memory join client-side — works for small reference tables but bloats payload and breaks down once you have multiple reference fields.
N+1: 1 + 100 round trips per page render. Not viable.
None of these scale, and all of them push complexity into every consumer of the query API. This is the kind of thing the framework should solve once.
EmDash already has hydrateEntryBylines and getBylinesForEntries which solve the exact same problem for byline credits — the proposal is to generalize that pattern to user-defined reference fields.
Proposed API (additive)
Two complementary options on CollectionFilter / getEmDashEntry options:
// Simple: expand all reference fields up to N levelsconst{ entries }=awaitgetEmDashCollection("liquors",{status: "published",depth: 1,});// entries[0].data.producer is now { id, slug, name, country, ... }// Granular (future extension): only expand chosen fieldsconst{ entries }=awaitgetEmDashCollection("liquors",{populate: {producer: true,sake_rice: {populate: {type: true}},},});
Type changes:
exportinterfaceCollectionFilter{// ... existing fields/** * Eagerly resolve reference fields. `depth: 1` expands top-level * references; `depth: 2` also expands references on the resolved * children, and so on. Defaults to 0 (current behavior — IDs only). */depth?: number;// populate?: PopulateSpec; // future extension, not in v1}
Behavior
depth: 0 (default): no behavior change. Reference fields stay as ID strings.
depth: N: walk the result rows, collect every reference-typed field, batch-fetch the targets in a single IN (...) query per target collection (chunked to respect D1's variable limit — see below), substitute objects in place, and recurse N - 1 times on the newly hydrated children.
Missing/dangling references → set to null (matches Strapi/Payload behavior).
Cycles → bounded by depth so no infinite recursion.
Out of scope for v1 (could be discussed/added later):
The populate object form. v1 lands depth only; populate is the granular successor that we can iterate on once depth is in.
Field selection (select: [...]).
Filtering on related fields (where: { producer: { country: "JP" } }).
Implementation sketch (v1: depth only)
The change is concentrated and reuses the existing batch-fetch pattern.
packages/core/src/loader.ts — extend CollectionFilter with depth?: number and thread it through loadCollection / loadEntry. After mapping rows to entries, call a new hydrateReferences(type, entries, depth) helper.
packages/core/src/query.ts — getEmDashCollection / getEmDashEntry accept depth, pass through to the loader, and skip the helper when depth === 0 (zero overhead default path).
New file: packages/core/src/query/hydrate-references.ts:
Read _emdash_fields once via the schema registry to get (collection, field) → { type: "reference", target_collection } map (cached per request via existing getRequestContext()).
Walk all entry data, collect Map<targetCollection, Set<referencedId>>.
For each target collection, run oneSELECT * FROM ec_<target> WHERE id IN (...) per chunk (chunk size = 90 to stay under D1's 100-variable limit, matching the workaround needed for issue #219).
Build Map<id, row> and substitute into the parent data.
If depth > 1, recurse on the hydrated child rows.
No DB schema changes. No migration. Purely a read-side feature.
The whole thing should be ~200-300 lines including tests.
Interaction with existing features
Bylines: orthogonal. hydrateEntryBylines runs after loadCollection, references would too. Order can be parallel (Promise.all) since they touch different tables.
i18n: hydrated children should respect the same locale context as the parent query. Easy via the existing ALS context.
Drafts/preview: hydrated children honor status === "published" unless preview mode is active (same rule as the parent fetch). Dangling refs to draft/unpublished targets become null.
Visual editing (entry.edit): I propose v1 does not wrap hydrated children with edit proxies, since they originate from a different collection. That can be added later if there's demand.
emdash-env.d.ts currently generates reference fields as string (zod-generator.ts:366). For v1 I'd leave that as-is (so existing code keeps compiling) and let consumers cast where they need the hydrated shape:
const{ entries }=awaitgetEmDashCollection("liquors",{depth: 1});constproducer=entries[0].data.producerasProducer;// user knows depth was 1
Improving generated types so depth: 1 returns Liquor & { producer: Producer } is doable but TypeScript-heavy and probably worth its own follow-up PR.
Alternatives considered
GraphQL-style field selection — much more powerful, much more work, and inconsistent with EmDash's REST/loader-first design. Out of scope.
Denormalization at write time — pushes complexity onto every plugin/import script that touches the data and fights the schema-as-source-of-truth design.
Leave it to userland — exactly what I'm doing today, and it doesn't compose. Every page that wants related data has to reimplement batched join logic.
Payload's depth is closest in spirit to what I'm proposing here — minimal API surface, sane default, easy to reason about.
Questions for maintainers
Is the additive depth API the right shape, or do you prefer to start with populate directly? My case for depth first: it's small, opt-in, and unblocks the common case immediately. populate can be layered on without breaking it.
Should hydrated child entries pass through getRequestContext()'s edit-mode wrapping, or stay raw?
Is there appetite for a follow-up that improves the generated types?
What I'm offering
If this gets a thumbs up, I'd open a PR with:
The loader.ts / query.ts changes
A new hydrate-references.ts module mirroring the byline pattern
Tests in packages/core/test/ covering: zero-depth no-op, single ref, multi ref, multi collection, dangling ref, depth=2 recursion, IN-clause chunking, locale propagation
A .changeset/ minor-version entry
Docs update under docs/src/content/docs/guides/querying-content.mdx
Happy to iterate on the API in this thread before writing any code.
(AI disclosure: this proposal was drafted with Claude Code. Implementation, tests, and review will be human-checked against the existing patterns in bylines/index.ts and follow the same style.)
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Summary
Today,
referencefields returned bygetEmDashCollection()/getEmDashEntry()are plain ID strings (zod-generator.ts:120,fields/reference.ts). Resolving them requires a follow-upgetEmDashEntry()call per row, which becomes the textbook N+1 problem when rendering list views (e.g. a blog index that wants the author name, a product list that needs the brand name).I'd like to propose a small, additive query option —
depth— that eagerly hydratesreferencefields in batch, with no impact on existing call sites.Why
Concrete case: I'm migrating a 9-collection / ~30k-entry sakepo product catalog from Strapi to EmDash. The catalog renders an
/liquorsindex page (100 items per page) where each card shows the product name + producer name. With Strapi'spopulate=producerthis is one query. With the current EmDash query API the choices are:None of these scale, and all of them push complexity into every consumer of the query API. This is the kind of thing the framework should solve once.
EmDash already has
hydrateEntryBylinesandgetBylinesForEntrieswhich solve the exact same problem for byline credits — the proposal is to generalize that pattern to user-definedreferencefields.Proposed API (additive)
Two complementary options on
CollectionFilter/getEmDashEntryoptions:Type changes:
Behavior
depth: 0(default): no behavior change. Reference fields stay as ID strings.depth: N: walk the result rows, collect every reference-typed field, batch-fetch the targets in a singleIN (...)query per target collection (chunked to respect D1's variable limit — see below), substitute objects in place, and recurseN - 1times on the newly hydrated children.null(matches Strapi/Payload behavior).depthso no infinite recursion.Out of scope for v1 (could be discussed/added later):
populateobject form. v1 landsdepthonly;populateis the granular successor that we can iterate on oncedepthis in.select: [...]).where: { producer: { country: "JP" } }).Implementation sketch (v1:
depthonly)The change is concentrated and reuses the existing batch-fetch pattern.
packages/core/src/loader.ts— extendCollectionFilterwithdepth?: numberand thread it throughloadCollection/loadEntry. After mapping rows to entries, call a newhydrateReferences(type, entries, depth)helper.packages/core/src/query.ts—getEmDashCollection/getEmDashEntryacceptdepth, pass through to the loader, and skip the helper whendepth === 0(zero overhead default path).packages/core/src/query/hydrate-references.ts:_emdash_fieldsonce via the schema registry to get(collection, field) → { type: "reference", target_collection }map (cached per request via existinggetRequestContext()).Map<targetCollection, Set<referencedId>>.SELECT * FROM ec_<target> WHERE id IN (...)per chunk (chunk size = 90 to stay under D1's 100-variable limit, matching the workaround needed for issue #219).Map<id, row>and substitute into the parentdata.depth > 1, recurse on the hydrated child rows.The whole thing should be ~200-300 lines including tests.
Interaction with existing features
hydrateEntryBylinesruns afterloadCollection, references would too. Order can be parallel (Promise.all) since they touch different tables.localecontext as the parent query. Easy via the existing ALS context.status === "published"unless preview mode is active (same rule as the parent fetch). Dangling refs to draft/unpublished targets becomenull.entry.edit): I propose v1 does not wrap hydrated children with edit proxies, since they originate from a different collection. That can be added later if there's demand.Type generation
emdash-env.d.tscurrently generatesreferencefields asstring(zod-generator.ts:366). For v1 I'd leave that as-is (so existing code keeps compiling) and let consumers cast where they need the hydrated shape:Improving generated types so
depth: 1returnsLiquor & { producer: Producer }is doable but TypeScript-heavy and probably worth its own follow-up PR.Alternatives considered
Comparison with peers
populate=*, deeply nested object formdepth: number(docs)*[_type=="post"]{..., author->}?fields=*,author.*Payload's
depthis closest in spirit to what I'm proposing here — minimal API surface, sane default, easy to reason about.Questions for maintainers
depthAPI the right shape, or do you prefer to start withpopulatedirectly? My case fordepthfirst: it's small, opt-in, and unblocks the common case immediately.populatecan be layered on without breaking it.depthmatches Payload;populate: true | objectmatches Strapi. I'm flexible.getRequestContext()'s edit-mode wrapping, or stay raw?What I'm offering
If this gets a thumbs up, I'd open a PR with:
loader.ts/query.tschangeshydrate-references.tsmodule mirroring the byline patternpackages/core/test/covering: zero-depth no-op, single ref, multi ref, multi collection, dangling ref, depth=2 recursion, IN-clause chunking, locale propagation.changeset/minor-version entrydocs/src/content/docs/guides/querying-content.mdxHappy to iterate on the API in this thread before writing any code.
(AI disclosure: this proposal was drafted with Claude Code. Implementation, tests, and review will be human-checked against the existing patterns in
bylines/index.tsand follow the same style.)All reactions