-
Notifications
You must be signed in to change notification settings - Fork 1
Grouped Files
What this page covers: opening the grouped-file backend with Storages.createGroupedFile, the
key-major on-disk layout (one file per key, holding every collection that shares that key —
the inverse of Local Files' one-file-per-entity layout), why it's ideal for "everything about one
entity-root in one file" workloads, the fact that — like local files — it accepts a JSON or YAML
codec, how queries work as a correct-but-slow full scan (no real index, yet still validating
index declarations), crash-safe atomic writes, and the two things it can't do — no
transactions, no index acceleration.
📌 Note — switching to grouped files is a one-line change at construction. From
storage.repository(...)onward your code is identical to every other backend. See Choosing a Backend for the full capability matrix.
import br.com.finalcraft.everydatabase.*;
import br.com.finalcraft.everydatabase.codec.JacksonYamlCodec;
import br.com.finalcraft.everydatabase.modules.groupedfile.GroupedFileConfig;
import java.nio.file.Paths;
import java.util.Optional;
import java.util.UUID;
// 1. Describe the entity once — key type FIRST, entity type second.
EntityDescriptor<UUID, PlayerData> PLAYERS = EntityDescriptor.builder(UUID.class, PlayerData.class)
.collection("PlayerData") // becomes a sub-node in each key file
.keyExtractor(PlayerData::getUuid)
.codec(new JacksonYamlCodec<>(PlayerData.class)) // YAML or JSON — both allowed here
.build();
// 2. Open the backend — just point it at a directory.
GroupedFileStorage storage = Storages.createGroupedFile(
new GroupedFileConfig(Paths.get("playerdata")));
storage.init().join(); // creates the base directory if absent
// 3. Use it — exactly like any other backend.
Repository<UUID, PlayerData> repo = storage.repository(PLAYERS);
UUID id = UUID.randomUUID();
repo.save(new PlayerData(id, "Alice", 100)).join(); // writes playerdata/<uuid>.yml
Optional<PlayerData> alice = repo.find(id).join(); // -> Optional[Alice]
storage.close().join();createGroupedFile returns the concrete GroupedFileStorage type. The PlayerData entity is the
same plain Jackson POJO used throughout the wiki (Quick Start).
📌 Note — every I/O call returns a
CompletableFuture..join()is shown for brevity; compose withthenApply/thenComposein real code. There are no blocking variants — see The Async API.
A single constructor — just the base directory where the per-key files live:
import java.nio.file.Paths;
new GroupedFileConfig(Paths.get("playerdata"));| Argument | Type | Meaning |
|---|---|---|
baseDirectory |
Path |
root directory; each key becomes one file directly under it |
There are no formatting knobs: the container format follows the codec (below), and JSON output is always indented for readability.
Where Local Files is collection-major (<base>/<collection>/<key>.json — one file per entity),
grouped files invert the layout to key-major: one file per key, directly under the base
directory, holding every collection that shares that key as a sub-node:
playerdata/
_schema/
layout.json # reserved: the container format of this directory
migrations.json # reserved: applied migration versions (isolated in a sub-dir)
5f1e8400-e29b-41d4-a716-446655440000.yml
9a2c….yml
Each key file is a single structured document, one top-level node per collection:
# playerdata/5f1e8400-….yml
PlayerData: # collection 1
uuid: "5f1e8400-e29b-41d4-a716-446655440000"
name: "Alice"
score: 100
AuthMe: # collection 2, same UUID key
passwordHash: "…"Collections co-locate when their keys stringify to the same value.
🧭 When this shines — a player whose data is spread across many logical collections, loaded on join and persisted on quit as a single read/write per key. One file is the whole entity-root.
The filename is the sanitized key plus the codec's extension (.yml / .json). FileKeyNames.safeStem
handles path separators, case-differing names on case-insensitive filesystems, reserved Windows device
names, and over-long keys: it appends a stable hash suffix (hash-truncating an over-long stem to a
short prefix plus that hash), so distinct keys stay one file per key, portable, with the file and
its per-key lock always colliding-or-not together. The key contract is otherwise the usual one —
a stable, unique toString() of ≤ 255 characters — see
Entities, Keys & Collections. Collection names must match
^[a-zA-Z][a-zA-Z0-9_]*$.
📌 Note — writes are crash-atomic. Each save is a read-modify-write of the whole key file, guarded by a per-key write lock (shared storage-wide so two collections of the same key never lose each other's update), then published via a sibling
.tmp+ATOMIC_MOVE(falling back when the filesystem can't do an atomic move). A crash mid-write never leaves a truncated file — at worst an orphan.tmp. This guards against torn files, not power loss: the write is not fsynced, so the very last save can still be lost on a power cut if the OS hasn't flushed its page cache. Deleting a key's last remaining collection removes the now-empty file.
Like Local Files, grouped files treat the payload structurally, so this backend accepts a
JSON or YAML codec. There is no format option — the container format is taken from the codec on
the descriptor: JacksonJsonCodec yields .json files, JacksonYamlCodec yields readable .yml
files.
import br.com.finalcraft.everydatabase.codec.JacksonJsonCodec;
import br.com.finalcraft.everydatabase.codec.JacksonYamlCodec;
.codec(new JacksonYamlCodec<>(PlayerData.class)) // -> playerdata/<uuid>.yml
.codec(new JacksonJsonCodec<>(PlayerData.class)) // -> playerdata/<uuid>.json (indented)
⚠️ Gotcha — all collections sharing one base directory must agree on one format. They write the same physical files, so mixing a JSON codec and a YAML codec under the sameGroupedFileConfigfails fast withIllegalStateException. A codec that is neither JSON nor YAML (an opaque/binary codec) can't be embedded into the structured aggregate document and is rejected withIllegalArgumentException. Pick one format per base directory.
Since 1.2.0 the directory remembers. The format is written to _schema/layout.json on first use,
so reopening a YAML directory with a JSON codec fails instead of quietly reporting an empty
collection and starting a parallel set of .json files beside the .yml ones holding the data. A
directory that predates the file has its format inferred from the extensions on disk and written
down; one that already holds both formats refuses to open, listing how many files of each it found —
that is the fingerprint of a mismatch that already happened, and only you can say which set to keep.
Grouped files and Local Files are the only two backends that accept YAML — SQL, Mongo, and
in-memory require a JSON codec (codec.isJsonCodec()). See Codecs.
Like local files, grouped files have no real index. Queries are answered by a full scan — every
key file is read, this collection's sub-node is extracted, and the entity is matched in memory. It's
O(total keys) and slow on large datasets (the scan reads files of unrelated collections too), but
it returns the correct result, so a query that works here keeps working when you swap to SQL or
Mongo.
repo.findBy("score", 100).join();
repo.query(Query.range("score", 50, null)).join(); // score >= 50 (full scan, correct result)Crucially, the backend still validates index declarations before scanning:
⚠️ Gotcha — querying a field that was not declared as anIndexHintthrowsIllegalArgumentExceptionhere too — even though the scan could answer it. This is deliberate: it stops a query that works on grouped files from silently breaking when the storage is swapped for SQL/Mongo (which genuinely require the declaration). Declare the field with.index(IndexHint.<type>("..."))or@Indexed. See Indexing & Queries.
📌 Note — a scan reads each key file individually, so it is point-in-time consistent per file, not a whole-store snapshot. A key created or deleted while the scan runs may be transiently missing (or double-counted) — the scan sees each file as of the moment it reads it, never a frozen view of the whole directory. For an exact scan (e.g. a
count()you'll act on), run it against a quiescent store.
This is the one backend where a stored file can belong to no collection. A key file aggregates every collection sharing its key, so if the document itself will not parse, nothing inside it can be attributed: skipping it under-reports the collection being read, and counting it inflates every other collection in the directory.
Because neither number is true, the reads that answer with a number or a set refuse:
GroupedFile: cannot count 'ec_accounts': key file 'a3f1c2.json' is not a readable document,
so it belongs to no collection and any answer would be a guess. Use scanAll() to list every
unreadable file, then repair or remove it.
| Read | On an unparseable key file |
|---|---|
count() |
throws IllegalStateException, naming the file |
keys(...) |
throws IllegalStateException, naming the file |
all() / query(...)
|
skips it, logs a WARN — they hand back entities, so omitting one they cannot read is their contract everywhere |
scanAll(...) |
reports it as a failed ScanRow — this is the diagnosis, run it when a count fails |
find(key) / exists(key)
|
throws for that key only; other keys are unaffected |
⚠️ Gotcha — do not confuse this with a poisoned row: a key file that parses fine and declares the collection, whose payload does not decode. That row is unambiguously yours, socount()counts it,all()skips it, andcount() != all().count()is the intended tell-tale. Only a file broken as a document triggers the refusal above.
💡 Tip — a boot-time
count()on a grouped-file store can now fail. If your startup treats the count as infallible (a guard, a health check), decide explicitly what an unreadable store means to it — usually "refuse to start" is the right answer, since the alternative is running on data you cannot fully see.
GroupedFileStorage does not implement tx.TransactionalStorage. There is no inTransaction.
⚠️ Gotcha — a singlesaveis atomic (the.tmp+ATOMIC_MOVEwrite of the whole key file), and because one key file holds every collection of that key, writing several collections for the same key in onesavecycle lands atomically together. But there is no multi-key transaction — you can't commit/rollback a group of different-key writes as a unit. If you need ACID across multiple entities, use a SQL backend (MySQL & MariaDB / PostgreSQL / H2) or Mongo with a replica set. See Transactions.
Optimistic locking isn't enforced either — a versioned descriptor's versions(...) reports 0 for
existing keys (like H2), so it degrades to plain upsert. See Optimistic Locking and
Choosing a Backend.
Unlike transactions, migrations work. GroupedFileStorage implements schema.SchemaAwareStorage;
applied versions are tracked in a reserved _schema/migrations.json ledger under the base directory,
and migrations are forward-only. The ledger lives in its own _schema/ sub-directory (never
collides with a key file) and is written atomically (.tmp + move) — a truncated ledger would read
back as "nothing applied" and re-run every migration over already-migrated data. Extend
GroupedFileMigration and override executeOnStorage(GroupedFileStorage):
import br.com.finalcraft.everydatabase.modules.groupedfile.GroupedFileMigration;
class V1_SeedAdmins extends GroupedFileMigration {
@Override public String version() { return "001"; }
@Override public String description() { return "seed admin profiles"; }
@Override protected void executeOnStorage(GroupedFileStorage storage) {
// mutate via repositories obtained from `storage`
}
}
storage.register(new V1_SeedAdmins()).migrate().join();MigrationContext.getNativeClient(...) also exposes the GroupedFileStorage and its base Path. See
Schema Migrations.
🧭 Decision — choose grouped files when your data is naturally key-major: many logical collections that all key on the same identity (e.g. a player UUID), and you want the whole entity-root in one human-readable file loaded and saved as a unit. One file per key (YAML or indented JSON), crash-atomic, zero ops. The trade-offs are the same as local files: no transactions, no real index (full-scan queries), no optimistic-locking enforcement — fine for small or cold datasets and single-server deployments. If you'd rather have one file per entity per collection, use Local Files; for large or write-heavy data, or anything needing ACID, prefer a SQL backend or Mongo.
A common pattern: ship small deployments on grouped/local files, let operators flip to MariaDB/Mongo for large ones and move the live data with Moving Data Between Backends — no code changes, source untouched.
| Capability | Grouped Files |
|---|---|
| Transactions | ❌ |
| Schema Migrations | ✅ (tracked in _schema/migrations.json) |
| Indexing & Queries | |
| Optimistic Locking | ❌ (degrades to upsert) |
| Change feed | ✅ (OS watch service — sees edits made outside the app) |
| Key-major reads/writes | ✅ KeyMajorStorage
|
| Persistence | Durable (one file per key) |
Every key file lives directly under the base directory, and every scan reads all of them. So a
base directory that accumulates collections keyed by unrelated things — player UUIDs, account ids,
free-form cooldown names — is paying twice: the collections never share a meaningful key, an
accidental collision would put two of them in one file behind one lock, and each pays for the other's
files on every count(), all(), query() and keys().
Keeping them apart is a second storage over a second directory. There is no config knob for it, and none is needed — a directory already is the unit of separation:
Storage players = Storages.createGroupedFile(new GroupedFileConfig(base.resolve("player")));
Storage accounts = Storages.createGroupedFile(new GroupedFileConfig(base.resolve("account")));Each gets its own directory, listing, lock map, layout file and migration ledger.
⚠️ Key files in a sub-directory make the open fail. Only the base directory is read, so those files would be invisible: the format would be inferred as if the directory were empty, the collections would report nothing, and the first save would write a second copy beside data nobody can see. The error names the sub-directory — open a storage over it, or move the files up into the base first.
Every collection of one key already lives in one file, behind one lock, so grouped files implement
the KeyMajorStorage capability. Check for it with instanceof and fall back on backends that store
collections apart:
if (storage instanceof KeyMajorStorage kms) {
KeyBundle bundle = kms.loadKey(uuid, PLAYER_DATA, ECONOMY, HOMES).join(); // one parse
PlayerData data = bundle.get(PLAYER_DATA).orElseGet(PlayerData::new);
…
kms.batchKey(uuid, b -> b // one atomic move
.put(PLAYER_DATA, data).put(ECONOMY, eco).put(HOMES, homes)).join();
} else {
PlayerData data = storage.repository(PLAYER_DATA).find(uuid).join().orElseGet(PlayerData::new);
…
}Beyond the saved I/O, the batch is the only version that is atomic: N separate saves can be interrupted between two of them and leave the key half-updated.
⚠️ Atomicity is per key and nothing more. Grouped files still do not implement Transactions — there is no rollback, no isolation beyond the per-key lock, and no way to span two keys.
- Choosing a Backend — the capability matrix and data-at-rest formats across all backends.
- Local Files — the collection-major sibling (one file per entity).
- Quick Start — the minimal describe → open → use → close round-trip.
-
The Async API — the
CompletableFuturemodel every call shares. -
Codecs —
JacksonYamlCodec(file backends only) vs the JSON-required backends. - Indexing & Queries — why declarations are validated even without a real index.
-
Schema Migrations —
GroupedFileMigration,_schema/migrations.json, forward-only. - Entities, Keys & Collections — the key/collection contract and filename sanitisation.
- Transactions — why grouped files opt out, and the backends that don't.
- Cross-Process Cache Sync — no push feed in v1; version polling (deletes only).
- Moving Data Between Backends — migrate a grouped-file store into SQL/Mongo (codec change).
EveryDatabase · Home · made by Petrus Pradella
Getting Started
Core Concepts
Working with Data
Backends
- Choosing a Backend
- MySQL & MariaDB
- PostgreSQL
- H2
- MongoDB
- Local Files
- Grouped Files
- In-Memory
- Benchmarks
Manager Module
- Caching & References
- Typed References (Ref)
- Caching Managers
- Cache Policies & Freshness
- Cross-Process Cache Sync
- Write-Back & Conflict Resolution
- Payload Schema Evolution
- One Entity, Many Databases
Operations
Advanced
Reference
Contributing