-
Notifications
You must be signed in to change notification settings - Fork 0
Easy Migrate Without Code Change
ToplingDB is a fork of RocksDB. To reduce development costs, avoid fragmenting the ecosystem, and continue benefiting from the RocksDB community, ToplingDB keeps changes to the RocksDB codebase to a minimum. It does so through a Side Plugin architecture: Factory/Product classes dynamically inject third-party implementations, while new code remains outside the original RocksDB codebase. See Motivation To Solution for the motivation and architecture behind this design.
A typical configuration file, such as mytopling.json, defines dozens of modules:
- Plugins and extensions: integrate custom comparators, table formats, MemTables, merge operators, compaction filters, and more through configuration or application code. See AnyPlugin and CompactionFilterFactory As SidePlugin.
- Advanced storage engines: CSPP MemTable, a Crash-Safe Parallel Patricia Trie with more than 30 times the concurrent-write performance of SkipList; ToplingZipTable, which combines CO-Index and PA-Zip for high in-memory compression; SingleFastTable, which provides maximum performance without compression; and DispatcherTable, which lets each LSM level use a different TableFactory.
- Distributed compaction: offload compaction to a remote worker cluster. See Distributed Compaction.
- Embedded WebView: run a built-in HTTP server that provides an observability dashboard. See WebView.
- Prometheus monitoring: expose metrics through the Statistics module and display them in Grafana. See Displaying ToplingDB Metrics with Grafana.
- Online configuration updates: adjust mutable database- and ColumnFamily-level options through the web interface. Features such as Dyna MemTable can switch between SkipList, CSPP, and other implementations without restarting the process.
-
Configuration templates: reuse configuration fragments through
${var}references and${template}inheritance. See Template Object.
This is the most complete way to use ToplingDB. It requires a well-structured configuration file and an understanding of Topling-specific components and their parameters. For a new project, we recommend integrating SidePluginRepo directly into the application. See Using ToplingDB from Scratch for a complete introduction. An existing RocksDB application can instead use Easy Migrate: no application code changes are required; only the build dependency needs to be changed to ToplingDB.
To lower migration costs, ToplingDB provides the TOPLINGDB_EASY_MIGRATE_CONF environment variable.
Point this variable to a JSON or YAML configuration file. With zero application code changes and only the build dependency changed to ToplingDB, the application can use every ToplingDB performance enhancement enabled in that configuration, including CSPP MemTable, ToplingZipTable, and DispatcherTable. The precise set of enabled components still depends on the YAML or JSON file.
export TOPLINGDB_EASY_MIGRATE_CONF=/path/to/easy_conf.jsonThe Easy Migrate configuration format is identical to that used by SidePluginRepo/full integration. Setting the environment variable is enough to load it transparently into the application.
When http.auto_start_http is enabled, loading the Easy Migrate configuration starts the same embedded web service used by WebView. Use the database normally—read-only and secondary instances count as normal use—and each active instance automatically appears in the web dashboard under its database name. No additional integration code is needed for the operations dashboard. See Optional Embedded Web Service below for the configuration.
- New projects: prefer full SidePluginRepo integration. It loads configuration in-process, manages database lifecycles, and naturally supports multi-database orchestration and language bindings. Start with Using ToplingDB from Scratch and see Home for an overview.
-
Existing RocksDB applications: use Easy Migrate. Set
TOPLINGDB_EASY_MIGRATE_CONF, change the build dependency to ToplingDB, and leave the application code untouched. The same JSON/YAML file can enable the web dashboard, monitoring, online configuration changes, and the path-based configuration selection described in this article.
Both approaches use the same configuration format. At startup, custom comparators, table formats, MemTables, merge strategies, compaction filters, and other extensions can be integrated as described in AnyPlugin and the relevant guides. These extensions are fully compatible with Easy Migrate. Once configured, distributed compaction can also be enabled, provided the remote compaction cluster understands the same configuration semantics.
If an existing application needs to register already-running instances with SidePluginRepo for orchestration without changing its main business path, see Migrating Complex DB-Opening Code. This is a compatibility technique for legacy integration, not a replacement for the recommendation to use SidePluginRepo directly in new projects.
Easy Migrate uses exactly the same configuration format as SidePluginRepo/full integration. If SidePluginRepo is later brought under in-process application control, the same file can be reused as-is. Custom compaction filters, distributed compaction, and other extensions can be configured as described in their respective guides. They are not mutually exclusive with Easy Migrate, as long as the workers share the same configuration semantics.
The DBOptions and CFOptions sections can reference other components defined in the file. DBOptions.write_buffer_manager references an instance in the top-level WriteBufferManager module. CFOptions.memtable_factory and CFOptions.table_factory can respectively reference Topling CSPP MemTable, under MemTableRepFactory, and DispatcherTable, which is first defined as a named entry under TableFactory and then referenced from the ColumnFamily options. sample-conf/db_bench_enterprise.yaml is the authoritative, complete example of the matching WriteBufferManager, cspp, and dispatch definitions and references. Other db_bench_*.yaml files in the same directory are also useful examples. Both JSON and YAML are supported; see Yaml File.
As with full integration, the configuration file can contain a top-level http section. See Home and WebView for settings such as document_root and listening_ports. The application-facing behavior is as follows:
-
http.auto_start_http: the embedded web service starts automatically after the Easy Migrate configuration is loaded only when this value is true. If it is false or omitted, no port is opened, preventing accidental exposure of an operations endpoint. Its semantics are the same as when the complete Topling service is started from a configuration file. - Instances visible in the dashboard: use the database normally, including in read-only or secondary mode, and each instance appears in the web dashboard automatically. No application change is required merely to make the dashboard recognize it. Full SidePluginRepo integration can also register instances explicitly with the repository, which is the recommended path for new projects. See Migrating Complex DB-Opening Code for the legacy-integration technique.
Example excerpt in YAML:
http:
document_root: /dev/shm/mydb_web
listening_ports: '2011'
auto_start_http: true # Start the embedded web service only when this is trueThe YAML below is an excerpt. See db_bench_enterprise.yaml for complete definitions and parameters for WriteBufferManager, MemTableRepFactory (CSPP), and TableFactory (DispatcherTable).
# See db_bench_enterprise.yaml at the link above for complete examples of
# MemTableRepFactory, TableFactory, and WriteBufferManager.
MemTableRepFactory:
cspp:
class: cspp
params:
mem_cap: 1G
TableFactory:
fast:
class: SingleFastTable
params:
indexType: MainPatricia
keyPrefixLen: 0
zip:
class: ToplingZipTable
params:
localTempDir: /dev/shm
sampleRatio: 0.01
dispatch:
class: DispatcherTable
params:
default: fast
readers:
SingleFastTable: fast
ToplingZipTable: zip
level_writers: [fast, fast, zip, zip, zip, zip, zip, zip]
WriteBufferManager:
wbm:
class: WriteBufferManager
params:
buffer_size: 2560M
DBOptions:
default:
write_buffer_manager: "${wbm}" # Top-level WriteBufferManager instance "wbm"
max_background_compactions: 32
max_subcompactions: 1
max_background_flushes: 2
max_total_wal_size: 5G
allow_mmap_reads: false
manual_wal_flush: true
two_write_queues: true
CFOptions:
default:
memtable_factory: "${cspp}" # Topling CSPP MemTable (MemTableRepFactory)
table_factory: dispatch # References TableFactory.dispatch
num_levels: 7
write_buffer_size: 2560M
max_write_buffer_number: 4
target_file_size_base: 32M
level_compaction_dynamic_level_bytes: true
disable_auto_compactions: trueIf every database shares one set of options, define one configuration block named default under both DBOptions and CFOptions. This is also the final fallback name when path matching finds nothing. See the minimal Sui example or the repository's sample-conf directory.
This raises a fundamental question:
When several database instances with different paths run in the same process or on the same machine, how can each select the right configuration?
The answer is namespace resolution.
Consider three typical instances with different paths. They may run in the same process or be distributed across the same set of nodes:
/data/order/prod/orders/db
/data/order/prod/items/db
/data/analytics/prod/metrics/db
Different workloads often need different combinations of DBOptions, which apply at the database level, and CFOptions, which apply at the ColumnFamily level. You do not need to memorize the ownership of every setting; the important point is that both groups may need tuning.
-
orders/db, as a path suffix: write-heavy. Typical changes include larger MemTables or write buffers, usually in CFOptions, and greater background-compaction parallelism, usually in DBOptions. -
items/db: read-heavy and write-light. It will often favor a faster SST/table format and block-related settings, usually in CFOptions. -
metrics/db: bulk ingestion. It will often favor throughput and a different compaction strategy, potentially involving both CFOptions and DBOptions.
With a traditional integration, application code often assembles different Options through path-dependent branches. Easy Migrate instead records these choices in one global JSON/YAML file. Each instance opens normally, but selects the appropriate section from that shared file according to its own database path, followed by the ColumnFamily-name rules described in Section 5.
The system therefore needs a stable way to distinguish instances using only the path string and the CF name. It must also let a set of paths with the same suffix but different prefixes share one configuration by default.
Typical intent: when paths partition a database or separate business lines in one process, databases of the same kind often share a suffix such as orders/db or items/db, while prefixes such as the mount root, tenant, or cluster directory vary by partition. In that case, maintain one short suffix key such as orders/db and share its DBOptions/CFOptions across all matching instances. When one prefix needs an exception, define a longer qualified path that includes the distinguishing prefix. The rule “try the longer name first and stop at the first match” lets it override the shared setting.
Grouping paths by a shared suffix closely resembles how a programming language resolves a symbol by trying a more fully qualified namespace before a more general name. This guide therefore treats a path as a namespace.
Easy Migrate assumes zero application code changes, with only the build dependency changed to ToplingDB. The configuration cannot see semantic labels such as “order database” or “audit database.” It can use only the database path known to the engine and the ColumnFamily name. The built-in Default CF always has the fixed name "default", as explained in Sections 5 and 6. A path itself often encodes data-center, tenant, business-line, and environment hierarchies. The central question is therefore: how can that one path, together with the CF name, express configuration intent at every required level?
Think of a database path as a qualified name ordered from outside to inside. The left side generally contains variable partition or business prefixes, while the right-hand suffix identifies a stable kind of database. A longer qualified name, containing more left-hand segments, targets a specific prefix; a shorter suffix key is shared across prefixes. Configuration lookup begins with the entire path, removes one segment at a time from the left, and stops as soon as a matching entry is found. In other words, a prefix-specific name is tried before the shared suffix name.
For example, /data/order/prod/orders/db produces the following candidates, in the same order shown in the next section:
/data/order/prod/orders/db, data/order/prod/orders/db, order/prod/orders/db, prod/orders/db, orders/db, and db, followed when necessary by the configuration block named default.
Each path segment maps to one namespace level. In /data/order/prod/orders/db, data is the outermost level, order and prod are nested intermediate levels, and orders/db is the innermost entity name.
Lookup starts with the longest qualified name. Following the rule in Section 3.3, it removes the leftmost segment to produce progressively shorter suffix keys, trying each in turn until one matches.
Given a path, remove one component at a time from the left, including the root slash /.
For /data/order/prod/orders/db:
/data/order/prod/orders/db Step 1: full path, most specific
data/order/prod/orders/db Step 2: remove the root slash /
order/prod/orders/db Step 3: remove data
prod/orders/db Step 4: remove order
orders/db Step 5: remove prod
db Step 6: basename only
"default" Step 7: final fallback if nothing above matches
Why shorten the path from the left instead of walking up the directory tree from the end?
The left side represents the outer deployment topology—data center, business line, and environment—while the right side identifies the database itself. Removing the rightmost segment first would discard the part that best distinguishes one database from another, leaving prefixes that cannot express common requirements such as different databases under the same business line.
A path with N levels provides N namespace levels, and configuration can be defined at any of them. Lookup always starts at the most specific level and stops at the first match.
Suppose the following configuration is defined:
{
"CFOptions": {
"orders/db": {
"write_buffer_size": "2560M",
"max_write_buffer_number": 4
},
"prod/orders/db": {
"target_file_size_base": "64M"
}
}
}Tracing the CFOptions lookup for /data/order/prod/orders/db—the sample JSON contains only CFOptions—gives:
Step 1: /data/order/prod/orders/db no match
Step 2: data/order/prod/orders/db no match
Step 3: order/prod/orders/db no match (the key is not defined)
Step 4: prod/orders/db match: target_file_size_base=64M
Lookup stops there. The write_buffer_size from orders/db is not used, because prod/orders/db has already matched.
Easy Migrate begins with the key corresponding to the entire path. At each step it removes the leftmost path component, thereby trying progressively shorter suffix-shaped keys. The first defined key ends the lookup. Shorter suffix keys are not tried, and fields from multiple path keys are not merged. This realizes the intent described in Section 3.1: use a short suffix key as the cross-prefix default, and add a longer qualified name only when a partition or business prefix needs an exception.
This rule applies only to keys on the same shortening chain. A key that never occurs in an instance's candidate sequence does not get “shadowed”; it simply never participates in lookup. For example, if the instance path ends in .../orders, its chain can never contain orders/db.
The example below deliberately uses paths in the same .../orders/db family, placing the shared suffix key orders/db and the prefix-specific override sensitive/audit/orders/db on the same lookup chain.
The section shared across partitions is stored under the suffix-only key orders/db:
{
"CFOptions": {
"orders/db": {
"write_buffer_size": "256M",
"max_write_buffer_number": 4
}
}
}The audit database lives at /data/sensitive/audit/orders/db. To override a few fields for this prefix while remaining on the same lookup chain, define the longer key sensitive/audit/orders/db, including the distinguishing partition segments. Once it matches, lookup stops before reaching the shared orders/db suffix:
{
"CFOptions": {
"orders/db": {
"write_buffer_size": "256M",
"max_write_buffer_number": 4
},
"sensitive/audit/orders/db": {
"disable_auto_compactions": true
}
}
}Lookup for /data/sensitive/audit/orders/db proceeds until:
/data/sensitive/audit/orders/db no match
data/sensitive/audit/orders/db no match
sensitive/audit/orders/db match: disable_auto_compactions=true
Lookup stops here. The shorter suffix keys audit/orders/db, orders/db, db, and so on are not tried, so write_buffer_size from the shared orders/db entry is not merged into the result. To reuse most fields from orders/db while overriding a few for one prefix, explicitly inherit the shared fragment with template / ${template} as described in Template Object. Do not rely on multiple path matches being combined automatically.
For any database path, repeatedly remove its leftmost component, as described in Section 3.3, and look for a configuration key equal to each resulting string. This covers every available specificity level. For /data/order/prod/orders/db, the root slash and five path components produce:
Step Configuration key
─────────────────────────────────
1 /data/order/prod/orders/db full path, most specific
2 data/order/prod/orders/db root slash removed
3 order/prod/orders/db data removed
4 prod/orders/db order removed
5 orders/db prod removed
6 db basename only, least specific
7 "default" final fallback
To apply different settings to a ColumnFamily, append :<ColumnFamily name> to each path candidate and use the rules in the next section.
ColumnFamily configuration adds one dimension to DBOptions lookup: the ColumnFamily name. Every RocksDB database has a Default Column Family, whose name is fixed by the engine as "default" rather than chosen by the application like a custom ColumnFamily name.
For a ColumnFamily named mycf, lookup has two phases. Within each phase, it follows the same long-to-short path-suffix chain as Section 3.3, appends :mycf to each key, and finally tries the bare name mycf without a path.
Phase 1: configuration for the named CF
Append :mycf to each candidate on the Section 3.3 path chain, then try bare mycf:
/data/order/prod/orders/db:mycf
data/order/prod/orders/db:mycf
order/prod/orders/db:mycf
prod/orders/db:mycf
orders/db:mycf
db:mycf
mycf bare-name fallback
Phase 2: fall back to the built-in Default CF configuration
If Phase 1 finds no match and the current ColumnFamily mycf is not the built-in "default" CF, repeat the same search with the suffix default. Here, "default" is the fixed name of the Default Column Family, allowing a custom CF to inherit CFOptions written for the default CF:
/data/order/prod/orders/db:default
data/order/prod/orders/db:default
order/prod/orders/db:default
prod/orders/db:default
orders/db:default
db:default
default bare-name fallback
When the target is already the built-in Default CF, whose name can only be "default", Phase 1 has covered the entire Section 3.3 path-suffix chain. Phase 2 must not repeat the same fallback-to-default search.
This design provides fine-grained control. A specific ColumnFamily under a specific path can have custom settings, while application-defined ColumnFamilies with no matching entry can reuse the configuration written for the built-in Default CF under the same path.
For dynamic creation and deletion of ColumnFamilies under full integration, see Dynamic Create Drop ColumnFamily.
The string "default" has three meanings in a configuration file. Do not confuse them when writing keys:
-
A configuration-block name under
DBOptionsorCFOptions: the block used when no more specific match is found, commonly serving as a global default. -
A literal directory component in a path: for example,
defaultin/data/prod/default. -
The name of RocksDB's built-in Default Column Family: this is fixed by the engine as
"default", not chosen by the application as a custom ColumnFamily name.
In practice: if the path basename is already default, do not add a redundant fallback with the same name. When writing a rule for the built-in Default CF, run the first-phase path × :default chain only; do not “fall back to the default CF” a second time.
During initial evaluation, or when all instances have similar workloads, define only default:
{
"DBOptions": {
"default": {
"max_background_compactions": 8
}
},
"CFOptions": {
"default": {
"write_buffer_size": "2560M",
"max_write_buffer_number": 4
}
}
}Every database instance matches default, and every ColumnFamily matches default. This is the lowest-cost migration: write one configuration file, set one environment variable, and restart.
To tune different kinds of databases, use multi-component path suffixes as keys. The orders/db and items/db entries below follow the shared-suffix pattern from Section 3.1:
{
"CFOptions": {
"default": {
"write_buffer_size": "256M",
"max_write_buffer_number": 4
},
"orders/db": {
"write_buffer_size": "5120M",
"max_write_buffer_number": 6
},
"items/db": {
"target_file_size_base": "64M"
}
}
}-
/data/any/prefix/orders/dbmatchesorders/dband gets a larger MemTable/write buffer, primarily throughwrite_buffer_sizeandmax_write_buffer_numberin this example. -
/data/any/prefix/items/dbmatchesitems/dband gets larger SST files. - Every other instance matches
defaultand uses the standard 256 MB MemTable.
In a more complex deployment, use longer qualified paths, whose keys contain more path components, to distinguish business lines or intermediate directory levels. This follows the Section 3.1 pattern of adding a longer qualifier only when necessary. Suppose paths are organized by business line:
/data/order/db/orders/db
/data/order/db/items/db
/data/analytics/db/metrics/db
Keys containing the business-line component—order/db/... and analytics/db/... below—allow each instance to match its corresponding entry:
{
"CFOptions": {
"order/db/orders/db": {
"max_write_buffer_number": 8,
"level_compaction_dynamic_level_bytes": false
},
"order/db/items/db": {
"max_write_buffer_number": 2
},
"analytics/db/metrics/db": {
"disable_auto_compactions": true,
"level_compaction_dynamic_level_bytes": false
}
}
}- The lookup chain for
/data/order/db/orders/dbmatchesorder/db/orders/db. - The lookup chain for
/data/order/db/items/dbmatchesorder/db/items/db. - The lookup chain for
/data/analytics/db/metrics/dbmatchesanalytics/db/metrics/db.
Different business lines can therefore use different settings while configuration remains centralized. Updating one JSON/YAML file affects every corresponding instance.
Suppose an audit instance must disable background compaction and use only one exact key on its path chain. Once that key matches, shorter suffixes are not tried:
{
"CFOptions": {
"compliance/audit/db": {
"disable_auto_compactions": true,
"level_compaction_dynamic_level_bytes": true
}
}
}The compliance/audit/db entry matches /data/compliance/audit/db. Lookup stops at that match, so shorter suffix entries on the same chain are not read. A shared key with a different path shape would never have been in this instance's candidate sequence and is therefore unrelated to this match.
Sometimes a database or ColumnFamily should remain completely unaffected by Easy Migrate because the instance already has the desired options and should not be overridden by the global file.
Use the “stop at the first match” rule and define an empty object {} at the appropriate level:
{
"DBOptions": {
"special/db": {},
"default": {
"max_background_compactions": 8
}
},
"CFOptions": {
"special/db:mycf": {},
"default": {
"write_buffer_size": "256M"
}
}
}For an instance at /data/special/db, lookup matches the empty special/db object and does not merge anything below it. Database-level options retain their preexisting behavior. The same applies to the mycf ColumnFamily: once special/db:mycf matches the empty object, lookup does not fall back to the CF configuration block named default.
- Whether the dashboard, monitoring, and online configuration changes are available depends on whether the YAML/JSON contains
http, Statistics, and related settings, and whetherauto_start_httpis enabled. It does not depend on whether Easy Migrate is used. - Once the dashboard is enabled, database instances in use by the process appear there automatically. Application changes are generally unnecessary merely for display.
- Follow the relevant guides for custom plugins and distributed compaction. New projects should use SidePluginRepo directly; see Using ToplingDB from Scratch. For an existing application that needs to register instances with the repository manually, see Migrating Complex DB-Opening Code.
- New projects: use full SidePluginRepo integration as the primary path. See Using ToplingDB from Scratch and Compile And Install.
- Existing applications: begin with Easy Migrate for evaluation. No application code changes are required; only the build dependency needs to be changed to ToplingDB. If SidePluginRepo is later managed in-process, the same configuration file can be reused instead of starting over.
Phase 1: Validate Quickly with Easy Migrate
export TOPLINGDB_EASY_MIGRATE_CONF=./easy_conf.jsonMeasure ToplingDB on the existing workload. If the configuration enables components such as CSPP MemTable and ToplingZipTable, their performance benefits are available in this phase.
Phase 2: Add Observability
While continuing to use Easy Migrate, add http, auto_start_http: true, Statistics, and related sections to the same file. This enables WebView and Prometheus/Grafana with zero application code changes. For a long-term integration that manages SidePluginRepo or language bindings in-process, follow the recommended new-project path in Using ToplingDB from Scratch, rather than treating a convenience path for legacy applications as the permanent architecture.
Phase 3: Enable Online Configuration Updates
Use the embedded web service's dynamic configuration support to tune parameters without restarting. The namespace mechanism in this article naturally separates different business lines.
Phase 4: Add Distributed Compaction
When compaction becomes a bottleneck, configure Dcompact to offload it to a worker cluster. Custom compaction filters and similar extensions can be integrated as described in CompactionFilterFactory As SidePlugin and used together with Easy Migrate. The worker cluster must understand the same configuration semantics when performing remote compaction.
Each phase can stand on its own; there is no need to adopt every phase at once.
Start with the database's full path. Remove one path component at a time from the left and look for a configuration key with the resulting name. Stop at the first match; do not merge it with configuration blocks at shorter suffixes on the same chain. Use a block named default as the global fallback. For a ColumnFamily, append :<ColumnFamily name> to each candidate on the same path-suffix chain. An application-defined ColumnFamily may also fall back along that chain to :default, the fixed name of the built-in Default CF. For the architectural background, see Motivation To Solution.