Skip to content

Embeddium

BonsUnleashed edited this page Oct 7, 2026 · 4 revisions

Embeddium

Minecraft 1.20.1 / Forge: this page documents that build and its measurements. For the separate 170-control Minecraft 1.21.1 port, see Minecraft 1.21.1 NeoForge.

Thirteen controls for Embeddium 0.3.31+mc1.20.1 (mod id embeddium), all client-side. They cover chunk upload bookkeeping, vertex serializer lookups, hidden faces of connected-texture models, random block-model variants, sprite ticks and, since 1.0.30, section searches that are replayed while nothing changed, draw commands that are reused and merged, and a faster sort for translucent quads. These gains come on top of the work Embeddium already does; the hidden-face control reduced the tested chunk-meshing workload by 25 % with byte-identical vertices and indices. Embedded Block Entities, Oculus and Valkyrien Skies also modify RenderSectionManager; their call sites are preserved and the transformed class was audited. Every switch checks the code it would change against the tested build; another build is left untouched with one WARN line (see How the patches are applied).

Key Target Side Since Kind
embeddium_direct_upload_preparation Embeddium 0.3.31+mc1.20.1 CLIENT 1.0.14 opt
embeddium_draw_batch_cache Embeddium 0.3.31 (Minecraft 1.20.1) CLIENT 1.0.30 opt
embeddium_entity_sort_radix Embeddium 0.3.31 CLIENT 1.0.30 opt
embeddium_lazy_completed_jobs Embeddium 0.3.31+mc1.20.1 CLIENT 1.0.12 opt
embeddium_merged_draws Embeddium 0.3.31 (Minecraft 1.20.1) CLIENT 1.0.30 opt
embeddium_pending_upload_sum Embeddium 0.3.31+mc1.20.1 CLIENT 1.0.12 opt
embeddium_search_replay Embeddium 0.3.31 (Minecraft 1.20.1) CLIENT 1.0.30 opt
embeddium_section_cache_prefix_cleanup Embeddium 0.3.31+mc1.20.1 CLIENT 1.0.26 opt
embeddium_serializer_lookup_snapshot Embeddium 0.3.31+mc1.20.1 CLIENT 1.0.21 opt
embeddium_sprite_tick_frame_times Embeddium 0.3.31 (Minecraft 1.20.1) CLIENT 1.0.29 opt
embeddium_upload_classification Embeddium 0.3.31+mc1.20.1 CLIENT 1.0.12 opt
embeddium_visible_faces_first Embeddium 0.3.31 with Fusion 1.3.14+a CLIENT 1.0.26 opt
embeddium_weighted_pick_table Embeddium 0.3.31 (Minecraft 1.20.1 weighted block models) CLIENT 1.0.26 opt

embeddium_upload_classification

Since: 1.0.12 · Script (1.0.19 and earlier): furious10-classification.js Patched: me.jellysquid.mods.sodium.client.render.chunk.region.RenderRegionManager.uploadMeshes Mixin (since 1.0.20): RenderRegionManagerClassifyMixin

What upstream did. For each region upload, Embeddium created one stream-filtered list of the outputs that need a full upload, performed those uploads, then created a second stream-filtered list of index-only resort uploads. Two stream pipelines per region, every frame that has finished chunk builds.

What the patch does. Collects each list directly with ordered loops. The two passes are deliberately kept on opposite sides of the full-upload call, because an output's classification flag is mutable and may change during the first pass; turning this into a single early partition would be wrong. The result lists keep the unmodifiable contract of Stream.toList(). The fast path is limited to ordinary ChunkBuildOutput instances; custom subclasses fall back to the original pipeline.

What stays the same. Region iteration, the private upload calls, the read-only list contract, element identity and order, and classification changes between the phases. Native tests covered full, index-only and mixed uploads, buffer growth, failure paths, custom outputs, and the EBE and Oculus integrations.

Measured. classification-0 87.04 → 6.99 ns (−92.0 %, 448 → 0 B); classification-8 151.23 → 61.63 ns (−59.2 %, 504 → 160 B); classification-64 744.53 → 456.16 ns (−38.7 %, 1,032 → 768 B). These isolate the list processing, not GPU upload time.


embeddium_lazy_completed_jobs

Since: 1.0.12 · Script (1.0.19 and earlier): furious10-collector.js Patched: me.jellysquid.mods.sodium.client.render.chunk.RenderSectionManager (a new private lazy collector used by uploadChunks) Mixin (since 1.0.20): RenderSectionManagerMixin

What upstream did. collectChunkBuildResults allocated an ArrayList before polling its concurrent queue, and uploadChunks returned immediately when that list came back empty. On every frame with no completed chunk build, that is an allocation for nothing.

What the patch does. Adds a separate private lazy helper for uploadChunks only: it polls once first, uses an internal empty sentinel on a miss, and allocates the list only on the first result. The research had missed that destroy also calls the original collector; that method and the original collector/shutdown contract are left intact, and no public collection-returning API changes.

What stays the same. Exactly the same poll sequence, unwrap and exception order, upload processing and deletion. No separate queue-size or peek check, no new wait and no batch-size cap.

Measured. collector-idle 11.69 → 8.39 ns (−28.2 %, 32 → 0 B). This is primarily an idle-allocation cleanup (12 ms inclusive in the reference profile).


embeddium_pending_upload_sum

Since: 1.0.12 · Script (1.0.19 and earlier): furious10-arena.js Patched: me.jellysquid.mods.sodium.client.gl.arena.GlBufferArena.upload Mixin (since 1.0.20): GlBufferArenaPendingSumMixin

What upstream did. When the first upload attempt does not fit and the arena must grow, Embeddium built a second stream to sum the byte lengths of the remaining uploads.

What the patch does. Replaces only that sum with an ordered local long accumulator. The original LinkedList, the upload attempts, staging flushes, capacity policy, retry and final exception are untouched; signed long overflow, the division by stride and the final int cast are preserved.

Measured. arena-size-sum 187.37 → 78.06 ns (−58.3 %, 280 → 0 B). This runs only on the capacity-growth path, so the expected benefit in play is small.


embeddium_direct_upload_preparation

Since: 1.0.14 · Scripts (1.0.19 and earlier): furious14-arena.js plus a coordinated extension of furious10-classification.js; each switch is checked separately and all on/off combinations were tested Patched: the private mesh and resort upload routes of RenderRegionManager, feeding GlBufferArena.upload Mixins (since 1.0.20): GlBufferArenaMappedUploadsMixin, RenderRegionManagerDirectUploadMixin

What upstream did. To hand vertex and index uploads to the arena, Embeddium built stream pipelines over its own local ArrayLists (map, filter out null index buffers, collect) into a LinkedList inside the arena. There are three such owned routes: vertex uploads, mesh index uploads and resort index uploads.

What the patch does. Adds an internal list-based route that constructs the same ordered LinkedList directly. The existing public stream entry point stays for other callers.

What stays the same. When the arena buffer is captured, the separation between the vertex and index phases, omission of null indices, incremental removal of successful uploads, retry order, resize and staging-buffer flushes, and the bytes that reach the GPU (verified by readback after growth). This is separate from embeddium_upload_classification and embeddium_pending_upload_sum, which remain unchanged in behaviour.

Measured. Empty filtered index-upload preparation through the real arena/staging path 168.47 → 74.94 ns (−55.5 %, 400 → 88 B); non-empty GPU upload and resize behaviour were checked separately for correctness.


embeddium_serializer_lookup_snapshot

Since: 1.0.21 · Mixin: SerializerLookupSnapshotMixin Patched: me.jellysquid.mods.sodium.client.render.vertex.serializers.VertexSerializerRegistryImpl.find

What upstream did. Every copy between two vertex formats (text glyphs, entity cuboids, sprite expanders) looks up its serializer in a map guarded by a StampedLock, taking and releasing a read lock for each lookup.

What the patch does. Entries are only ever added to that map, never replaced or removed. A hit is now answered from an immutable copy of the map published through a volatile field; a miss runs the original locked lookup, which also refreshes the copy when the map has grown (Oculus writes its four serializers straight into the map, which the size check picks up). The method body is otherwise Embeddium's own (LGPL-3.0).

What stays the same. The serializer object returned: 400 randomized traces and a four-reader / one-writer stress test of 8 million lookups always returned the same object, and in the 1.0.21 client run every cached key and a missing key answered the same through the patch as through the locked map.

Measured. 7.6 → 5.7 ns per warm, uncontended lookup on the transformed class (−26 %; the review model measured 7.2 → 4.2 ns). In two client profiles the lookup was 0.17 % / 0.21 % of the render thread.


embeddium_section_cache_prefix_cleanup

Since: 1.0.26 · Target: Embeddium 0.3.31+mc1.20.1 · Side: CLIENT · Kind: opt

Mixins: SectionCachePrefixCleanupMixin

Embeddium keeps up to 512 cloned chunk sections for its chunk builders and, every frame, checked every one of them for having gone unused for 5 seconds. The cache is ordered by last use (a section that is used again moves to the end), so the expired sections are always the oldest ones at its front: the check now removes from the front and stops at the first section that is still in use. The same sections are removed in the same order. If the clock ever runs backwards, or the order is not confirmed, the original full check runs until a check of the whole cache confirms the order again.

Measured: 3.1 million offline checks against Embeddium's own class, 0 differences; a cleanup with 512 cached sections 772 -> 45 ns. The full check was 2.0% of the render thread while flying over new terrain.


embeddium_visible_faces_first

Since: 1.0.26 · Target: Embeddium 0.3.31 with Fusion 1.3.14+a · Side: CLIENT · Kind: opt

Mixins: VisibleFacesFirstMixin, EntryModelsMixin, VariantAllowlistMixin

Embeddium builds each face's quads first and asks afterwards whether the face can be seen, throwing the quads away when it cannot. For Fusion's connected-texture models that ran the whole texture pipeline for faces hidden behind other blocks. For Fusion's random-offset model groups (the way Fusion's Embeddium renderer hands over its model-modifier blocks) whose parts are all vanilla simple models, Fusion's base models or vanilla weighted models of those two, a hidden face now returns no quads straight away, and a visible face is not asked twice. Hidden faces never draw anything, and a visible face's quads do not depend on what was skipped, because the random source is reset for every face. Every other model, plain vanilla and Fusion models outside such a group included, keeps the original order.

Measured: in game with this pack's resource packs, meshing the 733 non-empty sections around the player took 3.9 s instead of 5.2 s per full rebuild (median of 6 alternating rounds), and every section's vertex and index data came out byte-identical.


embeddium_weighted_pick_table

Since: 1.0.26 · Target: Embeddium 0.3.31 (Minecraft 1.20.1 weighted block models) · Side: CLIENT · Kind: opt

Mixins: WeightedPickTableMixin

A block with weighted random models (xali's sand has 28 variants, netherrack 36) picks its variant for every face and render-layer query by walking the list until the drawn number runs out. The model now keeps the running totals of its weights, built once, and finds the variant with a short search. The random draw is the same call as before; weights are never negative, so the first total above the drawn number is exactly where the walk stopped, and a negative draw still picks the first entry. With Embeddium's own weighted-model change turned off (its mixin.features.model setting) there is nothing to replace, and the models stay exactly as Minecraft ships them (1.0.30-1.0.33 stopped the game there).

Measured: in a harness running Embeddium's and Fusion's own classes, a mix of real variant lists picked in 34 ns instead of 81 (2.4x faster) and netherrack's 36 variants 1.6x faster; lists that nearly always stop at the first entry stay about the same; no allocation. Picking was about 2-3% of chunk meshing time.

In-game section comparisons found no differing vertex/index data. The isolated in-game timing difference was inside the rig's noise; the method-level speedup above is not a demonstrated whole-meshing or FPS gain. The separate BonsFusion changes are not part of this JAR.


embeddium_sprite_tick_frame_times

Since: 1.0.29 · Target: Embeddium 0.3.31 (Minecraft 1.20.1) · Side: CLIENT · Kind: opt

Mixin: SpriteTickerFrameTimesMixin

Invisible animated textures advance from a copy of their frame times. With Embeddium's "animate only visible textures", every animated texture that was not drawn since the last tick still steps its frame counter once per tick. To read the current frame's length Embeddium walks four separate objects (animation, frame list, its array, frame), and with a pack's 2,000-3,000 animated textures they are cold in the CPU cache every tick. The frame lengths never change (the list and its entries are immutable), so each texture now copies them once into a small array and steps from that: the same counters, in the same order, to the same values. Textures on screen, the option turned off, and any list that is not one of the immutable kinds still use Embeddium's code.

Measured: offline on the transformed classes 3.6 million checks over 3,000 texture pairs and 400 ticks (423,468 steps on the new path) all identical; 94-138 -> 49-79 ns per texture per tick with a cold cache and 17-27 -> 6-14 ns warm (2,500-10,000 textures, the client's ZGC and ParallelGC), no allocation. Embeddium's step was 2.5-7.3% of the render thread in the 1.0.26 client recording.


embeddium_draw_batch_cache

Since: 1.0.30 · Target: Embeddium 0.3.31 (Minecraft 1.20.1) · Side: CLIENT · Kind: opt

Mixins: DefaultChunkRendererBatchCacheMixin, SectionRenderDataStorageVersionMixin, ChunkRenderListAccessMixin

Terrain draw commands are reused while nothing they come from changes. Every frame, and with Oculus shadows twice, Embeddium rebuilds each region's list of draw commands from its sections' mesh data. Each region pass now keeps the commands of its last two builds with everything they were built from: its mesh data's version (bumped by every upload, removal, resort and buffer move), the list of sections, the pass, the face-culling setting and the camera's position relative to the region's sections as Embeddium's face test sees it. When all of these are equal the kept commands are copied back instead of rebuilt; they are the same commands, byte for byte. A build whose inputs changed costs a little more than before (the copy that is kept).

Idea: Sodium 0.7 draw-batch caching, idea text only

Measured: 2,690,603 checks over 300,000 builds (uploads, removals, resorts, buffer moves, list and camera changes, both passes, switch toggles), 0 mismatches; unchanged regions 80-89% less time (hot 503 -> 58 us, cold 1,498 -> 262 us per frame for 432 regions), changed regions +14-28%; 2.0% of the render thread with a still camera (1.6% shadow pass)


embeddium_merged_draws

Since: 1.0.30 · Target: Embeddium 0.3.31 (Minecraft 1.20.1) · Side: CLIENT · Kind: opt

Mixin: DefaultChunkRendererMergeMixin

Touching terrain draw commands are merged. Embeddium issues one draw command per face direction of every section, even when consecutive commands read vertex ranges that touch (a section's face directions are stored back to back). In the solid and cutout passes such commands are now issued as one, which draws the same triangles in the same order with the same vertex numbers. Sorted translucent passes are left alone. A shader pack that reads per-draw numbers (gl_PrimitiveID, gl_DrawID, gl_BaseVertex) keeps the unmerged commands; the pack's files are checked once when it is loaded. With embeddium_draw_batch_cache on, the merge runs only when a region's commands are rebuilt.

Idea: Sodium 0.7 combined draw commands, idea text only

Measured: 1,000,000 command lists (311 million commands, edge cases) and 120,701 real builds, each merge equal to an independent triangle-by-triangle expansion, 0 mismatches; draw commands 54,418 -> 30,680 (main view) and 93,107 -> 15,640 (shadow pass) per frame in the bench world; the merge costs 3.4-4.3 ns per command when it runs; in game with a shader pack and the graphics card at 96% load, about 1% more frames per second (ahead in all three paired runs)


embeddium_search_replay

Since: 1.0.30 · Target: Embeddium 0.3.31 (Minecraft 1.20.1) · Side: CLIENT · Kind: opt

Mixins: OcclusionCullerReplayMixin, RenderSectionManagerEpochMixin, ViewportAccessMixin, RenderSectionInfoEpochMixin

A section search whose inputs are all unchanged replays the previous one. Every frame Embeddium searches the chunk sections outward from the camera to find the ones to draw, and with Oculus shadows it searches twice (shadow pass and main view). With a still view it repeats the same walk. A search depends only on the section links, the sections' visibility data, the camera section, the occlusion setting, the distance test's inputs and the frustum's own numbers. Once a view stands still, the switch records a search (the sections it reported, in order, with their answers) under a key of all those inputs, bit for bit. A later search with exactly that key stamps the same sections with the new frame and reports them to Embeddium's own list builder in the same order, so render lists, rebuild queues and the translucency check come out as before. Any change, an unknown frustum or list builder, or another mod's mixin on these classes runs the full search; a view that keeps moving is only checked now and then.

Idea: Sodium 0.6.1 graph-search gating and Iris's shadow-list reuse, idea text only

Measured: 120,209,862 checks against the shipped class over 12,000 searches (still, moving, turning, sun steps, link and build-state changes, loads, switch toggles), 0 mismatches; still view 58-68% less time per search (hot 65 -> 21 and 106 -> 37 us, cold 193 -> 78 and 378 -> 158 us); moving view +0.5-0.6 us per search hot, +4-6 us cold (1-3%); the search is 14.3% of the render thread with a still camera in the 1.0.26 recording (10.6% shadow pass)


embeddium_entity_sort_radix

Since: 1.0.30 · Target: Embeddium 0.3.31 · Side: CLIENT · Kind: opt

Mixin: VertexSorterRadixMixin

Translucent quad sort by packed keys. Embeddium orders every translucent quad buffer that is sorted on upload (translucent entities, items, beacon beams; with Oculus each batched segment) with a merge sort over an index array that compares squared distances through the index. For keys without NaN that is exactly the stable "farthest first" order. From 64 quads on, the switch packs each distance and its index into one long and sorts those with a byte-wise radix sort, which gives that same order; shorter buffers and any NaN key still go through Embeddium's own sort, and only the returned array is allocated.

Idea: Sodium 0.7.0's faster vertex sorting (radix sort), idea text only

Measured: 218,751 checks + 2 x 265,634 (both packed-sort paths), 0 mismatches; 1.5-1.7x at 96 quads, 2-2.8x at 256, 3-3.9x at 1,024, 3.5-5.8x at 16,384 (sort incl. distances, ParallelGC and ZGC), a third less allocation; below 64 quads unchanged (one length test); 0.25% of the render thread in an entity-heavy scene before, near 0 elsewhere


Candidates that were measured and withheld

For completeness, two Embeddium ideas were fully implemented, passed their correctness tests, and were not shipped because native timing regressed or did not improve:

  • Embeddium scalar section bounds (ClonedChunkSection.copyBlockEntities, 1.0.12 batch): HotSpot already removed the temporary BoundingBox; the final run was +23.2 % slower with equal allocation.
  • Embeddium sparse draw commands (DefaultChunkRenderer.addDrawCommands, 1.0.14 batch): identical command bytes and faster sparse sections, but dense batches regressed about 4.6 % (ZGC) and 13.6 % (repeat).

Their patches and switches are absent from every release; the original methods stay active.

Bons and Furious

Minecraft 1.20.1 / Forge 1.0.34

Compatibility

Controls by mod

Links

Clone this wiki locally