Skip to content

Benchmarks

Arnel Robles edited this page Sep 28, 2026 · 1 revision

Benchmark Results

Core Mapping Performance

BenchmarkDotNet 0.13.12, .NET 8, five warmup iterations and twenty measured, on two architectures because one is not evidence of anything portable. Reproduce with dotnet run -c Release --project tests/Mapsicle.Benchmarks -- --core, which is the job the CI gate runs.

The arm64 figures are the median of six runs on an idle 4-core Ampere VM. BenchmarkDotNet's error column describes how consistent the iterations were inside one process, which is not the same as whether the run reproduces. On that machine it does not, closely: repeating the identical commit moved the Mapperly collection row by 36 percent, and Mapperly is a source generator this project has never touched. A single run of any of these rows is one sample.

Single object, five properties:

Runtime Manual Mapsicle AutoMapper Mapperly Mapsicle vs AutoMapper
x64 Linux (CI runner) 18.6 ns 57.8 ns 83.2 ns 18.8 ns 1.44x faster
arm64 Linux (Ampere VM) 23.0 ns 101.2 ns 131.3 ns 28.2 ns 1.30x faster

Mapperly is not a competitor, it is a different trade, and it wins the one this table measures. At 18.2 ns against hand-written code's 18.3 ns it is not close to manual, it is indistinguishable from it, because a source generator emits ordinary C# assignments at compile time and leaves no delegate, no cache lookup and no indirection at runtime. Mapsicle and AutoMapper both build an expression tree, compile it, cache it, look it up and invoke through it. That apparatus is the entire 2.5x to 3x gap, and no runtime mapper can close it, because the apparatus is what makes it a runtime mapper.

What you buy with it: Mapperly needs a partial class with a [Mapper] attribute and a declared method for every pair, all known at compile time. It cannot map a Dictionary<string, object> into a type chosen at runtime, or a collection whose items turn out to have different runtime types, because there is nothing for it to generate against. Mapsicle needs no configuration and resolves types as it meets them.

If you need the compiler to prove every mapping exists, choose Mapperly. That guarantee is real and Mapsicle does not offer it: a pair it cannot generate warns and falls back rather than failing the build.

The table below measures Mapsicle's runtime engine, which is what an undeclared pair uses. Declare the pair and it is level with hand written code instead, which is measured in Compile-time mapping further down. Read the undeclared row as the floor rather than the whole story.

Other scenarios:

Scenario Mapsicle AutoMapper Mapperly vs AutoMapper
Flattening, x64 64.9 ns 101.2 ns 21.1 ns 1.56x faster
Flattening, arm64 110.0 ns 140.2 ns 33.0 ns 1.28x faster
Collection (100), x64 2,175 ns 2,618 ns 1,933 ns 1.20x faster
Collection (100), arm64 4,311 ns 4,696 ns 2,894 ns 1.09x faster
Collection (10,000), arm64 481 us 1,134 us 322 us 2.36x faster
Deep nesting (15 levels), arm64 626 ns 5,145 ns 282 ns 8.22x faster

Allocation per operation matches hand-written code for single objects and flattening (48 B and 56 B, the destination and nothing else). On a collection Mapsicle allocates 5,656 B against AutoMapper's 6,992 B, about 19 percent less, and the same as source-generated Mapperly.

About the collection rows. A List<T> is mapped by a loop compiled for its element type, which is where most of that number comes from. Before that loop existed these rows were 1.07x slower on x64 and 1.04x faster on arm64, which is to say parity. Arrays, and lists whose element type is object, an interface or abstract, keep the older loop and its older cost, because a loop compiled for a type no element actually is sends every element down a slower path. If collection throughput at this size is what your workload is bounded by, Mapperly is still faster than both, though at 2,175 ns against its 1,933 the gap on x64 is now about 12 percent.

At ten thousand the picture changes, and not because the per-element cost changed. AutoMapper allocates 742 KB there against Mapsicle's 560 KB, enough to reach generation 2 collections while Mapsicle stays in 0 and 1. The 2.36x is mostly that.

Three things worth more than the table:

  • The job length changes the answer. These numbers come from five warmup iterations and twenty measured. The same commits under BenchmarkDotNet's ShortRun, three and three, put the arm64 collection row at 1.05x slower rather than 1.05x faster, because three warmup iterations do not get compiled delegates to steady state and the mapper that compiles more pays for it. On a hosted runner ShortRun gave an interval of plus or minus 43 percent of the mean. The gate used to run that job. It does not now.
  • Mapsicle.Fluent is not on this table. A complex object through the fluent configuration API is about 1.03x AutoMapper, and a hundred of them about 1.22x, because the fluent path maps a collection element by element rather than through the compiled loop. The static API is what the rows above measure.
  • An earlier version of this table claimed 2.1x on single objects and rough parity on collections. Neither held when the benchmark was re-run. It went unnoticed because CI measured the comparison, printed it and exited zero regardless. CI now fails when the comparison changes direction, and it reads BenchmarkDotNet's own summary rather than a stopwatch loop, so the published numbers and the gated numbers come from one source.

Edge Case Performance

Scenario Mapsicle AutoMapper Mapperly Notes
Deep Nesting (15 levels) 626 ns 5,145 ns 282 ns All safe; Mapsicle 8.22x AutoMapper
Circular References Handled by default Opt in via PreserveReferences() Opt in via UseReferenceHandling Only Mapsicle needs no configuration
Large Collection (10K) 0.48 ms 1.13 ms 0.32 ms Mapsicle 2.36x AutoMapper
Parallel (1000 threads) ✅ Thread-safe ✅ Thread-safe ✅ Thread-safe All thread-safe
Cold Start Medium Slow None Mapperly pre-compiled

Performance Optimizations (v1.1+)

Optimization Improvement Status
TypedMapperCache<T,D> Zero-allocation generic cache ✅ NEW
MapTo<TSource,TDest>() Strongly-typed mapping, no boxing ✅ NEW
Skip depth tracking for simple No overhead for flat types ✅ NEW
Lock-free cache reads Eliminates contention ✅
Collection mapper caching +20% for collections (v1.1) ✅
PropertyInfo caching +15% faster cold starts ✅
Primitive fast path Skips depth tracking ✅
Cached compiled actions No runtime reflection ✅
LRU cache option Memory-bounded in long-run apps ✅
Collection pre-allocation Capacity hints for known sizes ✅

Memory & Cache Statistics (v1.1+)

// Enable memory-bounded caching
Mapper.UseLruCache = true;
Mapper.MaxCacheSize = 1000;  // Default

// Monitor cache performance
var stats = Mapper.CacheInfo();
Console.WriteLine($"Cache entries: {stats.Total}");
Console.WriteLine($"Hit ratio: {stats.HitRatio:P1}");  // Only when LRU enabled
Console.WriteLine($"Hits: {stats.Hits}, Misses: {stats.Misses}");
Feature Mapsicle (Unbounded) Mapsicle (LRU) AutoMapper
Memory Bounded ❌ ✅ ❌
Cache Statistics Entry count only Full stats ❌
Configurable Limit ❌ ✅ ❌
Lock-Free Reads ✅ ✅ Partial

The claims are gated

The comparison above is not only published, it is checked. dotnet run -c Release --project tests/Mapsicle.Benchmarks -- --quick runs Mapsicle against AutoMapper and returns a non-zero exit code when a ratio moves outside its bound. CI runs it on every pull request.

Two more gates guard the other half of the pitch:

  • core-has-no-dependencies packs Mapsicle and fails if the nuspec declares a single dependency.
  • licence-boundary fails if anything under src/ references AutoMapper, which is RPL-1.5 or a paid licence. It is compared against in tests/, which is never packed.

Run Benchmarks Yourself

cd tests/Mapsicle.Benchmarks
dotnet run -c Release              # Full suite
dotnet run -c Release -- --quick   # Smoke test
dotnet run -c Release -- --edge    # Edge cases only

Clone this wiki locally