Celestial v1.0 is released 馃帀
What Changed
- feat: add benchamrk eval for the project #2
- feat: add output table for benchmark_test #3
- feat: Add GitHub Actions workflow for Go project #5
- feat: update test result with avg score #6
v1.0 benchmark
Note
Celestial benchmark was tested on 13th Gen Intel(R) Core(TM) i7-13620H
Benchmark Workload GoProcs Workers Batch AvgTasks Samples AvgNs/task AvgTasks/s AvgSpeedup AvgEfficiency
BenchmarkAtomicSliceDispatcher/CPUWork CPUWork 16 15 - 36880523 5 31.61 31.65M - -
BenchmarkAtomicSliceDispatcher/TinyWork TinyWork 16 15 - 77799407 5 15.48 64.59M - -
BenchmarkBatchDispatcher/CPUWork CPUWork 16 15 64 40916848 5 28.68 34.92M - -
BenchmarkBatchDispatcher/TinyWork TinyWork 16 15 64 1000000000 5 0.43 2.34B - -
BenchmarkChannelDispatcher/CPUWork CPUWork 16 15 - 2974709 5 401.80 2.49M - -
BenchmarkChannelDispatcher/TinyWork TinyWork 16 15 - 4058515 5 293.10 3.41M - -
BenchmarkEvaluationScore/AtomicSliceCPUWork CPUWork 16 15 - 32768 5 - - 7.81x 0.52
BenchmarkEvaluationScore/BatchCPUWork CPUWork 16 15 64 32768 5 - - 7.47x 0.50
BenchmarkEvaluationScore/ChannelCPUWork CPUWork 16 15 - 32768 5 - - 0.81x 0.05
BenchmarkSequential/CPUWork CPUWork 16 1 - 3935243 5 303.02 3.30M - -
BenchmarkSequential/TinyWork TinyWork 16 1 - 1000000000 5 1.10 905.32M - -
deficiency
Channel distribution is more than an order of magnitude slower than Atomic/Batch. The reason is that each task goes through:
jobs channel -> worker -> results channel -> consumer
The overhead associated with channel synchronization, scheduling, and result transmission is very high.
Optimization approach: The core API should prioritize RunSlice, which is atomic index distribution. Run is suitable for streaming tasks but not for high-frequency, small tasks.
Translated with DeepL.com (free version)