perf: reduce TensorView metadata allocations#33
Merged
Conversation
voltjia
force-pushed
the
perf/reduce-tensor-view-allocations
branch
from
July 22, 2026 23:29
f3cde2a to
b1187cc
Compare
voltjia
marked this pull request as ready for review
July 23, 2026 02:00
TensorView metadata allocations
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
TensorView::Shape/TensorView::Stridesconstructor overloads so owned metadata can be moved intoTensorViewwithout redundant copies.Motivation
TensorViewis intended to be a lightweight framework-neutral tensor metadata layer, but common rvalue, initializer-list,operator[], andT()paths created temporary vectors and then copied them into the result. Non-empty rank-2 views therefore performed three or four metadata allocations where two owned buffers are sufficient.The existing public
Shape = std::vector<Size>andStrides = std::vector<Stride>aliases, generic constructors, object data members, and behavior are preserved in this first phase.Related issue: N/A - this performance work is not linked to an existing issue.
Type of Change
feat- new feature / new backend capability / new public APIfix- bug fixperf- performance improvement without behavior changerefactor- code restructuring without behavior changetest- adding or fixing tests onlydocs- documentation onlybuild/ci- build system or CI configurationchore- tooling, formatting, or other non-code changes!in the Conventional Commits prefix or aBREAKING CHANGE:footer)Platforms Affected
WITH_CPU)WITH_NVIDIA)WITH_ILUVATAR)WITH_HYGON)WITH_METAX)WITH_MOORE)WITH_CAMBRICON)WITH_ASCEND)Smoke Build and Test Result
CPU Release build on the Linux x86_64 NVIDIA validation host (GCC 11.4.0):
NVIDIA Release build in the existing
accelerator-dev/nvidia:latestimage on 8x A100-SXM4-80GB (CUDA 13.1.80, GCC 13.3.0):Test Results on Supported Platforms
10/10passed, including allocation, performance, install, and installed-consumer tests8/8smoke tests passed; the final clang-format-only amend rebuilt successfully in the same imageFull `ctest` output
The amended head rebuilt successfully in the same NVIDIA image.
Benchmark / Performance Impact
Harness:
tests/performance/perf_tensor_view.cc, Release CPU backend on the same Linux x86_64 validation host, GCC 11.4.0, rank 2, shape[32, 64], float32 where applicable, 1,000 warmup iterations, 200,000 measured iterations per sample, and 7 samples. Values below are mean nanoseconds per operation.operator[]T()numel()controlThe allocation regression test verifies the corresponding metadata allocation reductions: rvalue full metadata, initializer-list,
operator[], andT()paths decrease from four allocations to two; rvalue default-stride construction decreases from three to two. Copy and move construction remain at two and zero allocations, respectively.Notes for Reviewers
ShapeandStridesremainstd::vector, generic tensor-like/range constructors remain available, andTensorViewdata member types and order are unchanged.operator newinterposition.TensorViewcorrectness fixes are deliberately out of scope.git diff --checkand both clang-format 21 CI checks pass.