You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Cost of evaluating an analytical field via a user-specified python function
#1322
I made claude write a benchmark to see the cost of the current way to call python from c++ for init functions, compared with equivalent plain c++ and more optimized approach of calling the python function. Here is the kind of obtained results for evaluating a cos(x)sin(y)-like function :
Times in microseconds per field fill, percentages relative to cpp_direct.
2d / sincos
variant
64
128
256
512
nodes
4.7k
17.6k
67.9k
266.8k
cpp_direct
63.0
232.3
895.9
3526
cpp_via_vectors
70.6 (+12%)
261.4 (+13%)
1014 (+13%)
5213 (+48%)
py_now
254.9 (+305%)
950.7 (+309%)
4140 (+362%)
22038 (+525%)
py_np
73.8 (+17%)
270.4 (+16%)
1056 (+18%)
5953 (+69%)
py_np_cached
65.3 (+3.6%)
238.0 (+2.4%)
928.6 (+3.6%)
3756 (+6.5%)
py_np_cached_flat
64.0 (+1.6%)
234.1 (+0.8%)
911.8 (+1.8%)
3656 (+3.7%)
py_np_out
63.6 (+0.9%)
232.6 (+0.1%)
901.8 (+0.7%)
3581 (+1.5%)
3d / sincos
variant
16
32
64
nodes
8.4k
48.0k
319.1k
cpp_direct
167.6
959.2
6368
cpp_via_vectors
188.4 (+12%)
1079 (+13%)
8963 (+41%)
py_now
691.8 (+313%)
4230 (+341%)
37943 (+496%)
py_np
197.4 (+18%)
1120 (+17%)
9973 (+57%)
py_np_cached
174.4 (+4.1%)
997.6 (+4.0%)
6918 (+8.6%)
py_np_cached_flat
169.6 (+1.2%)
972.7 (+1.4%)
6688 (+5.0%)
py_np_out
169.2 (+1.0%)
957.2 (-0.2%)
6526 (+2.5%)
This shows that the best approach to wrap python compatible with the current API py_np_cached_flat is always a few percent more costly than plain cpp. py_np_out is slightly more efficient but is not compatible with the current API: the former requires to define a function that modify an argument it receives in-place, whereas current API requires to define a function that returns the result.
The question is do we agree to pay those extra few percent in order to enable the user to specify the external B0 field via python ? For a time-dependan external fieldt, there is 2 vecfields (B0 and its time-derivative) per Runge-Kutta iteration (if we want to be rigorous) to evaluate.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
I made claude write a benchmark to see the cost of the current way to call python from c++ for init functions, compared with equivalent plain c++ and more optimized approach of calling the python function. Here is the kind of obtained results for evaluating a cos(x)sin(y)-like function :
Times in microseconds per field fill, percentages relative to
cpp_direct.2d / sincos
cpp_directcpp_via_vectorspy_nowpy_nppy_np_cachedpy_np_cached_flatpy_np_out3d / sincos
cpp_directcpp_via_vectorspy_nowpy_nppy_np_cachedpy_np_cached_flatpy_np_outThis shows that the best approach to wrap python compatible with the current API
py_np_cached_flatis always a few percent more costly than plain cpp.py_np_outis slightly more efficient but is not compatible with the current API: the former requires to define a function that modify an argument it receives in-place, whereas current API requires to define a function that returns the result.The question is do we agree to pay those extra few percent in order to enable the user to specify the external B0 field via python ? For a time-dependan external fieldt, there is 2 vecfields (B0 and its time-derivative) per Runge-Kutta iteration (if we want to be rigorous) to evaluate.
I would say yes, but what do you think ?
All reactions