Hello Go maintainers,
Related to #61052 #18138, #7514, but bringing more data.
This was discussed in #57175 recently, raising an issue for a future reference.
We have recently built capabilities to measure the cost of stack growth across our fleet. This cost is significant (3.9%, core weighted) vs a traditionally costly thing like GC, which costs us 7.3%.
We would like to solve this, we have some ideas, but we hope that by providing the data we can land a solution that is acceptable witihin OSS. We’re happy to contribute once you advise us on the best approach.
Fleet wide costs of runtime.newstack:
| Scope |
Cores of the fleet |
Copystack cost |
More than 3% |
More than 2% |
More than 1% |
Average |
| Top10 |
33% |
30% |
8/10 |
10/10 |
10/10 |
4.3% |
| Top20 |
48% |
42% |
12/20 |
14/20 |
18/20 |
3.7% |
| Top100 |
82% |
80% |
62/100 |
73/100 |
88/100 |
3.7% |
| ~600 |
100% |
100% |
334/600 |
406/600 |
509/600 |
3.9% |
Top 13 services goroutine stack distribution:
| Service |
Mean |
P50 |
P75 |
P90 |
P99 |
Max |
Starting stack size |
| 1 |
14.2 KB |
12.1 KB |
21.1 KB |
22.2 KB |
23.3 KB |
25.1 KB |
4kb |
| 2 |
12.5 KB |
10.6 KB |
20.1 KB |
25.9 KB |
31.4 KB |
31.6 KB |
4kb |
| 3 |
15.5 KB |
13 KB |
23.2 KB |
24.2 KB |
24.8 KB |
24.9 KB |
4kb |
| 4 |
16.0 KB |
12 KB |
26.9 KB |
30.7 KB |
37.8 KB |
38.0 KB |
4kb |
| 5 |
19.0 KB |
19.3 KB |
23.6 KB |
32.3 KB |
42.9 KB |
43.1 KB |
4kb |
| 6 |
17.0 KB |
11.9 KB |
21.6 KB |
36.1 KB |
72.5 KB |
73.4 KB |
4kb |
| 7 |
13.1 KB |
12 KB |
16.6 KB |
22.1 KB |
56.5 KB |
56.9 KB |
|
| 8 |
10.1 KB |
7.1 KB |
12.0 KB |
21.1 KB |
22.9 KB |
23.1 KB |
|
| 9 |
6.9 KB |
6.1 KB |
11.7 KB |
12.4 KB |
13.1 KB |
13.6 KB |
|
| 10 |
5.0 KB |
3.5 KB |
5.4 KB |
10.5 KB |
22.5 KB |
22.6 KB |
2kb |
| 11 |
24.9 KB |
26.6 KB |
30.7 KB |
31.3 KB |
31.4 KB |
32.2 KB |
4kb |
| 12 |
6.4 KB |
2.7 KB |
7.0 KB |
21.2 KB |
21.4 KB |
21.4 KB |
2kb |
| 13 |
12.9 KB |
9.9 KB |
22.2 KB |
22.7 KB |
24.5 KB |
38.4 KB |
4kb |
Top13 services goroutine stack categorization:
| Service |
≥ 4KB |
≥ 8KB |
≥ 16KB |
≥ 24KB |
≥ 32KB |
≥ 48KB |
≥ 64KB |
| 1 |
99.5% |
73.9% |
47.1% |
0.5% |
|
|
|
| 2 |
99.75% |
61.4% |
27.2% |
15.2% |
|
|
|
| 3 |
100% |
73.6% |
44.0% |
13.6% |
|
|
|
| 4 |
98% |
67.4% |
39.9% |
32.4% |
4.7% |
|
|
| 5 |
97.2% |
82.0% |
66.5% |
24.9% |
10.5% |
|
|
| 6 |
98.4% |
74.8% |
42.6% |
22.1% |
10.9% |
3.5% |
1.2% |
| 7 |
98.4% |
65.9% |
34.1% |
4.1% |
3.3% |
3.3% |
|
| 8 |
99.1% |
49.6% |
13.3% |
|
|
|
|
| 9 |
61.1% |
34.3% |
|
|
|
|
|
| 10 |
46.6% |
12.7% |
2.5% |
|
|
|
|
| 11 |
100% |
91.3% |
84.5% |
80.7% |
0.6% |
|
|
| 12 |
42% |
24.0% |
10.0% |
|
|
|
|
| 13 |
90% |
75.5% |
36.2% |
1.9% |
0.2% |
|
|
For now, we have a per-service knob that overwrites those values and have been rolling overrides to either 16kb or 32kb for the top services. Those changes worked as expected, we see dramatic CPU drop and corresponding stack memory increase - in our case it’s usually <200mb of memory and it’s acceptable for those services as their heap sizes are much larger.
For a long term plan we’re considering several ideas, we might be trying some of them internally.
| Ideas |
Comment |
| New stats and that’s it (number of growth/shrink, histogram?) |
|
| Emit copystack cost, similar to the GC cost. |
|
| Min-size knob |
Dumb ideas 😞 |
| Override min 4k for everyone? |
| “High memory mode” similar to GOMEMLIMIT? |
| Round up to n+1 size |
| P75 instead of mean would be good enough |
Reduces growth ops by ~50% in our fleet. |
| Use new metrics: growth vs goroutine start to adjust dynamically |
Complicated (?) |
| Start higher: allow for faster shrinking |
“Optimistic presizing” |
| Start higher: shrink for “stable” goroutines |
| Limit scope: Exclude long running goroutines |
“Excluding outliers” |
| Limit scope: Only include goroutines that had stack growth in the last N cycles |
| Limit scope: Only include goroutines that started (or died?) in the last N cycles |
Hello Go maintainers,
Related to #61052 #18138, #7514, but bringing more data.
This was discussed in #57175 recently, raising an issue for a future reference.
We have recently built capabilities to measure the cost of stack growth across our fleet. This cost is significant (3.9%, core weighted) vs a traditionally costly thing like GC, which costs us 7.3%.
We would like to solve this, we have some ideas, but we hope that by providing the data we can land a solution that is acceptable witihin OSS. We’re happy to contribute once you advise us on the best approach.
Fleet wide costs of runtime.newstack:
Top 13 services goroutine stack distribution:
Top13 services goroutine stack categorization:
For now, we have a per-service knob that overwrites those values and have been rolling overrides to either 16kb or 32kb for the top services. Those changes worked as expected, we see dramatic CPU drop and corresponding stack memory increase - in our case it’s usually <200mb of memory and it’s acceptable for those services as their heap sizes are much larger.
For a long term plan we’re considering several ideas, we might be trying some of them internally.