Inclusion of additional cell_methods statistics for climate normal and anomaly products #471
Replies: 13 comments 7 replies
|
Thanks for your comment. Those three bullet points have different kinds of answer!
Best wishes Jonathan |
|
These three points are all non-trivial and worth considering carefully. They can all be treated with the current conventions, but possibly could be done better if we made enhancements. I can see there's an argument for allowing A trend is an intensive quantity, like a mean. The default cell method for intensive quantities is I don't think we ought to introduce the numerical percentage into the Jonathan |
|
Hi @Roebeling and @JonathanGregory standard_errorI'm fine with introducing a new cell method. The aforementioned uncertainty proposal suggests deprecating the use of the standard_error standard name modifier for the use of expressing uncertainty, and in general I'd be happy if we didn't introduce any more standard name modifiers! (Edit August 2026 - the uncertainty proposal is now available: cf-convention/cf-conventions#660) The standard error of a sampling distribution is the standard deviation of a statistic calculated from that sample, so I don't think that the name tendencyIf we didn't have any I think I agree with Jonathan's analysis, but find the use of a "tendency" cell method confusing - it's not intuitive to me that the cell method does not indicate that a tendency of the physical quantity (itself tendency) has been carried out, and we'd have to restrict its use to particular standard names. For me "point" makes more sense, as it is implied anyway and doesn't require any special treatment - i.e. percentileI agree with Jonathan on the use of a coordinate or scalar coordinate variable. |
|
Dear Rob @Roebeling Sorry for slow reply. I have been on holiday. Percentiles@TomLav is correct about what I meant. He shows an example with multiple percentiles. The method could also be used for a single percentile e.g. or Either way, the special case of Your use-case is of climatological percentiles. To indicate this, I think that @TomLav's example should have I agree that the natural choice for the standard name would be Standard errorI agree with @davidhassell that TrendsThe description of The trend (from regression against time) of X is an estimate of the time-mean tendency of X within the interval of interest, assuming that the tendency However, if there are temporal variations in X not due to the underlying constant trend, the instantaneous The regression trend does not come from one instantaneous value ( Therefore I'd like to suggest a new cell method of Best wishes Jonathan |
|
Just to let you know that the aforementioned uncertainty proposal is now available: cf-convention/cf-conventions#660 |
|
Dear Rob @Roebeling @davidhassell and I have discussed my suggestion at the end of my last posting. As a result, I'd like to replace the name I suggested for a new cell method. Instead, I suggest we introduce a cell method of If you think that the Best wishes Jonathan |
|
Dear @JonathanGregory Thanks for the percentile explaination. Following your suggestion to the comment of @TomLav the idea became clear to me. Meanwhile I already experimented a bit, and now implememted the percentile part as you suggest in your first option ie Example: Metadata of percentile variables variables: Cheers, Rob |
|
I for one can't quite get comfortable with the the method being called "intensive". To my mind "intensive" doesn't care how much of something you have as long as any additional substance has the same properties as the original bit. For a trend, this would not be the case. Suppose you have a time interval with a trend x and a second interval where the time-series is identical to that in the first time interval (so again a trend x). If you then connect the two intervals and compute the trend, you probably get something about half of the trend in each of the individual intervals (i.e., nearer x/2). How should I look at this to make sense of characterizing the trend as an "intensive" property? |
|
Dear Karl I agree that "intensive" and "extensive" are not self-explanatory words, and therefore not ideal for CF vocab. We do already use them in App E—we always have as far as I remember—because they're scientific terminology e.g. described by Wikipedia. We could certainly include definitions like Wikipedia has, but I'm sure you know all this. Can you think of better words to use which characterise this distinction? I agree with your interpretation of "intensive", and I understand your example, but I think there's an obvious counter-example. Suppose you have a timeseries with a fairly constant trend in it, and divide it in two. The trend in each of the two shorter series will be about the same as the trend in the longer series. This fits the definition of "intensive". It's like the mean. If you have an intensive quantity with no trend, you can estimate the mean from any section, but if there's random variability, the estimate will be less accurate if you chop up the dataset. I think that's the sort of situation in which your example would most likely occur. Best wishes Jonathan |
|
Could we simply use the generic term "property" or "statistic" to point out which dimension(s) the standard name is referring to? so a "tendency" representative of a time-interval (and interpretable as a "trend" over that time interval) would have cell_methods: "time: property". (or "time: statistic"). Of course maybe you don't need a cell_methods at all in this case, since trend presumably implies "over time". We would encourage further explanation of the "property" (or "statistic") by asking the data writers to include some explanatory text in parentheses: e.g., for a variable like |
|
Dear Jonathan, Karl, David, On the cell_methods question, I think Karl's concern is the decisive one, and I'd like to add a structural observation and a suggestion. Every current There is also a difficulty with the proposed Appendix E wording. If we quote IUPAC's "magnitude is independent of the size of the system", the motivating case does not satisfy it: for two equal blocks of identical linear series, each of slope x, the regression slope of the concatenation tends to x/4 (exactly 5/21 for I would suggest that the property actually at stake is not intensivity but aggregability. For Whatever keyword is adopted, the estimator would need recording as well, since there are many alternatives for calculating a trend -- OLS, non-parametric (e.g. Theil–Sen), and many more, that give materially different values for the same cell; and it may be worth noting that a trend is only well defined along an ordered coordinate, so Best wishes, |
|
Dear Karl, Lars, Rob et al. In view of your discomfort, I agree that Your comments make me realise that we're using those words in a more general way than the physics and chemisty definitions involving properties of matter. Can we think of better words to use to explain the idea? I think the distinction is well defined: If a quantity is "extensive" with respect to a dimension, it must tend to zero as the size of the cell does (in that dimension), whereas if it is "intensive" a non-zero value is meaningful for a cell of zero extent. Temperature and air pressure are intensive in time, so their default cell method is A tendency, interpreted as a time-derivative, is an intensive value in time, in this sense. A linear trend over time is an estimate of a tendency. A trend is mostly useful as a statistic when it's reasonable to regard the variable as being the sum of an underlying constant rate of change (= the tendency), some random noise and maybe some periodic variation. If this is true, and you know the time-derivative at a sufficient temporal resolution to resolve all the random wiggles, over a long enough time to cover many complete periods of variation, the time-mean of the tendency over the finer cells would be a good estimate of the underlying tendency. By that argument, it would be correct to put One reason for its not being obvious is that in practice we don't have data of such length and high temporal resolution that it can be done that way. Instead we take time-means over intervals to reduce the noise and then calculate the trend. I appreciate Lars's point that a trend is not "aggregatable", meaning you can't get it from the finer cells i.e. the annual means alone. I think the reason is that the linear trend of X is not a statistic of X alone; it's a multivariate statistic that depends on the simultaneous variation of X and t; the annual values of X aren't enough to calculate it because you need to know what years (t) they come from. "Trend" couldn't be a cell method because it isn't a statistic solely of the variation of the data variable within the cells, like all the others are. To compute dX/dt you need t as well as X. I think that's the fundamental reason why dX/dt must have a different standard name from X, rather being regarded as a statistic of X, and it's consistent with a related reason, that its If that is correct, maybe we need some generic cell method keyword to indicate that it depends on the variation of X and the extent of the cell, but it's not extensive and can't be evaluated at a point within the cell. Perhaps the word There's another problem which I had overlooked, regarding "within years". It's natural to describe Rob's use-case as a climatological statistic: first you mean within years e.g. monthly means, then calculate the trend over years. However, the quantity meaned within years is not the I guess we could make some exception for I agree with Lars that it may be useful or necessary to record in more detail how the trend was computed. Since this operation creates a new variable, with a new standard name, I think this extra information should be stored in some other attrbute than the Best wishes Jonathan |

Uh oh!
There was an error while loading. Please reload this page.
Topic for discussion
At EUMETSAT we are developing a climate normals and anomaly service. The data are stored in CF compliant NetCDF files. In our files we provide several time-series statistics that, as far as we can judge, are not part of the statistics supported in CF Conventions version 1.13. It concerns the statistics
From our perspective, it would be useful if these statistics could be included in the cell_methods statistics listed in Appendix E (cell_methods). Is our observation correct, and would it be worthwhile to add these statistics to the next version of the cell_methods table?
Thank you!
202604_CF_Request_Cell_Methods_Statistics.pdf
All reactions