****************************************************************

# Data exploration

In [1]:
import this

The Zen of Python, by Tim Peters

Beautiful is better than ugly.
Explicit is better than implicit.
Simple is better than complex.
Complex is better than complicated.
Flat is better than nested.
Sparse is better than dense.
Readability counts.
Special cases aren't special enough to break the rules.
Although practicality beats purity.
Errors should never pass silently.
Unless explicitly silenced.
In the face of ambiguity, refuse the temptation to guess.
There should be one-- and preferably only one --obvious way to do it.
Although that way may not be obvious at first unless you're Dutch.
Now is better than never.
Although never is often better than *right* now.
If the implementation is hard to explain, it's a bad idea.
If the implementation is easy to explain, it may be a good idea.
Namespaces are one honking great idea -- let's do more of those!


### Quick-note on project directory

The main root dir `~/3dcorrection` is structured as follow:
* `data/` contains raw and preprocessed data. 
    * `raw/` is actually a symbolic link to the same repo for all candidates, DO NOT TOUCH IT!
    * `processed/` will be created when data is preprocessed and will contain all transformed data
* 

In [1]:
import os

root_path = os.path.join('/', 'root', 'bootcamps')

data_path = os.path.join(root_path, 'data')
cache_path = os.path.join(data_path, 'cache')
raw_data_path = os.path.join(data_path, 'raw')
processed_data_path = os.path.join(data_path, 'processed')

### The 3D Correction Use-Case

The European Centre for Medium-range Weather Forecasts (ECMWF) has developed a series of model giving the current best accurate parametrization scheme available—among those, SPARTACUS delivers **radiation** prediction over the globe. Because it is demanding in computations, a simpler, degraded model called TRIPLECLOUD is developed to satisfy the production environment constraints. 

Like most climate models, to leverage hardware acceleration, the choice is made to split the globe in blocks—this has the immediate consequence of losing the spatial correlation for a gain in parallelization. 

The unit block is a column that express values throughout the vertical dimension over a set of levels. Each level is

Now let's load the raw data we'll be using throughout this hands-on. Take a look at the [source notebook](https://git.ecmwf.int/projects/MLFET/repos/maelstrom-radiation/browse/climetlab_maelstrom_radiation/radiation.py) for a more info on the variables.

In [21]:
!./run -s 3_build_graphs.py

[35m[1mMetaflow 2.5.0[0m[35m[22m executing [0m[31m[1mBuildGraphsFlow[0m[35m[22m[0m[35m[22m for [0m[31m[1muser:mluser[0m[35m[22m[K[0m[35m[22m[0m
[35m[22mValidating your flow...[K[0m[35m[22m[0m
[32m[1m    The graph looks good![K[0m[32m[1m[0m
[35m[22mRunning pylint...[K[0m[35m[22m[0m
[32m[1m    Pylint is happy![K[0m[32m[1m[0m
[35m2022-02-12 17:30:50.493 [0m[1mWorkflow starting (run-id 1644687050490150):[0m
[35m2022-02-12 17:30:50.508 [0m[32m[1644687050490150/start/1 (pid 17001)] [0m[1mTask is starting.[0m
[35m2022-02-12 17:30:51.234 [0m[32m[1644687050490150/start/1 (pid 17001)] [0m[1mForeach yields 424 child steps.[0m
[35m2022-02-12 17:30:51.234 [0m[32m[1644687050490150/start/1 (pid 17001)] [0m[1mTask finished successfully.[0m
[35m2022-02-12 17:30:51.249 [0m[32m[1644687050490150/slice_and_save/2 (pid 17037)] [0m[1mTask is starting.[0m
[35m2022-02-12 17:30:51.258 [0m[32m[1644687050490150/slice_and_save/3 (p

In [20]:
from metaflow import Flow, namespace

namespace('user:mluser')
flow = Flow('BuildGraphsFlow')
runs = list(flow)
run0 = runs[0]
run0.data.name

print(runs)

[Run('BuildGraphsFlow/1644685159054978'), Run('BuildGraphsFlow/1644685106881761'), Run('BuildGraphsFlow/1644684847330717'), Run('BuildGraphsFlow/1644684805003096'), Run('BuildGraphsFlow/1644684739348654'), Run('BuildGraphsFlow/1644684714030761')]


In [28]:
run = Flow('BuildGraphsFlow').latest_run
steps = list(run.steps())
steps[-1].tasks

<bound method Step.tasks of Step('BuildGraphsFlow/1644687050490150/start')>

In [18]:
import climetlab as cml
import dask
import dask.array as da
from glob import glob
import numpy as np
import os.path as osp
import xarray as xr

import config

cml.settings.set("cache-directory", cache_path)

cmlds = cml.load_dataset(
    'maelstrom-radiation', 
    dataset='3dcorrection', 
    raw_inputs=False, 
    timestep=list(range(0, 3501, config.params['timestep'])), 
    minimal_outputs=False,
    patch=list(range(0, 16, 1)),
    hr_units='K d-1',
)

xr_array = cmlds.to_xarray()
xr_array

                                    

Unnamed: 0,Array,Chunk
Bytes,70.39 MiB,397.50 kiB
Shape,"(1085440, 17)","(16960, 6)"
Count,1729 Tasks,384 Chunks
Type,float32,numpy.ndarray
"Array Chunk Bytes 70.39 MiB 397.50 kiB Shape (1085440, 17) (16960, 6) Count 1729 Tasks 384 Chunks Type float32 numpy.ndarray",17  1085440,

Unnamed: 0,Array,Chunk
Bytes,70.39 MiB,397.50 kiB
Shape,"(1085440, 17)","(16960, 6)"
Count,1729 Tasks,384 Chunks
Type,float32,numpy.ndarray

Unnamed: 0,Array,Chunk
Bytes,14.96 GiB,106.36 MiB
Shape,"(1085440, 137, 27)","(16960, 137, 12)"
Count,6080 Tasks,1024 Chunks
Type,float32,numpy.ndarray
"Array Chunk Bytes 14.96 GiB 106.36 MiB Shape (1085440, 137, 27) (16960, 137, 12) Count 6080 Tasks 1024 Chunks Type float32 numpy.ndarray",27  137  1085440,

Unnamed: 0,Array,Chunk
Bytes,14.96 GiB,106.36 MiB
Shape,"(1085440, 137, 27)","(16960, 137, 12)"
Count,6080 Tasks,1024 Chunks
Type,float32,numpy.ndarray

Unnamed: 0,Array,Chunk
Bytes,1.12 GiB,8.93 MiB
Shape,"(1085440, 138, 2)","(16960, 138, 1)"
Count,768 Tasks,128 Chunks
Type,float32,numpy.ndarray
"Array Chunk Bytes 1.12 GiB 8.93 MiB Shape (1085440, 138, 2) (16960, 138, 1) Count 768 Tasks 128 Chunks Type float32 numpy.ndarray",2  138  1085440,

Unnamed: 0,Array,Chunk
Bytes,1.12 GiB,8.93 MiB
Shape,"(1085440, 138, 2)","(16960, 138, 1)"
Count,768 Tasks,128 Chunks
Type,float32,numpy.ndarray

Unnamed: 0,Array,Chunk
Bytes,571.41 MiB,8.93 MiB
Shape,"(1085440, 138, 1)","(16960, 138, 1)"
Count,320 Tasks,64 Chunks
Type,float32,numpy.ndarray
"Array Chunk Bytes 571.41 MiB 8.93 MiB Shape (1085440, 138, 1) (16960, 138, 1) Count 320 Tasks 64 Chunks Type float32 numpy.ndarray",1  138  1085440,

Unnamed: 0,Array,Chunk
Bytes,571.41 MiB,8.93 MiB
Shape,"(1085440, 138, 1)","(16960, 138, 1)"
Count,320 Tasks,64 Chunks
Type,float32,numpy.ndarray

Unnamed: 0,Array,Chunk
Bytes,563.12 MiB,8.80 MiB
Shape,"(1085440, 136, 1)","(16960, 136, 1)"
Count,320 Tasks,64 Chunks
Type,float32,numpy.ndarray
"Array Chunk Bytes 563.12 MiB 8.80 MiB Shape (1085440, 136, 1) (16960, 136, 1) Count 320 Tasks 64 Chunks Type float32 numpy.ndarray",1  136  1085440,

Unnamed: 0,Array,Chunk
Bytes,563.12 MiB,8.80 MiB
Shape,"(1085440, 136, 1)","(16960, 136, 1)"
Count,320 Tasks,64 Chunks
Type,float32,numpy.ndarray

Unnamed: 0,Array,Chunk
Bytes,4.14 MiB,66.25 kiB
Shape,"(1085440,)","(16960,)"
Count,192 Tasks,64 Chunks
Type,float32,numpy.ndarray
"Array Chunk Bytes 4.14 MiB 66.25 kiB Shape (1085440,) (16960,) Count 192 Tasks 64 Chunks Type float32 numpy.ndarray",1085440  1,

Unnamed: 0,Array,Chunk
Bytes,4.14 MiB,66.25 kiB
Shape,"(1085440,)","(16960,)"
Count,192 Tasks,64 Chunks
Type,float32,numpy.ndarray

Unnamed: 0,Array,Chunk
Bytes,4.14 MiB,66.25 kiB
Shape,"(1085440,)","(16960,)"
Count,192 Tasks,64 Chunks
Type,float32,numpy.ndarray
"Array Chunk Bytes 4.14 MiB 66.25 kiB Shape (1085440,) (16960,) Count 192 Tasks 64 Chunks Type float32 numpy.ndarray",1085440  1,

Unnamed: 0,Array,Chunk
Bytes,4.14 MiB,66.25 kiB
Shape,"(1085440,)","(16960,)"
Count,192 Tasks,64 Chunks
Type,float32,numpy.ndarray

Unnamed: 0,Array,Chunk
Bytes,571.41 MiB,8.93 MiB
Shape,"(1085440, 138)","(16960, 138)"
Count,192 Tasks,64 Chunks
Type,float32,numpy.ndarray
"Array Chunk Bytes 571.41 MiB 8.93 MiB Shape (1085440, 138) (16960, 138) Count 192 Tasks 64 Chunks Type float32 numpy.ndarray",138  1085440,

Unnamed: 0,Array,Chunk
Bytes,571.41 MiB,8.93 MiB
Shape,"(1085440, 138)","(16960, 138)"
Count,192 Tasks,64 Chunks
Type,float32,numpy.ndarray

Unnamed: 0,Array,Chunk
Bytes,571.41 MiB,8.93 MiB
Shape,"(1085440, 138)","(16960, 138)"
Count,192 Tasks,64 Chunks
Type,float32,numpy.ndarray
"Array Chunk Bytes 571.41 MiB 8.93 MiB Shape (1085440, 138) (16960, 138) Count 192 Tasks 64 Chunks Type float32 numpy.ndarray",138  1085440,

Unnamed: 0,Array,Chunk
Bytes,571.41 MiB,8.93 MiB
Shape,"(1085440, 138)","(16960, 138)"
Count,192 Tasks,64 Chunks
Type,float32,numpy.ndarray

Unnamed: 0,Array,Chunk
Bytes,571.41 MiB,8.93 MiB
Shape,"(1085440, 138)","(16960, 138)"
Count,192 Tasks,64 Chunks
Type,float32,numpy.ndarray
"Array Chunk Bytes 571.41 MiB 8.93 MiB Shape (1085440, 138) (16960, 138) Count 192 Tasks 64 Chunks Type float32 numpy.ndarray",138  1085440,

Unnamed: 0,Array,Chunk
Bytes,571.41 MiB,8.93 MiB
Shape,"(1085440, 138)","(16960, 138)"
Count,192 Tasks,64 Chunks
Type,float32,numpy.ndarray

Unnamed: 0,Array,Chunk
Bytes,571.41 MiB,8.93 MiB
Shape,"(1085440, 138)","(16960, 138)"
Count,192 Tasks,64 Chunks
Type,float32,numpy.ndarray
"Array Chunk Bytes 571.41 MiB 8.93 MiB Shape (1085440, 138) (16960, 138) Count 192 Tasks 64 Chunks Type float32 numpy.ndarray",138  1085440,

Unnamed: 0,Array,Chunk
Bytes,571.41 MiB,8.93 MiB
Shape,"(1085440, 138)","(16960, 138)"
Count,192 Tasks,64 Chunks
Type,float32,numpy.ndarray

Unnamed: 0,Array,Chunk
Bytes,567.27 MiB,8.86 MiB
Shape,"(1085440, 137)","(16960, 137)"
Count,1664 Tasks,64 Chunks
Type,float32,numpy.ndarray
"Array Chunk Bytes 567.27 MiB 8.86 MiB Shape (1085440, 137) (16960, 137) Count 1664 Tasks 64 Chunks Type float32 numpy.ndarray",137  1085440,

Unnamed: 0,Array,Chunk
Bytes,567.27 MiB,8.86 MiB
Shape,"(1085440, 137)","(16960, 137)"
Count,1664 Tasks,64 Chunks
Type,float32,numpy.ndarray

Unnamed: 0,Array,Chunk
Bytes,567.27 MiB,8.86 MiB
Shape,"(1085440, 137)","(16960, 137)"
Count,1664 Tasks,64 Chunks
Type,float32,numpy.ndarray
"Array Chunk Bytes 567.27 MiB 8.86 MiB Shape (1085440, 137) (16960, 137) Count 1664 Tasks 64 Chunks Type float32 numpy.ndarray",137  1085440,

Unnamed: 0,Array,Chunk
Bytes,567.27 MiB,8.86 MiB
Shape,"(1085440, 137)","(16960, 137)"
Count,1664 Tasks,64 Chunks
Type,float32,numpy.ndarray


The returned object is a ClimateLab dataset Xarray Dataset

Let's check the content of the downloaded file

most operations are computed lazily in dask/xarray when needed and if possible on every chunk, treated and seen 'as if' it was a continuous array

In [3]:
xr_array.sca_inputs

Unnamed: 0,Array,Chunk
Bytes,70.39 MiB,397.50 kiB
Shape,"(1085440, 17)","(16960, 6)"
Count,1729 Tasks,384 Chunks
Type,float32,numpy.ndarray
"Array Chunk Bytes 70.39 MiB 397.50 kiB Shape (1085440, 17) (16960, 6) Count 1729 Tasks 384 Chunks Type float32 numpy.ndarray",17  1085440,

Unnamed: 0,Array,Chunk
Bytes,70.39 MiB,397.50 kiB
Shape,"(1085440, 17)","(16960, 6)"
Count,1729 Tasks,384 Chunks
Type,float32,numpy.ndarray


In [4]:
features = [
    'sca_inputs',
    'col_inputs',
    'hl_inputs',
    'inter_inputs',
    'flux_dn_sw',
    'flux_up_sw',
    'flux_dn_lw',
    'flux_up_lw',
]

for feat in features:
    print(f'{feat}: {xr_array[feat].data}')

sca_inputs: dask.array<concatenate, shape=(1085440, 17), dtype=float32, chunksize=(16960, 6), chunktype=numpy.ndarray>
col_inputs: dask.array<concatenate, shape=(1085440, 137, 27), dtype=float32, chunksize=(16960, 137, 12), chunktype=numpy.ndarray>
hl_inputs: dask.array<concatenate, shape=(1085440, 138, 2), dtype=float32, chunksize=(16960, 138, 1), chunktype=numpy.ndarray>
inter_inputs: dask.array<transpose, shape=(1085440, 136, 1), dtype=float32, chunksize=(16960, 136, 1), chunktype=numpy.ndarray>
flux_dn_sw: dask.array<concatenate, shape=(1085440, 138), dtype=float32, chunksize=(16960, 138), chunktype=numpy.ndarray>
flux_up_sw: dask.array<concatenate, shape=(1085440, 138), dtype=float32, chunksize=(16960, 138), chunktype=numpy.ndarray>
flux_dn_lw: dask.array<concatenate, shape=(1085440, 138), dtype=float32, chunksize=(16960, 138), chunktype=numpy.ndarray>
flux_up_lw: dask.array<concatenate, shape=(1085440, 138), dtype=float32, chunksize=(16960, 138), chunktype=numpy.ndarray>


In [6]:
dataset_size = xr_array.dims['column']
num_shards = 53 * 2 ** 3
shard_size = dataset_size // num_shards

data = {}
# all this is lazy
for feat in features:
    array = xr_array[feat].data
    array = da.rechunk(array, chunks=(shard_size, *array.shape[1:]))
    data.update({feat: array})
    print(f'{feat}: {array}')

sca_inputs: dask.array<rechunk-merge, shape=(1085440, 17), dtype=float32, chunksize=(2560, 17), chunktype=numpy.ndarray>
col_inputs: dask.array<rechunk-merge, shape=(1085440, 137, 27), dtype=float32, chunksize=(2560, 137, 27), chunktype=numpy.ndarray>
hl_inputs: dask.array<rechunk-merge, shape=(1085440, 138, 2), dtype=float32, chunksize=(2560, 138, 2), chunktype=numpy.ndarray>
inter_inputs: dask.array<rechunk-merge, shape=(1085440, 136, 1), dtype=float32, chunksize=(2560, 136, 1), chunktype=numpy.ndarray>
flux_dn_sw: dask.array<rechunk-merge, shape=(1085440, 138), dtype=float32, chunksize=(2560, 138), chunktype=numpy.ndarray>
flux_up_sw: dask.array<rechunk-merge, shape=(1085440, 138), dtype=float32, chunksize=(2560, 138), chunktype=numpy.ndarray>
flux_dn_lw: dask.array<rechunk-merge, shape=(1085440, 138), dtype=float32, chunksize=(2560, 138), chunktype=numpy.ndarray>
flux_up_lw: dask.array<rechunk-merge, shape=(1085440, 138), dtype=float32, chunksize=(2560, 138), chunktype=numpy.ndarra

In [7]:
def broadcast_features(array: da.Array):
    a = da.repeat(array, 138, axis=-1)
    a = da.moveaxis(a, -2, -1)
    return a

def pad_tensor(array: da.Array):
    a = da.pad(array, ((0, 0), (1, 1), (0, 0)))
    return a

In [15]:
from typing import Dict

# still lazy
@dask.delayed
def perform_feature_engineering(data: Dict[str, da.Array]):
    
    print("feature engineering x")
    x = da.concatenate([
        data['hl_inputs'],
        pad_tensor(data['inter_inputs']),
        broadcast_features(data['sca_inputs'][..., np.newaxis])
    ], axis=-1)
    
    print("feature engineering y")
    y = da.concatenate([
        data['flux_dn_sw'][..., np.newaxis],
        data['flux_up_sw'][..., np.newaxis],
        data['flux_dn_lw'][..., np.newaxis],
        data['flux_up_lw'][..., np.newaxis],
    ], axis=-1)

    print(f"x of shape: {x.shape}")
    print(f"y of shape: {y.shape}")

In [17]:
from pprint import pprint

x, y = perform_feature_engineering(data).compute()
out_file = osp.join(processed_data_path, 'feats.h5')
x.to_hdf5(out_file, '/x')
y.to_hdf5(out_file, '/y')

feature engineering x


AttributeError: 'numpy.ndarray' object has no attribute 'chunks'

Traceback
---------
  File "/usr/local/lib/python3.8/dist-packages/dask/local.py", line 220, in execute_task
    result = _execute_task(task, data)
  File "/usr/local/lib/python3.8/dist-packages/dask/core.py", line 119, in _execute_task
    return func(*(_execute_task(a, cache) for a in args))
  File "/tmp/ipykernel_235239/3926166962.py", line 10, in perform_feature_engineering
  File "/tmp/ipykernel_235239/3661535310.py", line 2, in broadcast_features
  File "/usr/local/lib/python3.8/dist-packages/dask/array/creation.py", line 743, in repeat
    cchunks = cached_cumsum(a.chunks[axis], initial_zero=True)


In [None]:
x.shape

In [9]:
dask.config.set(scheduler='processes')

out_dir = osp.join(processed_data_path, 'feats_npy')
da.to_npy_stack(osp.join(out_dir, 'x'), x, axis=0)
da.to_npy_stack(osp.join(out_dir, 'y'), y, axis=0)