****************************************************************

# Data exploration

In [1]:
import this

The Zen of Python, by Tim Peters

Beautiful is better than ugly.
Explicit is better than implicit.
Simple is better than complex.
Complex is better than complicated.
Flat is better than nested.
Sparse is better than dense.
Readability counts.
Special cases aren't special enough to break the rules.
Although practicality beats purity.
Errors should never pass silently.
Unless explicitly silenced.
In the face of ambiguity, refuse the temptation to guess.
There should be one-- and preferably only one --obvious way to do it.
Although that way may not be obvious at first unless you're Dutch.
Now is better than never.
Although never is often better than *right* now.
If the implementation is hard to explain, it's a bad idea.
If the implementation is easy to explain, it may be a good idea.
Namespaces are one honking great idea -- let's do more of those!


### Quick-note on project directory

The main root dir `~/3dcorrection` is structured as follow:
* `data/` contains raw and preprocessed data. 
    * `raw/` is actually a symbolic link to the same repo for all candidates, DO NOT TOUCH IT!
    * `processed/` will be created when data is preprocessed and will contain all transformed data
* 

In [1]:
import os

root_path = os.path.join('/', 'home', 'jupyter', 'bootcamps')

data_path = os.path.join(root_path, 'data')
cache_path = os.path.join(data_path, 'cache')
raw_data_path = os.path.join(data_path, 'raw')
processed_data_path = os.path.join(data_path, 'processed')

### The 3D Correction Use-Case

The European Centre for Medium-range Weather Forecasts (ECMWF) has developed a series of model giving the current best accurate parametrization scheme available—among those, SPARTACUS delivers **radiation** prediction over the globe. Because it is demanding in computations, a simpler, degraded model called TRIPLECLOUD is developed to satisfy the production environment constraints. 

Like most climate models, to leverage hardware acceleration, the choice is made to split the globe in blocks—this has the immediate consequence of losing the spatial correlation for a gain in parallelization. 

The unit block is a column that express values throughout the vertical dimension over a set of levels. Each level is

Now let's load the raw data we'll be using throughout this hands-on. Take a look at the [source notebook](https://git.ecmwf.int/projects/MLFET/repos/maelstrom-radiation/browse/climetlab_maelstrom_radiation/radiation.py) for a more info on the variables.

## Download dataset

In [2]:
import climetlab as cml
import dask
import dask.array as da
from glob import glob
import numpy as np
import os.path as osp
import xarray as xr

import config

step = 250

cml.settings.set("cache-directory", config.cache_data_path)

cmlds = cml.load_dataset(
    'maelstrom-radiation', 
    dataset='3dcorrection', 
    raw_inputs=False, 
    timestep=list(range(0, 3501, step)), 
    minimal_outputs=False,
    patch=list(range(0, 16, 1)),
    hr_units='K d-1',
)

xr_array = cmlds.to_xarray()
xr_array

By downloading data from this dataset, you agree to the terms and conditions defined at https://apps.ecmwf.int/datasets/licences/general/ If you do not agree with such terms, do not download the data. 


                                                  

Unnamed: 0,Array,Chunk
Bytes,263.96 MiB,397.50 kiB
Shape,"(4070400, 17)","(16960, 6)"
Count,6481 Tasks,1440 Chunks
Type,float32,numpy.ndarray
"Array Chunk Bytes 263.96 MiB 397.50 kiB Shape (4070400, 17) (16960, 6) Count 6481 Tasks 1440 Chunks Type float32 numpy.ndarray",17  4070400,

Unnamed: 0,Array,Chunk
Bytes,263.96 MiB,397.50 kiB
Shape,"(4070400, 17)","(16960, 6)"
Count,6481 Tasks,1440 Chunks
Type,float32,numpy.ndarray

Unnamed: 0,Array,Chunk
Bytes,56.09 GiB,106.36 MiB
Shape,"(4070400, 137, 27)","(16960, 137, 12)"
Count,22800 Tasks,3840 Chunks
Type,float32,numpy.ndarray
"Array Chunk Bytes 56.09 GiB 106.36 MiB Shape (4070400, 137, 27) (16960, 137, 12) Count 22800 Tasks 3840 Chunks Type float32 numpy.ndarray",27  137  4070400,

Unnamed: 0,Array,Chunk
Bytes,56.09 GiB,106.36 MiB
Shape,"(4070400, 137, 27)","(16960, 137, 12)"
Count,22800 Tasks,3840 Chunks
Type,float32,numpy.ndarray

Unnamed: 0,Array,Chunk
Bytes,4.19 GiB,8.93 MiB
Shape,"(4070400, 138, 2)","(16960, 138, 1)"
Count,2880 Tasks,480 Chunks
Type,float32,numpy.ndarray
"Array Chunk Bytes 4.19 GiB 8.93 MiB Shape (4070400, 138, 2) (16960, 138, 1) Count 2880 Tasks 480 Chunks Type float32 numpy.ndarray",2  138  4070400,

Unnamed: 0,Array,Chunk
Bytes,4.19 GiB,8.93 MiB
Shape,"(4070400, 138, 2)","(16960, 138, 1)"
Count,2880 Tasks,480 Chunks
Type,float32,numpy.ndarray

Unnamed: 0,Array,Chunk
Bytes,2.09 GiB,8.93 MiB
Shape,"(4070400, 138, 1)","(16960, 138, 1)"
Count,1200 Tasks,240 Chunks
Type,float32,numpy.ndarray
"Array Chunk Bytes 2.09 GiB 8.93 MiB Shape (4070400, 138, 1) (16960, 138, 1) Count 1200 Tasks 240 Chunks Type float32 numpy.ndarray",1  138  4070400,

Unnamed: 0,Array,Chunk
Bytes,2.09 GiB,8.93 MiB
Shape,"(4070400, 138, 1)","(16960, 138, 1)"
Count,1200 Tasks,240 Chunks
Type,float32,numpy.ndarray

Unnamed: 0,Array,Chunk
Bytes,2.06 GiB,8.80 MiB
Shape,"(4070400, 136, 1)","(16960, 136, 1)"
Count,1200 Tasks,240 Chunks
Type,float32,numpy.ndarray
"Array Chunk Bytes 2.06 GiB 8.80 MiB Shape (4070400, 136, 1) (16960, 136, 1) Count 1200 Tasks 240 Chunks Type float32 numpy.ndarray",1  136  4070400,

Unnamed: 0,Array,Chunk
Bytes,2.06 GiB,8.80 MiB
Shape,"(4070400, 136, 1)","(16960, 136, 1)"
Count,1200 Tasks,240 Chunks
Type,float32,numpy.ndarray

Unnamed: 0,Array,Chunk
Bytes,15.53 MiB,66.25 kiB
Shape,"(4070400,)","(16960,)"
Count,720 Tasks,240 Chunks
Type,float32,numpy.ndarray
"Array Chunk Bytes 15.53 MiB 66.25 kiB Shape (4070400,) (16960,) Count 720 Tasks 240 Chunks Type float32 numpy.ndarray",4070400  1,

Unnamed: 0,Array,Chunk
Bytes,15.53 MiB,66.25 kiB
Shape,"(4070400,)","(16960,)"
Count,720 Tasks,240 Chunks
Type,float32,numpy.ndarray

Unnamed: 0,Array,Chunk
Bytes,15.53 MiB,66.25 kiB
Shape,"(4070400,)","(16960,)"
Count,720 Tasks,240 Chunks
Type,float32,numpy.ndarray
"Array Chunk Bytes 15.53 MiB 66.25 kiB Shape (4070400,) (16960,) Count 720 Tasks 240 Chunks Type float32 numpy.ndarray",4070400  1,

Unnamed: 0,Array,Chunk
Bytes,15.53 MiB,66.25 kiB
Shape,"(4070400,)","(16960,)"
Count,720 Tasks,240 Chunks
Type,float32,numpy.ndarray

Unnamed: 0,Array,Chunk
Bytes,2.09 GiB,8.93 MiB
Shape,"(4070400, 138)","(16960, 138)"
Count,720 Tasks,240 Chunks
Type,float32,numpy.ndarray
"Array Chunk Bytes 2.09 GiB 8.93 MiB Shape (4070400, 138) (16960, 138) Count 720 Tasks 240 Chunks Type float32 numpy.ndarray",138  4070400,

Unnamed: 0,Array,Chunk
Bytes,2.09 GiB,8.93 MiB
Shape,"(4070400, 138)","(16960, 138)"
Count,720 Tasks,240 Chunks
Type,float32,numpy.ndarray

Unnamed: 0,Array,Chunk
Bytes,2.09 GiB,8.93 MiB
Shape,"(4070400, 138)","(16960, 138)"
Count,720 Tasks,240 Chunks
Type,float32,numpy.ndarray
"Array Chunk Bytes 2.09 GiB 8.93 MiB Shape (4070400, 138) (16960, 138) Count 720 Tasks 240 Chunks Type float32 numpy.ndarray",138  4070400,

Unnamed: 0,Array,Chunk
Bytes,2.09 GiB,8.93 MiB
Shape,"(4070400, 138)","(16960, 138)"
Count,720 Tasks,240 Chunks
Type,float32,numpy.ndarray

Unnamed: 0,Array,Chunk
Bytes,2.09 GiB,8.93 MiB
Shape,"(4070400, 138)","(16960, 138)"
Count,720 Tasks,240 Chunks
Type,float32,numpy.ndarray
"Array Chunk Bytes 2.09 GiB 8.93 MiB Shape (4070400, 138) (16960, 138) Count 720 Tasks 240 Chunks Type float32 numpy.ndarray",138  4070400,

Unnamed: 0,Array,Chunk
Bytes,2.09 GiB,8.93 MiB
Shape,"(4070400, 138)","(16960, 138)"
Count,720 Tasks,240 Chunks
Type,float32,numpy.ndarray

Unnamed: 0,Array,Chunk
Bytes,2.09 GiB,8.93 MiB
Shape,"(4070400, 138)","(16960, 138)"
Count,720 Tasks,240 Chunks
Type,float32,numpy.ndarray
"Array Chunk Bytes 2.09 GiB 8.93 MiB Shape (4070400, 138) (16960, 138) Count 720 Tasks 240 Chunks Type float32 numpy.ndarray",138  4070400,

Unnamed: 0,Array,Chunk
Bytes,2.09 GiB,8.93 MiB
Shape,"(4070400, 138)","(16960, 138)"
Count,720 Tasks,240 Chunks
Type,float32,numpy.ndarray

Unnamed: 0,Array,Chunk
Bytes,2.08 GiB,8.86 MiB
Shape,"(4070400, 137)","(16960, 137)"
Count,6240 Tasks,240 Chunks
Type,float32,numpy.ndarray
"Array Chunk Bytes 2.08 GiB 8.86 MiB Shape (4070400, 137) (16960, 137) Count 6240 Tasks 240 Chunks Type float32 numpy.ndarray",137  4070400,

Unnamed: 0,Array,Chunk
Bytes,2.08 GiB,8.86 MiB
Shape,"(4070400, 137)","(16960, 137)"
Count,6240 Tasks,240 Chunks
Type,float32,numpy.ndarray

Unnamed: 0,Array,Chunk
Bytes,2.08 GiB,8.86 MiB
Shape,"(4070400, 137)","(16960, 137)"
Count,6240 Tasks,240 Chunks
Type,float32,numpy.ndarray
"Array Chunk Bytes 2.08 GiB 8.86 MiB Shape (4070400, 137) (16960, 137) Count 6240 Tasks 240 Chunks Type float32 numpy.ndarray",137  4070400,

Unnamed: 0,Array,Chunk
Bytes,2.08 GiB,8.86 MiB
Shape,"(4070400, 137)","(16960, 137)"
Count,6240 Tasks,240 Chunks
Type,float32,numpy.ndarray


In [3]:
print(f"num of instants: {3500 // step} /3500")
print(f"size: {xr_array.nbytes / float(1 << 30):,.0f} GB")

num of instants: 14 /3500
size: 77 GB


The returned object is a ClimateLab dataset Xarray Dataset

Let's check the content of the downloaded file

most operations are computed lazily in dask/xarray when needed and if possible on every chunk, treated and seen 'as if' it was a continuous array

In [4]:
xr_array.sca_inputs

Unnamed: 0,Array,Chunk
Bytes,263.96 MiB,397.50 kiB
Shape,"(4070400, 17)","(16960, 6)"
Count,6481 Tasks,1440 Chunks
Type,float32,numpy.ndarray
"Array Chunk Bytes 263.96 MiB 397.50 kiB Shape (4070400, 17) (16960, 6) Count 6481 Tasks 1440 Chunks Type float32 numpy.ndarray",17  4070400,

Unnamed: 0,Array,Chunk
Bytes,263.96 MiB,397.50 kiB
Shape,"(4070400, 17)","(16960, 6)"
Count,6481 Tasks,1440 Chunks
Type,float32,numpy.ndarray


In [5]:
xr_array.col_inputs

Unnamed: 0,Array,Chunk
Bytes,56.09 GiB,106.36 MiB
Shape,"(4070400, 137, 27)","(16960, 137, 12)"
Count,22800 Tasks,3840 Chunks
Type,float32,numpy.ndarray
"Array Chunk Bytes 56.09 GiB 106.36 MiB Shape (4070400, 137, 27) (16960, 137, 12) Count 22800 Tasks 3840 Chunks Type float32 numpy.ndarray",27  137  4070400,

Unnamed: 0,Array,Chunk
Bytes,56.09 GiB,106.36 MiB
Shape,"(4070400, 137, 27)","(16960, 137, 12)"
Count,22800 Tasks,3840 Chunks
Type,float32,numpy.ndarray


In [6]:
features = [
    'sca_inputs',
    'col_inputs',
    'hl_inputs',
    'inter_inputs',
    'flux_dn_sw',
    'flux_up_sw',
    'flux_dn_lw',
    'flux_up_lw',
]

for feat in features:
    print(f'{feat}: {xr_array[feat].data}')

sca_inputs: dask.array<concatenate, shape=(4070400, 17), dtype=float32, chunksize=(16960, 6), chunktype=numpy.ndarray>
col_inputs: dask.array<concatenate, shape=(4070400, 137, 27), dtype=float32, chunksize=(16960, 137, 12), chunktype=numpy.ndarray>
hl_inputs: dask.array<concatenate, shape=(4070400, 138, 2), dtype=float32, chunksize=(16960, 138, 1), chunktype=numpy.ndarray>
inter_inputs: dask.array<transpose, shape=(4070400, 136, 1), dtype=float32, chunksize=(16960, 136, 1), chunktype=numpy.ndarray>
flux_dn_sw: dask.array<concatenate, shape=(4070400, 138), dtype=float32, chunksize=(16960, 138), chunktype=numpy.ndarray>
flux_up_sw: dask.array<concatenate, shape=(4070400, 138), dtype=float32, chunksize=(16960, 138), chunktype=numpy.ndarray>
flux_dn_lw: dask.array<concatenate, shape=(4070400, 138), dtype=float32, chunksize=(16960, 138), chunktype=numpy.ndarray>
flux_up_lw: dask.array<concatenate, shape=(4070400, 138), dtype=float32, chunksize=(16960, 138), chunktype=numpy.ndarray>


## Flattened data

In [7]:
dataset_len = xr_array.dims['column']
# print(utils.prime_factors(dataset_len))

num_shards = 53 * 2 ** 4
shard_size = dataset_len // num_shards

data = {}
# all this is lazy
for feat in features:
    array = xr_array[feat].data
    array = da.reshape(array, shape=(array.shape[0], -1))
    data.update({feat: array})

In [8]:
x = da.concatenate([
    data['hl_inputs'],
    data['inter_inputs'],
    data['sca_inputs'],
    data['col_inputs']
], axis=-1)

y = da.concatenate([
    data['flux_dn_sw'],
    data['flux_up_sw'],
    data['flux_dn_lw'],
    data['flux_up_lw'],
], axis=-1)

In [9]:
x = da.rechunk(x, chunks=(shard_size, *x.shape[1:]))
y = da.rechunk(y, chunks=(shard_size, *y.shape[1:]))

In [11]:
dask.config.set(scheduler='processes')

out_dir = osp.join(processed_data_path, f'flattened-{step}')
x_path = osp.join(out_dir, 'x')
y_path = osp.join(out_dir, 'y')

for path in [x_path, y_path]:
    os.makedirs(path, exist_ok=True)
    
da.to_npy_stack(x_path, x, axis=0)
da.to_npy_stack(y_path, y, axis=0)

In [15]:
def save_params(params, type_):
    import json
    
    params_path = osp.join(root_path, f'params-{step}.json')
    
    data = {}
    if osp.isfile(params_path):
        with open(params_path, 'r') as file:
            data = json.load(file)
            
    data.update({type_: params})
    
    with open(params_path, 'w') as file:
        json.dump(data, file)
        
save_params({
    "dtype": x.dtype.name,
    "timestep": step,
    "dataset_len": len(x),
    "num_shards": len(x.chunks[0]),
    "x_shape": x.chunksize,
    "y_shape": y.chunksize,
}, 'flattened')

## Feature engineering

In [35]:
dataset_size = xr_array.dims['column']
num_shards = 53 * 2 ** 4
shard_size = dataset_size // num_shards

data = {}
# all this is lazy
for feat in features:
    array = xr_array[feat].data
    array = da.rechunk(array, chunks=(shard_size, *array.shape[1:]))
    # print(type(array))
    data.update({feat: array})
    # print(f'{feat}: {array}')

In [32]:
def broadcast_features(array: da.Array):
    a = da.repeat(array, 138, axis=-1)
    a = da.moveaxis(a, -2, -1)
    return a

def pad_tensor(array: da.Array):
    a = da.pad(array, ((0, 0), (1, 1), (0, 0)))
    return a

In [33]:
from typing import Dict

# still lazy
x = da.concatenate([
    data['hl_inputs'],
    pad_tensor(data['inter_inputs']),
    broadcast_features(data['sca_inputs'][..., np.newaxis])
], axis=-1)

y = da.concatenate([
    data['flux_dn_sw'][..., np.newaxis],
    data['flux_up_sw'][..., np.newaxis],
    data['flux_dn_lw'][..., np.newaxis],
    data['flux_up_lw'][..., np.newaxis],
], axis=-1)

edge = data['col_inputs']

print(f"x of shape: {x.shape}")
print(f"y of shape: {y.shape}")
print(f"edge of shape: {edge.shape}")

x of shape: (4070400, 138, 20)
y of shape: (4070400, 138, 4)
edge of shape: (4070400, 137, 27)


### To a single HDF5 file

In [36]:
da.rechunk(x, chunks=(shard_size, *x.shape[1:]))
da.rechunk(y, chunks=(shard_size, *y.shape[1:]))

Unnamed: 0,Array,Chunk
Bytes,8.37 GiB,10.11 MiB
Shape,"(4070400, 138, 4)","(4800, 138, 4)"
Count,18192 Tasks,848 Chunks
Type,float32,numpy.ndarray
"Array Chunk Bytes 8.37 GiB 10.11 MiB Shape (4070400, 138, 4) (4800, 138, 4) Count 18192 Tasks 848 Chunks Type float32 numpy.ndarray",4  138  4070400,

Unnamed: 0,Array,Chunk
Bytes,8.37 GiB,10.11 MiB
Shape,"(4070400, 138, 4)","(4800, 138, 4)"
Count,18192 Tasks,848 Chunks
Type,float32,numpy.ndarray


In [37]:
from pprint import pprint

out_file = osp.join(processed_data_path, f'features-{step}.h5')
if osp.isfile(out_file):
    os.remove(out_file)
    
x.to_hdf5(out_file, '/x')
y.to_hdf5(out_file, '/y')
edge.to_hdf5(out_file, '/edge')

TypeError: h5py objects cannot be pickled

### To a stack of NumPY files

In [None]:
out_dir = osp.join(processed_data_path, f'features-{step}')

x_path = osp.join(out_dir, 'x')
y_path = osp.join(out_dir, 'y')
edge_path = osp.join(out_dir, 'edge')

for path in [x_path, y_path, edge_path]:
    os.makedirs(path, exist_ok=True)
    
da.to_npy_stack(x_path, x, axis=0)
da.to_npy_stack(y_path, y, axis=0)
da.to_npy_stack(edge_path, edge, axis=0)

### Saving parameters for later use

In [None]:
save_params({
    "dtype": x.dtype.name,
    "timestep": step,
    "dataset_len": len(x),
    "num_shards": len(x.chunks[0]),
    "x_shape": x.chunksize,
    "y_shape": y.chunksize,
    "edge_shape": edge.chunksize,
}, 'features')