Describe the bug
A systematic sweep of every dataclass field in deeptab/configs/ against every shipped architecture
turned up a further set of options that either crash or are silently not wired up. This complements
#423, which lists an earlier set; none of the items below are in it.
Every item was reproduced by execution.
AutoInt crashes when use_cls=True — the output head is sized without the CLS token
Where: deeptab/architectures/autoint.py (69, 116)
AutoInt computes n_inputs = sum(len(info) for info in feature_information) and builds self.head = nn.Linear(d_model * n_inputs, num_classes), but EmbeddingLayer honours config.use_cls and prepends a CLS token, so the flattened sequence is d_model * (n_inputs + 1) wide and the final matmul fails.
Observed: use_cls=False -> ok, head=Linear(in_features=160, out_features=1). use_cls=True -> RuntimeError: mat1 and mat2 shapes cannot be multiplied (13x192 and 160x1). 192 = 326 (5 features + CLS) vs the head's 160 = 325. AutoIntConfig documents use_cls as 'Whether to use a CLS token for pooling instead of averaging'.
Expected: n_inputs should be incremented by 1 when config.use_cls is set (SAINT does exactly this at deeptab/architectures/saint.py:66-67), so AutoInt(use_cls=True) trains normally.
Repro
import warnings; warnings.simplefilter('ignore')
import numpy as np, pandas as pd
from deeptab.configs import TrainerConfig
from deeptab.configs.models.autoint_config import AutoIntConfig
from deeptab.models import AutoIntRegressor
rng = np.random.default_rng(0)
X = pd.DataFrame({'n1': rng.normal(size=64), 'n2': rng.normal(size=64), 'n3': rng.uniform(size=64),
'c1': rng.choice(list('abc'), size=64), 'c2': rng.choice(list('xy'), size=64)})
y = (X['n1'] * 2).to_numpy()
for uc in (False, True):
est = AutoIntRegressor(model_config=AutoIntConfig(d_model=32, n_layers=1, n_heads=4, use_cls=uc),
trainer_config=TrainerConfig(max_epochs=1, batch_size=16, patience=1, lr_patience=1),
random_state=101)
try:
est.fit(X, y, max_epochs=1, batch_size=16, patience=1, accelerator='cpu', devices=1,
enable_progress_bar=False, enable_model_summary=False, logger=False)
print(uc, 'ok', est._task_model.estimator.head)
except Exception as e:
print(uc, 'RAISED', type(e).__name__, e)
Mambular's dilation is applied only to the backward conv, so it is a no-op by default and crashes when bidirectional=True
Where: deeptab/nn/blocks/mamba.py (338-356)
conv1d_fwd is constructed without dilation= while conv1d_bwd (built only when bidirectional=True) passes dilation=dilation; both keep padding=d_conv-1. Consequently MambularConfig.dilation is silently ignored under the default bidirectional=False, and enabling both dilation>1 and bidirectional=True crashes because the padding was never widened for the dilated kernel.
Observed: bidirectional=False dilation=1 -> conv1d_fwd.dilation=(1,), pred0=0.0357717164. bidirectional=False dilation=4 -> conv1d_fwd.dilation=(1,), pred0=0.0357717164 (bit-identical: dilation ignored). bidirectional=True dilation=1 -> ok. bidirectional=True dilation=4 -> RuntimeError: Calculated padded input size per channel: (11). Kernel size: (13). Kernel size can't be greater than actual input size.
Expected: dilation should be applied to conv1d_fwd as well, with padding computed as dilation*(d_conv-1) so both directions produce a valid, symmetric convolution. As shipped, MambularConfig.dilation (documented as 'Dilation factor for the convolution') either does nothing or crashes.
Repro
import warnings; warnings.simplefilter('ignore')
import numpy as np, pandas as pd
from deeptab.configs import TrainerConfig
from deeptab.configs.models.mambular_config import MambularConfig
from deeptab.models import MambularRegressor
rng = np.random.default_rng(0)
X = pd.DataFrame({'n1': rng.normal(size=64), 'n2': rng.normal(size=64), 'n3': rng.uniform(size=64),
'c1': rng.choice(list('abc'), size=64), 'c2': rng.choice(list('xy'), size=64)})
y = (X['n1'] * 2 + X['n2']).to_numpy()
for bidir in (False, True):
for dil in (1, 4):
cfg = MambularConfig(d_model=32, n_layers=1, d_state=8, dilation=dil, bidirectional=bidir)
est = MambularRegressor(model_config=cfg,
trainer_config=TrainerConfig(max_epochs=1, batch_size=16, patience=1, lr_patience=1),
random_state=101)
try:
est.fit(X, y, max_epochs=1, batch_size=16, patience=1, accelerator='cpu', devices=1,
enable_progress_bar=False, enable_model_summary=False, logger=False)
blk = est._task_model.estimator.mamba.layers[0].layers
print(bidir, dil, 'fwd.dilation=', blk.conv1d_fwd.dilation,
'pred0=', est.predict(X)[0])
except Exception as e:
print(bidir, dil, 'RAISED', type(e).__name__, e)
TabulaRNN's rnn_dropout (default 0.2) has no effect — every recurrent layer is built with num_layers=1
Where: deeptab/nn/blocks/common.py (1861-1874 (ConvRNN.init rnn_args))
ConvRNN stacks config.n_layers separate single-layer RNN/LSTM/GRU modules and passes dropout=self.rnn_dropout into each, but PyTorch discards the dropout argument when num_layers == 1, so TabulaRNNConfig.rnn_dropout is a documented, on-by-default option that regularises nothing.
Observed: rnn_dropout=0.0 -> [(1, 0.0), (1, 0.0), (1, 0.0)], pred0=0.5273975134. rnn_dropout=0.9 -> [(1, 0.9), (1, 0.9), (1, 0.0)], pred0=0.5273975134 — bit-identical predictions with dropout at 0.9. PyTorch also emits the standard 'dropout option adds dropout after all but last recurrent layer, so non-zero dropout expects num_layers greater than 1' warning, which is swallowed.
Expected: Either an explicit nn.Dropout(rnn_dropout) between the stacked single-layer RNNs, or a single rnn_layer(num_layers=n_layers, dropout=rnn_dropout). As shipped, every TabulaRNN user silently gets zero recurrent dropout despite the default of 0.2.
Repro
import warnings; warnings.simplefilter('ignore')
import numpy as np, pandas as pd
from deeptab.configs import TrainerConfig
from deeptab.configs.models.tabularnn_config import TabulaRNNConfig
from deeptab.models import TabulaRNNRegressor
rng = np.random.default_rng(0)
X = pd.DataFrame({'n1': rng.normal(size=64), 'n2': rng.normal(size=64), 'n3': rng.uniform(size=64),
'c1': rng.choice(list('abc'), size=64), 'c2': rng.choice(list('xy'), size=64)})
y = (X['n1'] * 2 + X['n2']).to_numpy()
for d in (0.0, 0.9):
est = TabulaRNNRegressor(
model_config=TabulaRNNConfig(d_model=32, n_layers=3, dim_feedforward=32, rnn_dropout=d),
trainer_config=TrainerConfig(max_epochs=1, batch_size=16, patience=1, lr_patience=1),
random_state=101)
est.fit(X, y, max_epochs=1, batch_size=16, patience=1, accelerator='cpu', devices=1,
enable_progress_bar=False, enable_model_summary=False, logger=False)
rnns = est._task_model.estimator.rnn.rnns
print('rnn_dropout=', d, [(r.num_layers, r.dropout) for r in rnns],
'pred0=', est.predict(X)[0])
ENODE hardcodes its prediction head, silently ignoring head_layer_sizes / head_skip_layers / head_activation / head_use_batch_norm
Where: deeptab/architectures/enode.py (7 (unused MLPhead import), 85-90)
ENODE imports MLPhead but builds tabular_head as a fixed nn.Sequential(Linear(d_model,d_model), ReLU, Dropout(head_dropout), Linear(d_model,num_classes)), so four of the five documented head_* fields on ENODEConfig have no effect at all.
Observed: For head_layer_sizes=[] and [16, 8] the head is the identical Sequential(Linear(8,8), ReLU(), Dropout(0.3), Linear(8,1)) and pred0=0.0899604410 in both cases. head_skip_layers=True and head_use_batch_norm=True likewise produce bit-identical predictions.
Expected: ENODE should build its head with MLPhead(input_dim=..., config=config, output_dim=num_classes) like NODE does (deeptab/architectures/node.py:86-90), so the documented head_* configuration takes effect.
Repro
import warnings; warnings.simplefilter('ignore')
import numpy as np, pandas as pd
from deeptab.configs import TrainerConfig
from deeptab.configs.models.enode_config import ENODEConfig
from deeptab.models import ENODERegressor
rng = np.random.default_rng(0)
X = pd.DataFrame({'n1': rng.normal(size=64), 'n2': rng.normal(size=64), 'n3': rng.uniform(size=64),
'c1': rng.choice(list('abc'), size=64), 'c2': rng.choice(list('xy'), size=64)})
y = (X['n1'] * 2 + X['n2']).to_numpy()
for hls in ([], [16, 8]):
est = ENODERegressor(
model_config=ENODEConfig(num_layers=1, layer_dim=8, depth=2, head_layer_sizes=hls,
head_skip_layers=True, head_use_batch_norm=True),
trainer_config=TrainerConfig(max_epochs=1, batch_size=16, patience=1, lr_patience=1),
random_state=101)
est.fit(X, y, max_epochs=1, batch_size=16, patience=1, accelerator='cpu', devices=1,
enable_progress_bar=False, enable_model_summary=False, logger=False)
print(hls, est._task_model.estimator.tabular_head, 'pred0=', est.predict(X)[0])
Three of the eight documented pooling_method values crash at model construction on every pooling-based architecture
Where: deeptab/core/base_model.py (148-176 (initialize_pooling_layers))
BaseModel.initialize_pooling_layers builds the learnable pooling submodules from config.dim_feedforward, an attribute only TabulaRNNConfig defines, so pooling_method in {learned_flatten, attention, gated} raises AttributeError on FTTransformer, TabTransformer, SAINT, Mambular and MambAttention; on TabulaRNN, learned_flatten instead dies because the caller passes n_inputs as a list.
Observed: avg/max/sum/last/cls succeed; learned_flatten, attention and gated each raise AttributeError: 'FTTransformerConfig' object has no attribute 'dim_feedforward'. Same for TabTransformerConfig, SAINTConfig, MambularConfig and MambAttentionConfig. TabulaRNN raises TypeError: empty(): argument 'size' failed to unpack the object at pos 2 with error "type must be tuple of ints, but got list" for learned_flatten, because deeptab/architectures/tabularnn.py:86 passes n_inputs=[len(info) for info in feature_information] (a list) into nn.Linear(n_inputs * config.dim_feedforward, ...).
Expected: Either the learnable pooling modes work (sized from config.d_model, which is the actual width of the (B, S, d_model) tensor pool_sequence receives), or pooling_method rejects them up front with a clear error. Silently shipping three implemented-but-unreachable pooling modes that only fail after preprocessing and model construction is a defect.
Repro
import warnings; warnings.simplefilter('ignore')
import numpy as np, pandas as pd
from deeptab.configs import TrainerConfig
from deeptab.configs.models.fttransformer_config import FTTransformerConfig
from deeptab.models import FTTransformerRegressor
rng = np.random.default_rng(0)
X = pd.DataFrame({'n1': rng.normal(size=64), 'n2': rng.normal(size=64),
'c1': rng.choice(list('abc'), size=64)})
y = (X['n1'] * 2).to_numpy()
for pm in ['avg', 'max', 'sum', 'last', 'cls', 'learned_flatten', 'attention', 'gated']:
cfg = FTTransformerConfig(d_model=32, n_layers=1, n_heads=4,
transformer_dim_feedforward=64, pooling_method=pm)
est = FTTransformerRegressor(model_config=cfg,
trainer_config=TrainerConfig(max_epochs=1, batch_size=16, patience=1, lr_patience=1),
random_state=101)
try:
est.fit(X, y, max_epochs=1, batch_size=16, patience=1, accelerator='cpu', devices=1,
enable_progress_bar=False, enable_model_summary=False, logger=False)
print(pm, 'ok')
except Exception as e:
print(pm, 'RAISED', type(e).__name__, e)
use_pscan=True crashes with TypeError: 'NoneType' object is not callable — the optional-dependency fallback is never taken
Where: deeptab/nn/blocks/mamba.py (320-329, 517-518)
When the optional mambapy package is absent, the constructor catches ImportError, prints an install hint and sets self.pscan = None, but selective_scan_seq still unconditionally calls self.pscan(...) whenever self.use_pscan is true, so Mambular/MambAttention with use_pscan=True die deep inside forward instead of falling back to the sequential scan that is sitting right there in the else branch.
Observed: TypeError: 'NoneType' object is not callable, raised at deeptab/nn/blocks/mamba.py:518 (hs = self.pscan(deltaA, BX)) inside selective_scan_seq. Same for MambAttention. (import mambapy -> ModuleNotFoundError on this machine; mambapy is not a declared dependency in pyproject.toml.)
Expected: The except-ImportError branch clearly intends graceful degradation, so it should also set self.use_pscan = False (falling back to the sequential scan), or raise a clear ImportError naming mambapy at construction time. Not a cryptic NoneType TypeError from inside the scan.
Repro
import warnings; warnings.simplefilter('ignore')
import numpy as np, pandas as pd
from deeptab.configs import TrainerConfig
from deeptab.configs.models.mambular_config import MambularConfig
from deeptab.models import MambularRegressor
rng = np.random.default_rng(0)
X = pd.DataFrame({'n1': rng.normal(size=64), 'n2': rng.normal(size=64),
'c1': rng.choice(list('abc'), size=64)})
y = (X['n1'] * 2).to_numpy()
est = MambularRegressor(model_config=MambularConfig(d_model=16, n_layers=1, d_state=8, use_pscan=True),
trainer_config=TrainerConfig(max_epochs=1, batch_size=16, patience=1, lr_patience=1),
random_state=101)
est.fit(X, y, max_epochs=1, batch_size=16, patience=1, accelerator='cpu', devices=1,
enable_progress_bar=False, enable_model_summary=False, logger=False)
AutoInt's fprenorm option is dead — the code reads a nonexistent config.prenorm
Where: deeptab/architectures/autoint.py (114)
AutoInt gates its final normalisation on getattr(config, 'prenorm', False), but the field on AutoIntConfig is named fprenorm, so last_norm is always None and setting fprenorm=True changes nothing.
Observed: fprenorm=False -> last_norm=None, pred0=-0.3628833294. fprenorm=True -> last_norm=None, pred0=-0.3628833294 (bit-identical). grep -n prenorm deeptab/configs/models/autoint_config.py shows only fprenorm; no config in the repo defines prenorm.
Expected: fprenorm=True should construct the nn.LayerNorm(d_model) and change the model's output. This is a distinct field from the kv_compression/activation items already filed under #423.
Repro
import warnings; warnings.simplefilter('ignore')
import numpy as np, pandas as pd
from deeptab.configs import TrainerConfig
from deeptab.configs.models.autoint_config import AutoIntConfig
from deeptab.models import AutoIntRegressor
rng = np.random.default_rng(0)
X = pd.DataFrame({'n1': rng.normal(size=64), 'n2': rng.normal(size=64), 'n3': rng.uniform(size=64),
'c1': rng.choice(list('abc'), size=64), 'c2': rng.choice(list('xy'), size=64)})
y = (X['n1'] * 2 + X['n2']).to_numpy()
for v in (False, True):
est = AutoIntRegressor(model_config=AutoIntConfig(d_model=32, n_layers=1, n_heads=4, fprenorm=v),
trainer_config=TrainerConfig(max_epochs=1, batch_size=16, patience=1, lr_patience=1),
random_state=101)
est.fit(X, y, max_epochs=1, batch_size=16, patience=1, accelerator='cpu', devices=1,
enable_progress_bar=False, enable_model_summary=False, logger=False)
print('fprenorm=', v, 'last_norm=', est._task_model.estimator.last_norm,
'pred0=', est.predict(X)[0])
NODEConfig.norm and ENODEConfig.norm are documented options that no code ever reads
Where: deeptab/architectures/node.py (whole file (no reference to norm); same for deeptab/architectures/enode.py)
Both NODEConfig and ENODEConfig declare norm: str | None = None with the docstring 'Type of normalization to use in the model', but neither architecture calls get_normalization_layer or otherwise reads the field, so setting it changes nothing and the model contains no normalisation module.
Observed: NODE: norm=None/LayerNorm/BatchNorm all give pred0=-0.1135944128 and norm_modules=[]. ENODE: all three give pred0=0.0899604410 and norm_modules=[]. grep -n norm deeptab/architectures/node.py returns nothing.
Expected: Either the field is wired to get_normalization_layer (as FTTransformer/SAINT/Mambular do) or it is removed from the config. Currently a user tuning norm on NODE/ENODE (including via HPO) is tuning a parameter that cannot change the model. This is a different field/model from the TabM norm item in #423.
Repro
import warnings; warnings.simplefilter('ignore')
import numpy as np, pandas as pd
from deeptab.configs import TrainerConfig
from deeptab.configs.models.node_config import NODEConfig
from deeptab.configs.models.enode_config import ENODEConfig
from deeptab.models import NODERegressor, ENODERegressor
rng = np.random.default_rng(0)
X = pd.DataFrame({'n1': rng.normal(size=64), 'n2': rng.normal(size=64), 'n3': rng.uniform(size=64),
'c1': rng.choice(list('abc'), size=64), 'c2': rng.choice(list('xy'), size=64)})
y = (X['n1'] * 2 + X['n2']).to_numpy()
for cls, cfgcls in ((NODERegressor, NODEConfig), (ENODERegressor, ENODEConfig)):
for nrm in (None, 'LayerNorm', 'BatchNorm'):
est = cls(model_config=cfgcls(num_layers=1, layer_dim=8, depth=2, norm=nrm),
trainer_config=TrainerConfig(max_epochs=1, batch_size=16, patience=1, lr_patience=1),
random_state=101)
est.fit(X, y, max_epochs=1, batch_size=16, patience=1, accelerator='cpu', devices=1,
enable_progress_bar=False, enable_model_summary=False, logger=False)
norm_mods = [n for n, _ in est._task_model.estimator.named_modules() if 'norm' in n.lower()]
print(cls.__name__, nrm, est.predict(X)[0], norm_mods)
BaseModelConfig.cat_encoding is validated and documented but never read by any model, embedding or preprocessing code
Where: deeptab/configs/core.py (84-86 (docstring), 106 (field), 111-118 (validation))
cat_encoding is a first-class BaseModelConfig field with strict validation against {'int','one-hot','linear'} and is described in the docs as something 'the preprocessing and embedding machinery rely on', but a repo-wide grep shows it is never consumed anywhere: the only other occurrences are dead entries in HPO fixed_params default dicts and the HPO search space.
Observed: All three encodings produce bit-identical predictions, on every one of the 14 architectures I swept (with and without use_embeddings=True) — 28 identical pairs. The grep confirms there is no read site.
Expected: cat_encoding should either drive the categorical encoding in EmbeddingLayer / the preprocessor, or be removed. As shipped it is a validated, documented dial that is wired to nothing, and docs/core_concepts/config_system.md:132 and docs/core_concepts/custom_models.md:37 both advertise it as live.
Repro
# 1) static: no consumer exists
# grep -rn cat_encoding --include='*.py' /Users/christophweisser/Desktop/Coding/DeepTab/deeptab
# -> only configs/core.py (declaration+validation), models/{regressor,classifier,lss}_base.py
# and _mixins/hpo.py fixed_params dicts, and hpo/search_space.py.
# 2) dynamic:
import warnings; warnings.simplefilter('ignore')
import numpy as np, pandas as pd
from deeptab.configs import TrainerConfig
from deeptab.configs.models.mlp_config import MLPConfig
from deeptab.models import MLPRegressor
rng = np.random.default_rng(0)
X = pd.DataFrame({'n1': rng.normal(size=64), 'n2': rng.normal(size=64),
'c1': rng.choice(list('abc'), size=64), 'c2': rng.choice(list('xy'), size=64)})
y = (X['n1'] * 2).to_numpy()
for enc in ('int', 'one-hot', 'linear'):
est = MLPRegressor(model_config=MLPConfig(layer_sizes=[32, 16], use_embeddings=True, cat_encoding=enc),
trainer_config=TrainerConfig(max_epochs=1, batch_size=16, patience=1, lr_patience=1),
random_state=101)
est.fit(X, y, max_epochs=1, batch_size=16, patience=1, accelerator='cpu', devices=1,
enable_progress_bar=False, enable_model_summary=False, logger=False)
print(enc, est.predict(X)[:3])
BaseModelConfig.batch_norm, layer_norm and activation are silently ignored by most architectures
Where: deeptab/configs/core.py (75-81 (docstrings), 99-102 (fields), 147-154 (cross-field warning))
batch_norm and layer_norm are documented as 'Whether to use batch/layer normalisation in the model body' and post_init even warns when both are enabled, implying they are live everywhere; in fact only MLP reads both and only MLP/TabM read batch_norm. Likewise activation ('Activation function used throughout the model body', re-documented per-model as e.g. 'Activation function for the transformer layers') is ignored by FTTransformer, TabTransformer, TabulaRNN, NODE, ENODE and NDTF. layer_norm_eps is ignored by every architecture that does not go through get_normalization_layer.
Observed: All five FTTransformer variants return bit-identical predictions. Across the full sweep: batch_norm=True and layer_norm=True were bit-identical no-ops for FTTransformer, TabTransformer, AutoInt, SAINT, NODE, ENODE, NDTF, TabulaRNN, Mambular, MambAttention, MambaTab, ResNet (12 of 14; MLP honours both, TabM crashes on batch_norm per #423). layer_norm_eps was a no-op for MLP, ResNet, NODE, ENODE, NDTF, AutoInt, TabM, MambaTab. activation was a no-op for FTTransformer, TabTransformer, TabulaRNN, NODE, ENODE, NDTF (AutoInt is already covered by #423).
Expected: Fields that a given architecture cannot honour should not be inherited/advertised for it — or the architectures should read them. Today a user (or the HPO search space) can set batch_norm/layer_norm/activation on almost any model and get exactly the same network back, with post_init even emitting a 'both batch_norm and layer_norm are set' warning for a pair of dials that do nothing.
Repro
import warnings; warnings.simplefilter('ignore')
import numpy as np, pandas as pd, torch.nn as nn
from deeptab.configs import TrainerConfig
from deeptab.configs.models.fttransformer_config import FTTransformerConfig
from deeptab.models import FTTransformerRegressor
rng = np.random.default_rng(0)
X = pd.DataFrame({'n1': rng.normal(size=64), 'n2': rng.normal(size=64), 'n3': rng.uniform(size=64),
'c1': rng.choice(list('abc'), size=64), 'c2': rng.choice(list('xy'), size=64)})
y = (X['n1'] * 2 + X['n2']).to_numpy()
base = dict(d_model=32, n_layers=2, n_heads=4, transformer_dim_feedforward=64)
for label, kw in [('default', {}), ('batch_norm=True', dict(batch_norm=True)),
('layer_norm=True', dict(layer_norm=True)),
('activation=ReLU', dict(activation=nn.ReLU())),
('activation=GELU', dict(activation=nn.GELU()))]:
est = FTTransformerRegressor(model_config=FTTransformerConfig(**base, **kw),
trainer_config=TrainerConfig(max_epochs=1, batch_size=16, patience=1, lr_patience=1),
random_state=101)
est.fit(X, y, max_epochs=1, batch_size=16, patience=1, accelerator='cpu', devices=1,
enable_progress_bar=False, enable_model_summary=False, logger=False)
print(f'{label:18s} pred0={est.predict(X)[0]:.10f}')
# static confirmation:
# grep -l 'hparams.batch_norm' deeptab/architectures/*.py -> mlp.py, tabm.py only
# grep -l 'hparams.layer_norm' deeptab/architectures/*.py -> mlp.py only
# grep -l 'hparams.activation' deeptab/architectures/*.py -> mlp.py, resnet.py, tabm.py, tabr.py only
AD_weight_decay=False is a silent no-op: the _no_weight_decay markers are never read by the optimizer builder
Where: deeptab/nn/blocks/mamba.py (407-414 (setter); deeptab/training/optimizers.py:386-406 (consumer))
MambaBlock tags A_log_fwd/D_fwd (and the bwd pair) with param._no_weight_decay = True when AD_weight_decay is False, but build_parameter_groups splits parameters purely on isinstance(mod, LayerNorm/BatchNorm/GroupNorm) or param_name.endswith('bias') and never inspects _no_weight_decay, so the A and D matrices always land in the weight-decayed group.
Observed: AD_weight_decay=True -> marker=None, in_decay_group=True, pred0=0.0298216827. AD_weight_decay=False -> marker=True, in_decay_group=True, pred0=0.0298216827. Bit-identical predictions even with weight_decay=0.5 and no_weight_decay_for_bias_and_norm=True, i.e. the flag has literally no effect on optimisation. grep -rn _no_weight_decay deeptab shows the attribute is only ever written (mamba.py), never read.
Expected: build_parameter_groups should route any parameter carrying _no_weight_decay=True into the zero-weight-decay group, so MambularConfig/MambAttentionConfig's documented AD_weight_decay=False actually exempts the A/D state-space matrices from weight decay.
Repro
import warnings; warnings.simplefilter('ignore')
import numpy as np, pandas as pd
from deeptab.configs import TrainerConfig
from deeptab.configs.models.mambular_config import MambularConfig
from deeptab.models import MambularRegressor
from deeptab.training.optimizers import build_parameter_groups
rng = np.random.default_rng(0)
X = pd.DataFrame({'n1': rng.normal(size=64), 'n2': rng.normal(size=64), 'n3': rng.uniform(size=64),
'c1': rng.choice(list('abc'), size=64), 'c2': rng.choice(list('xy'), size=64)})
y = (X['n1'] * 2 + X['n2']).to_numpy()
for ad in (True, False):
est = MambularRegressor(
model_config=MambularConfig(d_model=32, n_layers=1, d_state=8, AD_weight_decay=ad),
trainer_config=TrainerConfig(max_epochs=1, batch_size=16, patience=1, lr_patience=1,
weight_decay=0.5, no_weight_decay_for_bias_and_norm=True),
random_state=101)
est.fit(X, y, max_epochs=1, batch_size=16, patience=1, accelerator='cpu', devices=1,
enable_progress_bar=False, enable_model_summary=False, logger=False)
m = est._task_model.estimator
g = build_parameter_groups(m, weight_decay=0.5, no_weight_decay_for_bias_and_norm=True)
a = m.mamba.layers[0].layers.A_log_fwd
print('AD_weight_decay=', ad, 'marker=', getattr(a, '_no_weight_decay', None),
'in_decay_group=', any(p is a for p in g[0]['params']),
'pred0=', est.predict(X)[0])
Expected behavior
An option exposed in a config should either work or be rejected at construction. The silent no-ops are
the worse half: rnn_dropout=0.2 (a default), dilation, AD_weight_decay, cat_encoding,
batch_norm/layer_norm/activation and the ENODE head options all appear to be honoured while
having no effect on the model at all.
Screenshots
n/a
Desktop (please complete the following information):
- OS: macOS (Darwin 25.5.0, arm64)
- Python version: 3.11.15
- deeptab Version: 2.0.0 (main @ 4e6a359)
Additional context
torch 2.9.1, lightning 2.6.5, scikit-learn 1.9.0, numpy 2.4.6. Found in a second-pass review of v2.0.0
(seven independent lenses, each finding adversarially re-verified by a second reviewer, then re-run by
hand). Distinct from the already-filed #409-#426.
Describe the bug
A systematic sweep of every dataclass field in
deeptab/configs/against every shipped architectureturned up a further set of options that either crash or are silently not wired up. This complements
#423, which lists an earlier set; none of the items below are in it.
Every item was reproduced by execution.
AutoInt crashes when use_cls=True — the output head is sized without the CLS token
Where:
deeptab/architectures/autoint.py(69, 116)AutoInt computes n_inputs = sum(len(info) for info in feature_information) and builds self.head = nn.Linear(d_model * n_inputs, num_classes), but EmbeddingLayer honours config.use_cls and prepends a CLS token, so the flattened sequence is d_model * (n_inputs + 1) wide and the final matmul fails.
Observed: use_cls=False -> ok, head=Linear(in_features=160, out_features=1). use_cls=True -> RuntimeError: mat1 and mat2 shapes cannot be multiplied (13x192 and 160x1). 192 = 326 (5 features + CLS) vs the head's 160 = 325. AutoIntConfig documents use_cls as 'Whether to use a CLS token for pooling instead of averaging'.
Expected: n_inputs should be incremented by 1 when config.use_cls is set (SAINT does exactly this at deeptab/architectures/saint.py:66-67), so AutoInt(use_cls=True) trains normally.
Repro
Mambular's dilation is applied only to the backward conv, so it is a no-op by default and crashes when bidirectional=True
Where:
deeptab/nn/blocks/mamba.py(338-356)conv1d_fwd is constructed without dilation= while conv1d_bwd (built only when bidirectional=True) passes dilation=dilation; both keep padding=d_conv-1. Consequently MambularConfig.dilation is silently ignored under the default bidirectional=False, and enabling both dilation>1 and bidirectional=True crashes because the padding was never widened for the dilated kernel.
Observed: bidirectional=False dilation=1 -> conv1d_fwd.dilation=(1,), pred0=0.0357717164. bidirectional=False dilation=4 -> conv1d_fwd.dilation=(1,), pred0=0.0357717164 (bit-identical: dilation ignored). bidirectional=True dilation=1 -> ok. bidirectional=True dilation=4 -> RuntimeError: Calculated padded input size per channel: (11). Kernel size: (13). Kernel size can't be greater than actual input size.
Expected: dilation should be applied to conv1d_fwd as well, with padding computed as dilation*(d_conv-1) so both directions produce a valid, symmetric convolution. As shipped, MambularConfig.dilation (documented as 'Dilation factor for the convolution') either does nothing or crashes.
Repro
TabulaRNN's rnn_dropout (default 0.2) has no effect — every recurrent layer is built with num_layers=1
Where:
deeptab/nn/blocks/common.py(1861-1874 (ConvRNN.init rnn_args))ConvRNN stacks config.n_layers separate single-layer RNN/LSTM/GRU modules and passes dropout=self.rnn_dropout into each, but PyTorch discards the dropout argument when num_layers == 1, so TabulaRNNConfig.rnn_dropout is a documented, on-by-default option that regularises nothing.
Observed: rnn_dropout=0.0 -> [(1, 0.0), (1, 0.0), (1, 0.0)], pred0=0.5273975134. rnn_dropout=0.9 -> [(1, 0.9), (1, 0.9), (1, 0.0)], pred0=0.5273975134 — bit-identical predictions with dropout at 0.9. PyTorch also emits the standard 'dropout option adds dropout after all but last recurrent layer, so non-zero dropout expects num_layers greater than 1' warning, which is swallowed.
Expected: Either an explicit nn.Dropout(rnn_dropout) between the stacked single-layer RNNs, or a single rnn_layer(num_layers=n_layers, dropout=rnn_dropout). As shipped, every TabulaRNN user silently gets zero recurrent dropout despite the default of 0.2.
Repro
ENODE hardcodes its prediction head, silently ignoring head_layer_sizes / head_skip_layers / head_activation / head_use_batch_norm
Where:
deeptab/architectures/enode.py(7 (unused MLPhead import), 85-90)ENODE imports MLPhead but builds tabular_head as a fixed nn.Sequential(Linear(d_model,d_model), ReLU, Dropout(head_dropout), Linear(d_model,num_classes)), so four of the five documented head_* fields on ENODEConfig have no effect at all.
Observed: For head_layer_sizes=[] and [16, 8] the head is the identical Sequential(Linear(8,8), ReLU(), Dropout(0.3), Linear(8,1)) and pred0=0.0899604410 in both cases. head_skip_layers=True and head_use_batch_norm=True likewise produce bit-identical predictions.
Expected: ENODE should build its head with MLPhead(input_dim=..., config=config, output_dim=num_classes) like NODE does (deeptab/architectures/node.py:86-90), so the documented head_* configuration takes effect.
Repro
Three of the eight documented pooling_method values crash at model construction on every pooling-based architecture
Where:
deeptab/core/base_model.py(148-176 (initialize_pooling_layers))BaseModel.initialize_pooling_layers builds the learnable pooling submodules from config.dim_feedforward, an attribute only TabulaRNNConfig defines, so pooling_method in {learned_flatten, attention, gated} raises AttributeError on FTTransformer, TabTransformer, SAINT, Mambular and MambAttention; on TabulaRNN, learned_flatten instead dies because the caller passes n_inputs as a list.
Observed: avg/max/sum/last/cls succeed; learned_flatten, attention and gated each raise AttributeError: 'FTTransformerConfig' object has no attribute 'dim_feedforward'. Same for TabTransformerConfig, SAINTConfig, MambularConfig and MambAttentionConfig. TabulaRNN raises TypeError: empty(): argument 'size' failed to unpack the object at pos 2 with error "type must be tuple of ints, but got list" for learned_flatten, because deeptab/architectures/tabularnn.py:86 passes n_inputs=[len(info) for info in feature_information] (a list) into nn.Linear(n_inputs * config.dim_feedforward, ...).
Expected: Either the learnable pooling modes work (sized from config.d_model, which is the actual width of the (B, S, d_model) tensor pool_sequence receives), or pooling_method rejects them up front with a clear error. Silently shipping three implemented-but-unreachable pooling modes that only fail after preprocessing and model construction is a defect.
Repro
use_pscan=True crashes with TypeError: 'NoneType' object is not callable — the optional-dependency fallback is never taken
Where:
deeptab/nn/blocks/mamba.py(320-329, 517-518)When the optional mambapy package is absent, the constructor catches ImportError, prints an install hint and sets self.pscan = None, but selective_scan_seq still unconditionally calls self.pscan(...) whenever self.use_pscan is true, so Mambular/MambAttention with use_pscan=True die deep inside forward instead of falling back to the sequential scan that is sitting right there in the else branch.
Observed: TypeError: 'NoneType' object is not callable, raised at deeptab/nn/blocks/mamba.py:518 (
hs = self.pscan(deltaA, BX)) inside selective_scan_seq. Same for MambAttention. (import mambapy-> ModuleNotFoundError on this machine; mambapy is not a declared dependency in pyproject.toml.)Expected: The except-ImportError branch clearly intends graceful degradation, so it should also set self.use_pscan = False (falling back to the sequential scan), or raise a clear ImportError naming mambapy at construction time. Not a cryptic NoneType TypeError from inside the scan.
Repro
AutoInt's fprenorm option is dead — the code reads a nonexistent config.prenorm
Where:
deeptab/architectures/autoint.py(114)AutoInt gates its final normalisation on getattr(config, 'prenorm', False), but the field on AutoIntConfig is named fprenorm, so last_norm is always None and setting fprenorm=True changes nothing.
Observed: fprenorm=False -> last_norm=None, pred0=-0.3628833294. fprenorm=True -> last_norm=None, pred0=-0.3628833294 (bit-identical).
grep -n prenorm deeptab/configs/models/autoint_config.pyshows only fprenorm; no config in the repo definesprenorm.Expected: fprenorm=True should construct the nn.LayerNorm(d_model) and change the model's output. This is a distinct field from the kv_compression/activation items already filed under #423.
Repro
NODEConfig.norm and ENODEConfig.norm are documented options that no code ever reads
Where:
deeptab/architectures/node.py(whole file (no reference tonorm); same for deeptab/architectures/enode.py)Both NODEConfig and ENODEConfig declare
norm: str | None = Nonewith the docstring 'Type of normalization to use in the model', but neither architecture calls get_normalization_layer or otherwise reads the field, so setting it changes nothing and the model contains no normalisation module.Observed: NODE: norm=None/LayerNorm/BatchNorm all give pred0=-0.1135944128 and norm_modules=[]. ENODE: all three give pred0=0.0899604410 and norm_modules=[].
grep -n norm deeptab/architectures/node.pyreturns nothing.Expected: Either the field is wired to get_normalization_layer (as FTTransformer/SAINT/Mambular do) or it is removed from the config. Currently a user tuning
normon NODE/ENODE (including via HPO) is tuning a parameter that cannot change the model. This is a different field/model from the TabMnormitem in #423.Repro
BaseModelConfig.cat_encoding is validated and documented but never read by any model, embedding or preprocessing code
Where:
deeptab/configs/core.py(84-86 (docstring), 106 (field), 111-118 (validation))cat_encoding is a first-class BaseModelConfig field with strict validation against {'int','one-hot','linear'} and is described in the docs as something 'the preprocessing and embedding machinery rely on', but a repo-wide grep shows it is never consumed anywhere: the only other occurrences are dead entries in HPO
fixed_paramsdefault dicts and the HPO search space.Observed: All three encodings produce bit-identical predictions, on every one of the 14 architectures I swept (with and without use_embeddings=True) — 28 identical pairs. The grep confirms there is no read site.
Expected: cat_encoding should either drive the categorical encoding in EmbeddingLayer / the preprocessor, or be removed. As shipped it is a validated, documented dial that is wired to nothing, and docs/core_concepts/config_system.md:132 and docs/core_concepts/custom_models.md:37 both advertise it as live.
Repro
BaseModelConfig.batch_norm, layer_norm and activation are silently ignored by most architectures
Where:
deeptab/configs/core.py(75-81 (docstrings), 99-102 (fields), 147-154 (cross-field warning))batch_norm and layer_norm are documented as 'Whether to use batch/layer normalisation in the model body' and post_init even warns when both are enabled, implying they are live everywhere; in fact only MLP reads both and only MLP/TabM read batch_norm. Likewise
activation('Activation function used throughout the model body', re-documented per-model as e.g. 'Activation function for the transformer layers') is ignored by FTTransformer, TabTransformer, TabulaRNN, NODE, ENODE and NDTF. layer_norm_eps is ignored by every architecture that does not go through get_normalization_layer.Observed: All five FTTransformer variants return bit-identical predictions. Across the full sweep: batch_norm=True and layer_norm=True were bit-identical no-ops for FTTransformer, TabTransformer, AutoInt, SAINT, NODE, ENODE, NDTF, TabulaRNN, Mambular, MambAttention, MambaTab, ResNet (12 of 14; MLP honours both, TabM crashes on batch_norm per #423). layer_norm_eps was a no-op for MLP, ResNet, NODE, ENODE, NDTF, AutoInt, TabM, MambaTab.
activationwas a no-op for FTTransformer, TabTransformer, TabulaRNN, NODE, ENODE, NDTF (AutoInt is already covered by #423).Expected: Fields that a given architecture cannot honour should not be inherited/advertised for it — or the architectures should read them. Today a user (or the HPO search space) can set batch_norm/layer_norm/activation on almost any model and get exactly the same network back, with post_init even emitting a 'both batch_norm and layer_norm are set' warning for a pair of dials that do nothing.
Repro
AD_weight_decay=False is a silent no-op: the _no_weight_decay markers are never read by the optimizer builder
Where:
deeptab/nn/blocks/mamba.py(407-414 (setter); deeptab/training/optimizers.py:386-406 (consumer))MambaBlock tags A_log_fwd/D_fwd (and the bwd pair) with param._no_weight_decay = True when AD_weight_decay is False, but build_parameter_groups splits parameters purely on isinstance(mod, LayerNorm/BatchNorm/GroupNorm) or param_name.endswith('bias') and never inspects _no_weight_decay, so the A and D matrices always land in the weight-decayed group.
Observed: AD_weight_decay=True -> marker=None, in_decay_group=True, pred0=0.0298216827. AD_weight_decay=False -> marker=True, in_decay_group=True, pred0=0.0298216827. Bit-identical predictions even with weight_decay=0.5 and no_weight_decay_for_bias_and_norm=True, i.e. the flag has literally no effect on optimisation.
grep -rn _no_weight_decay deeptabshows the attribute is only ever written (mamba.py), never read.Expected: build_parameter_groups should route any parameter carrying _no_weight_decay=True into the zero-weight-decay group, so MambularConfig/MambAttentionConfig's documented AD_weight_decay=False actually exempts the A/D state-space matrices from weight decay.
Repro
Expected behavior
An option exposed in a config should either work or be rejected at construction. The silent no-ops are
the worse half:
rnn_dropout=0.2(a default),dilation,AD_weight_decay,cat_encoding,batch_norm/layer_norm/activationand the ENODE head options all appear to be honoured whilehaving no effect on the model at all.
Screenshots
n/a
Desktop (please complete the following information):
Additional context
torch 2.9.1, lightning 2.6.5, scikit-learn 1.9.0, numpy 2.4.6. Found in a second-pass review of v2.0.0
(seven independent lenses, each finding adversarially re-verified by a second reviewer, then re-run by
hand). Distinct from the already-filed #409-#426.