Skip to content

v2026.02.13.1

Choose a tag to compare

@feranick feranick released this 13 Feb 22:46
· 181 commits to master since this release

Changelog:

  1. Common:
    • Program exit when files are not found is now handled more gracefully.
    • Added several sample input files for reference.
    • Improved error handling.
  2. DataML_DF:
    • Save optimal parameters directly after optimization.
    • New operational mode (-g) to generate optimal values for missing parameters, given the target prediction.
    • Allow optimization to use custom parameter file.
    • Optimized default hyperparameters.
  3. DataML_DAE:
    • Renamed from AddAutoEncoder.
    • Added ability to select the type of activation function (linear or sigmoid) using activation in ini file. The recommended use is linear for under-complete DAE (to prevent underestimation during data generation), sigmoid for overcomplete.
    • Added ability to include dropout in the encoder via ini flag dropout.
    • Added ability to add post-generation noise (using flag postGenerationNoise to enable it, and postGenerationNoise to set the max error to be added). This is needed to generate data with the original variance (i.e. spread).
    • increase default minimum dimension in DAE to reduce forced linearity.
    • Suggested initial values for ini file for both the under-complete and over-complete DAE are provided in the README.md.
    • New operational modes: (-t) to train the DAE, (-g) to generate new samples from DAE using csv prompt, (-a) to augment learning data.
    • Added ability to save best models and stop at best model options with new ini flags: stopAtBest, saveBestModel,metricBestModel.
    • Improved handling of data with null labels, vs null features. This can now be handled separately using the new ini flag: excludeZeroLabels.
    • Bug fixes: plotting data, optimizer selection.
  4. CorrAnalysis:
    • Several bug fixes in color and shapes of the markers in plots (including a better differentiation between valid and training samples.
    • Fixed bug with newer version of Numpy, that require a copy of the data.
    • Improved color mapping and legends with plotSpecificColors is specified. If customColors is False, a palette with legend is created to show which colors corresponds to which category. If customColors is True, the behavior remains unchanged.
  5. DataML_Maker:
    • Revised purging of empty rows. When validation rows are explicitly stated, purging crashes. This applies the purging on the training matrix only after validation data has been already removed.
  6. DataML_BatchMaker:
    • Generate a batch CSV file for batch prediction from parameter file. Uses custom DataML_BatchMaker.ini.
  7. ConverParamLabels:
    • Convert numeric labels created from training file (m123,m134,m136) into a similar file but using the proper labels from the parameter file (MAT_PAR23, MAT_PAR34...)
  8. ConverNormLabels: Renamed from ConvertLabel.py
  9. libDataML
    • getPrediction (for DF) and getPredictionsTF (for TF) are not both in `libDataML.
  10. Web versions - Pyscript.
    • DataML_DF: Updated to pyscript v2026.2.1.
    • Revised and updated meta and manifest files for proper handling of icons, and proper installation as web app.
    • scikit_learn models need to be generated with v1.7.0
    • Modularized files not to be dependent in naming to a single project.
    • igc_ml.js is now ml.js and does not change across project (i.e. can be used across projects).
    • Only index.html needs to be customized for a specific project.
    • batchCSVmaker: Generate a batch CSV file for batch prediction from parameter file.
  11. Utilities:
    • MergeTrainTestDatasets.: Merge train/test datasets
    • CheckDataSet: Uses CorrAnalysis.ini to read through the pd dataframe and checks for errors.
    • SplitPerfFromCorrAnal.py: Split CorrAnalysis complete master results into individual Perf files.
    • AddSinglePerfCorrAnalMaster.py: Add individual Perf sheets into CorrAnalysis Result master XLSX.
    • CreateMasterDatasetExcel.py: Create Master Dataset from provided excel file. This version includes column splitting for complex "AA-BB" sample codes, but can/should be adapted to generic needs. it includes column/row management (removal, renaming, etc) and handling of NaN.
    • SplitColumnCodesExcel.py: Process provided Excel datasets to create a master dataset with column splitting for complex "AA-BB" sample codes. It includes column/row management (removal, renaming, etc) and handling of NaN.
    • CreateSubsetExcel.py: Create subsets based on condition, in this case whether a value in a specific column is equal to a given.
    • ConvertToTFLiteK3.py: Updated for latest version of TensorFlow.
  12. Scripts: Moved all bash scripts from Utilities into separate folder.
    • New scripts for batch submission to slurm.
    • RemoveMLabelPar.sh: New bash script to remove m from config.txt or config_numeric.txt or from a sequence of parameters.
    • RemoveReductionTags.sh: New bash script to remove tags in files generated when performing feature reduction.