-
Notifications
You must be signed in to change notification settings - Fork 39
Machine Learning Algorithms
Here is a list of the machine learning algorithms available in the Foundry. These algorithms are all implemented in Java and are designed to be used in applications and extended for research.
Supervised learning algorithms take input-output pairs to train a function that attempts generalize to produce outputs for new and unseen inputs.
Batch supervised learning is one of the most typical type of machine learning algorithm. They are given a collection of input-output pairs of examples and train a function to generalize from them.
- AdaBoost
- Bagging and Balanced Bagging
- Bayesian Linear Regression and Bayesian Robust Linear Regression
- Decision Tree
- IVoting and Balanced IVoting
- Kernel Regression
- Gaussian Process Regression
- Naive Bayes
- Nearest Neighbor andK-Nearest Neighbor
- Linear Regression and Multivariate Linear Regression
- Locally-weighted Regression
- Logistic Regression
- Perceptron
- Random Forests
- Regression Tree
- Robust Regression
- Support Vector Machine via Sequential Minimal Optimization (SMO), Successive Overrelaxation, and Primal Estimated Sub-Gradient Solver (PEGASOS) with a variety of kernels
These are supervised algorithms that can learn incrementally from a stream of data, commonly called online learning. Many have both linear and kernel forms.
- Adaptive Regularization of Weights (AROW)
- Adatron
- Aggressive Relaxed Online Maximum Margin Algorithm (AROMMA)
- Ballseptron
- Confidence Weighted Linear Classification
- Forgetron
- Margin Infused Relaxed Algorithm (MIRA)
- Online Bagging
- Online Perceptron
- Online Voted Perceptron
- Passive-Aggressive Perceptron (PA-I and PA-II)
- Projectron
- Ramp Loss Passive-Aggressive Perceptron (PA^R)
- Relaxed Online Maximum Margin Algorithm (ROMMA)
- Shifting Perceptron
- Stoptron
- Winnow
The unsupervised learning algorithms are used with data that is not labeled. Clustering algorithms are usually dependent on using a provided distance metric.
- Affinity Propagation
- Dirichlet Process Clustering
- Dirichlet Process Mixture Model
- Generalized Hebbian Algorithm
- Hidden Markov Model
- Hierarchical Agglomerative Clustering
- Gaussian Mixture Model
- K-Means Clustering and Generalized Expectation Maximization Hard Assignment Clustering
- Partitional Custering
- Markov Chain
- Principal Components Analysis (PCA) and Kernel Principal Components Analysis
- Thin Singular Value Decomposition
General optimization methods can usually work with a variety of learned function types. A common one would be a Generalized Linear Model (GLM) or Neural Network that can be used with various activation functions.
- Broyden-Fletcher-Goldfarb-Shanno (BFGS)
- Conjugate Gradient
- Davidon-Flecher-Powell (DFP)
- Direction Set (Powell's Method)
- Downhill Simplex (Nelder-Mead)
- Fletcher-Revees conjugate gradient
- Fletcher Xu Hybrid Estimation
- Gauss-Newton
- Genetic Algorithms
- Gradient Descent
- Least-squares
- Levenberg Marquardt
- Liu-Storey conjugate gradient
- Polack-Ribiere conjugate gradient
- Root Finding
- Simulated Annealing
These algorithms are in the statistics packages and can be used with a variety of distributions
- Importance Sampling
- Kalman Filtering
- Markov Chain Monte Carlo (MCMC)
- Monte Carlo Integration
- Metropolis Hastings
- Particle Filtering and Sampling Importance Resampling Particle Filter
- Rejection Sampling
These are algorithms in the Text package that are typically used for topic detection.
- Latent Semantic Analysis (LSA)
- Latent Dirichlet Allocation (LDA)
- Probabilistic Latent Semantic Analysis (pLSA)
These are just simple baseline learners that others can be compared against.