Core machine learning interview questions on the bias-variance trade-off, overfitting, regularization, and supervised vs unsupervised learning.
If our model is too simple and has very few parameters then it may have high bias and low variance. On the other hand if our model has large number of parameters then it’s going to have high variance and low bias. So we need to find the right/good balance without overfitting and underfitting the data. [src]
ML models learn a relationship between inputs (called training features) and outputs (called labels). After fitting on a training dataset, the model's performance on a held-out test set must be evaluated to determine how well it generalizes to unseen data.
Modern ML models have trainable parameters that are learned to build this input-output relationship. The more parameters a model has, the more complex a relationship it can learn between inputs and targets.
Underfitting occurs when the flexibility of a model is not sufficient to capture the underlying pattern in a training dataset. Overfitting, on the other hand, occurs when a model is too flexible and effectively memorizes the training data.
An example of underfitting is estimating a second-order polynomial (quadratic function) with a first-order polynomial (a simple line). Similarly, estimating a line with a tenth-order polynomial would be an example of overfitting.
Regularization is a technique that discourages learning a more complex or flexible model to combat overfitting. Examples - Ridge (L2 norm) - Lasso (L1 norm) The obvious disadvantage of ridge (L2) regression is model interpretability. It shrinks the coefficients for the least important predictors close to zero, but never makes them exactly zero. In other words, the final model includes all predictors, although some may have very little influence. In the case of lasso (L1) regularization, the penalty can force some coefficient estimates to be exactly equal to zero when the tuning parameter λ is sufficiently large. Therefore, lasso also performs variable selection and yields sparse models. [src]
In supervised learning, we train a model to learn the relationship between input data and output data. We need to have labeled data to be able to do supervised learning.
With unsupervised learning, we only have unlabeled data. The model learns a representation of the data. Unsupervised learning is frequently used to initialize the parameters of the model when we have a lot of unlabeled data and a small fraction of labeled data. We first train an unsupervised model and, after that, we use the weights of the model to train a supervised model.
In reinforcement learning, the model has some input data and a reward depending on the output of the model. The model learns a policy that maximizes the reward. Reinforcement learning has been applied successfully to strategic games such as Go and even classic Atari video games.
A generative model will learn categories of data while a discriminative model will simply learn the distinction between different categories of data. Discriminative models will generally outperform generative models on classification tasks.
Instance-based Learning: The system learns the examples by heart, then generalizes to new cases using a similarity measure.
Model-based Learning: Another way to generalize from a set of examples is to build a model of these examples, then use that model to make predictions. This is called model-based learning. [src]
The Turing test is a method to test the machine’s ability to match the human level intelligence. A machine is used to challenge the human intelligence that when it passes the test, it is considered as intelligent. Yet a machine could be viewed as intelligent without sufficiently knowing about people to mimic a human.