Deep Learning Interview Questions and Answers (2026)

Deep learning interview questions on ReLU, batch normalization, dropout, vanishing gradients, ResNets, LSTMs, transformers, autoencoders and GANs.

13 questions on this page

How does a neural network work?

A neural network is a function approximator that learns transformations between inputs and outputs by propagating information through layers of neurons. At its simplest, a neural network computes y = f(Wx + b), where W contains the weights, b contains the biases, and f is a nonlinear activation function. The weights and biases are learned during training through backpropagation.

Why is ReLU better and more often used than Sigmoid in Neural Networks?

[src1] [src2]

[src]

List different activation neurons or functions.

[src]

What is vanishing gradient?

As we add more and more hidden layers, back propagation becomes less and less useful in passing information to the lower layers. In effect, as information is passed back, the gradients begin to vanish and become small relative to the weights of the networks.

[src]

What are dropouts?

Dropout is a simple way to prevent a neural network from overfitting. It is the dropping out of some of the units in a neural network. It is similar to the natural reproduction process, where the nature produces offsprings by combining distinct genes (dropping out others) rather than strengthening the co-adapting of them.

[src]

What is batch normalization and why does it work?

Training Deep Neural Networks is complicated by the fact that the distribution of each layer's inputs changes during training, as the parameters of the previous layers change. The idea is then to normalize the inputs of each layer in such a way that they have a mean output activation of zero and standard deviation of one. This is done for each individual mini-batch at each layer i.e compute the mean and variance of that mini-batch alone, then normalize. This is analogous to how the inputs to networks are standardized. How does this help? We know that normalizing the inputs to a network helps it learn. But a network is just a series of layers, where the output of one layer becomes the input to the next. That means we can think of any layer in a neural network as the first layer of a smaller subsequent network. Thought of as a series of neural networks feeding into each other, we normalize the output of one layer before applying the activation function, and then feed it into the following layer (sub-network).

[src]

What is the significance of Residual Networks?

The main thing that residual connections did was allow for direct feature access from previous layers. This makes information propagation throughout the network much easier. One very interesting paper about this shows how using local skip connections gives the network a type of ensemble multi-path structure, giving features multiple paths to propagate throughout the network.

[src]

Define LSTM.

Long Short Term Memory – are explicitly designed to address the long term dependency problem, by maintaining a state what to remember and what to forget.

[src]

List the key components of LSTM.

[src]

List the variants of RNN.

[src]

What is the basic difference between LSTM and Transformers?

LSTMs (Long Short Term Memory) models consist of RNN cells designed to store and manipulate information across time steps more efficiently. In contrast, Transformer models contain a stack of encoder and decoder layers, each consisting of self attention and feed-forward neural network components.

[src]

What is Autoencoder, name few applications.

Auto encoder is basically used to learn a compressed form of given data. Few applications include - Data denoising - Dimensionality reduction - Image reconstruction - Image colorization

[src]

What are the components of GAN?

[src]

← OptimizationComputer Vision →