7 items
Activation functions give neural networks their nonlinear structure.
The rectified linear unit, usually called ReLU, is the most widely used activation function in modern deep learning.
ReLU is simple and effective, but it has one sharp weakness.
ReLU and its variants improved optimization in deep networks, but they still have limitations.
Many neural networks produce raw scores. These scores are called logits.
Activation functions control both the forward signal and the backward signal.
Activation functions should be chosen for the architecture, loss, initialization, normalization, and training scale. There is no universal best activation. The right choice depends on what the layer must do.