Yes, you read this blog title correctly – the PyTorch NLLLoss() function (“negative log likelihood”) for multi-class classification doesn’t actually compute a result. Bizarre. The bottom line: The NLLLoss(x,y) function expects x to be a tensor of three or more values, where each value is negative, and y to be a tensor with a single value that represents an index into x. The return result is just the negative of the value in x at the indicated index.
For example, suppose x = [[-2.5, -3.6, -1.7]] and y = [2] then NLLLoss(x,y) = 1.7 which is just the negative of x[2].
Conceptually, the code would look like:
def my_nll_loss(x, y): idx = y.item() # tensor to scalar result = x[0][idx] # fetch value in x return -1 * result.item() # negate and return
Weird.
In order to use NLLLoss(x,y) the x is assumed to be the log(softmax(logits)) where logits is some neural network raw computed output values. When you apply log(softmax(logits)) you are actually computing several possible negative log likelihoods from x. This isn’t at all obvious. Then the NLLLoss(x,y) function just uses y to select one of the negative values, flips it to positive, a viola! That is the loss value.
This is extremely confusing to scientists and engineers who are new to PyTorch. Because of this, PyTorch added a different mechanism. You can create a neural network with no activation at all, and then use the CrossEntropyLoss(x,y) function. This function will invisibly apply log_softmax() and then call the NLLLoss(x,y) function behind the scenes.
But this mechanism is conceptually confusing too, because unless you know the behind-the-scenes details, you would likely try to apply softmax() activation on your neural network output nodes and then feed that result to the CrossEntropyLoss() function. Strangely, this would not give you an error, but you would in effect be applying softmax() twice — once explicitly in the neural network and then once again invisibly in the automatic CrossEntropyloss() call. Your neural network would train, but training would be slow and give a poor result.
And sadly, there’s an additional tricky detail. I won’t dive down that rabbit hole, but briefly, log_softmax() is used instead of regular softmax() for engineering-computation reasons — to avoid arithmetic overflow — rather than for conceptual reasons. But that’s another story.
Details like this make learning how to use neural network code libraries like PyTorch and TensorFlow extremely difficult. This fact has motivated massive efforts to create systems that are easier to use. One such effort is Kera, a wrapper library on top of Tensorflow. Another effort is the development of various “AutoML” systems where a user just feeds data to the system and AutoML figures out the rest. And in fact, PyTorch itself is a Python abstraction layer on top of the C++ Torch library. There are dozens of other approaches to simplifying neural networks.
Low-level of abstraction approaches like PyTorch are needed for complex problems that need neural network customization. High-level of abstraction approaches like “AutoML” will be useful for simple problems where the user doesn’t have much knowledge of neural technologies.
Levels of abstraction.


.NET Test Automation Recipes
Software Testing
SciPy Programming Succinctly
Keras Succinctly
R Programming
Visual Studio Live
Microsoft MLADS Conference
DevIntersection Conference
Machine Learning Week
Ai4 Conference
G2E Conference
iSC West Conference
You must be logged in to post a comment.