TF-IDF in a Nutshell

TF-IDF in a Nutshell

In this article I’ll try to show in very clear examples how TF standing for ‘term frequency’ and it’s counterpart IDF standing for ‘inverse document frequency’ help to find what you’re looking for. As was said before according to the Luhn’s assumption documents containing more occurrences of a query term might be more relevant. Let’s rank the documents by their term frequencies (note here and in the area of Informaton Retrieval in general ‘frequency’ means just count, not 1/count as in Physics):

This solves the problem, now documents #5 and #2 are in the very top.

Source: manticoresearch.com