A New Measure of Similarity in Textual Analysis: Vector Similarity Metric versus Cosine Similarity Metric

ABSTRACT This paper proposes a new similarity metric, Vector Similarity Metric (VSM), which is as simple as the popular Cosine Similarity Metric (CSM). The CSM has a major deficiency. It yields the same value, irrespective of how different the two vectors are in their sizes so long as the angle between them is the same. This deficiency remains intact even when Natural Language Processing is used to associate semantic meanings to the words/phrases and when the term frequency is modified using Inverse Document Frequency. This deficiency becomes a serious concern when one is comparing the risk profile of one company with the risk profile of another company or investigating the changes in the risk profile of a company from one year to another. The VSM is based on the difference of the two vectors. The paper demonstrates the superiority of VSM over CSM analytically and through real-world examples.

Related Organizations

University of Kansas
United States

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	8
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Top 10%
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Average
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Average

Found an issue? Give us feedback

8

Top 10%

Average

Upload OA version

Are you the author of this publication? Upload your Open Access version to Zenodo!

It’s fast and easy, just two clicks!

uploadUpload now