
Social network has become a very popular way for internet users to communicate and interact online. Users spend a great deal of time on famous social networks (e.g. Facebook, Twitter, Sina Weibo, etc.), reading news, discussing events and posting their messages. Unfortunately, this popularity also attracts a significant amount of spammers who continuously expose malicious behaviors (e.g. Post messages containing commercial topics or URLs, following a larger amount of users, etc.), leading to great inconvenience on normal users' social activities. In this paper, a supervised machine learning based spammer filtering method is proposed. We first collected a dataset from Sina Weibo that includes 30,116 users and more than 16 million messages, then, construct a labeled dataset of users and manually classify users into spammers and non-spammers, after that, abstract a set of novel features from message content and users' social behavior, and apply into SVM based spammer classifier. Our experiments show that true positive rate of spammers and non-spammers could reach 99.1% and 99.9%.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 6 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Top 10% | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
