Bias correction in clustering coefficient estimation

descriptionPublicationkeyboard_double_arrow_right Article , Conference object 01 Dec 2017Publisher:IEEEJournal:2017 IEEE International Conference on Big Data (Big Data)

Authors: Roohollah Etemadi; Jianguo Lu;

doi: 10.1109/bigdata.2017.8257976

Bias correction in clustering coefficient estimation

- Summary
- Metrics

Abstract

Clustering coefficient (C) is an important structural property to understand the complex structure of a graph. Calculating C is a computationally intensive task. Thereby, sampling-based methods have attracted substantial research for estimating C, and the closely related metric, the number of triangles. Unfortunately, widely used estimators for C are biased. We quantify the bias using Taylor expansion and find that the bias can be determined by the number of shared wedges and triangles in the sample. Based on the understanding of the bias, we give a new estimator that corrects the bias. The results are derived analytically and verified extensively in 56 networks ranging in different size and structure. The experiments reveal that the bias ranges widely from data to data. The relative bias can be as high as 4% or can be negative. For most of the graphs, the bias is small, although every graph does have a bias as quantified by our analytical results. Negative or small biases occur in online social networks where clustering coefficient is typically high. Positive and large biases typically occur in Web graphs, where there are nodes with high degrees but few neighboring nodes connecting with each other.

Related Organizations

University of Windsor
Canada

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	3
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Average
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Average
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Average

Found an issue? Give us feedback

3

Average

Fields of Science

engineering and technology

electrical engineering, electronic engineering, information engineering

Fields of Science

engineering and technology

electrical engineering, electronic engineering, information engineering

Upload OA version

Are you the author of this publication? Upload your Open Access version to Zenodo!

It’s fast and easy, just two clicks!

uploadUpload now