
doi: 10.1109/fcst.2010.46
Recently, a large amount of work has been done in XML data mining. However, most of the existing work focuses on the snapshot XML data, while XML data is dynamic in practical application. In order to mine knowledge hidden in the frozen structures (FS) which are not changed or very little changed during the historical changing process of an XML document, we present a method for clustering XML documents via FS. Also, a novel algorithm called weighted cosine measure (WCM) improved from the traditional algorithm has been proposed, and using which we can calculate the similarity between two clusters. Otherwise, we propose a method using the agglomerative hierarchical during the cluster process. Experiments results on our new algorithm indicate that the proposed solution performs significantly. XML documents can be effectively clustered and the results of using the WCM are better than using the traditional cosine measure. Then, XML documents in each cluster have similar structures not often changed.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 2 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
