
The way data storage is viewed has been changing consistently. The current trend in data storage is Hadoop which provides a scalable data storage mechanism for storing extremely large amount of data and to handle data intensive scientific applications. It makes use of the MapReduce framework and stores the data in HDFS(Hadoop Distributed File System). In HDFS architecture metadata is handled by the NameNode. In this paper, we propose a novel and efficient mechanism for managing the metadata dynamically by classifying the metadata based on its Importance Factor(If) which is a measure of the data's criticality, frequency of access and the importance of the client using the data. Metadata management is divided into three different techniques based on the importance. To save the amount of metadata from being a constraint on the main memory of the NameNode, the concept of sequence files is employed. This approach leads to more efficient low latency metadata operations, at the same time reduces the bottleneck of the NameNode main memory.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 2 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
