
doi: 10.2139/ssrn.6226478
National Natural Science Foundation of China (NSFC) is one of the most important public grants supporting scientific research. However, little is known about its allocation and the major obstacle to analyze it is the lack of disambiguated grantees data. In this paper, taking advantage of a labeled data from NSFC, we use the supervised learning algorithms to get the first systematic disambiguation result of all NSFC grants from 1986 to 2023. Our algorithm achieves a high F1 score of 0.979. Based on the disambiguated data, we document several stylized facts of China's research system, including the rapid expansion of the funded scientist population followed by a recent slowdown in new entrants, persistently low institutional mobility, and sustained concentration of high-value funding in elite institutions despite partial decentralization at the aggregate level. We make our source codes publicly available at https://github.com/jiyuzhang-glitch/NSFC-disam to support future studies.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
