
Parks are essential to urban well-being, making park satisfaction crucial for sustainable city development. Traditional survey-based approaches to understand sentiment towards parks among residents are often costly, time-consuming, and limited in scale. Recent social media–based studies have scaled such research but predominantly focus on text and frequently overlook visual information and the joint effects of text–image representations. This study presents an automated multimodal framework using crowdsourced reviews from Google Maps to model park satisfaction by integrating textual and visual features. Using Singapore as a case study, we analysed 76,869 textual reviews and 184,322 images associated with them. The results show that multimodal models are more useful than text-only approaches, with textual sentiment, emotional attributes, and image temporal characteristics identified as the most influential factors. These findings highlight the importance of multimodal analysis for advancing park research and informing planning and policy practices.
Urban green spaces, Park perception, Crowdsourced data, Vision-language model.
Urban green spaces, Park perception, Crowdsourced data, Vision-language model.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
