
AstraQ-VL is a compact LLaVA-style astronomy vision-language model that combines a frozen CLIP ViT-L/14 vision encoder, a two-layer MLP connector, and Qwen2.5-1.5B-Instruct. Stage 1 trains only the connector, while Stage 2 warm-starts the connector and adds LoRA adapters to the language model.The paper compares the two stages on an untouched image-disjoint internal test set and evaluates their external behavior through zero-shot prompted closed-set classification on AstroVLBench Tasks 1–2 and single-reference solar-image captioning on DeepSDO. Stage 2 improves automatic reference alignment on the internal test set and the selected AstroVLBench aggregate, but its external behavior is task-dependent: it regresses on FIRST and does not improve the DeepSDO overlap measures. All reported outcomes are automatic measurements and do not establish scientific correctness or factual reliability.
multimodal learning, Qwen, LLaVA, automatic evaluation, AstraQ-VL, CLIP, vision-language models, Astrophysics, LoRA, astronomy, Machine Learning, Artificial Intelligence, Computer Science, parameter-efficient fine-tuning
multimodal learning, Qwen, LLaVA, automatic evaluation, AstraQ-VL, CLIP, vision-language models, Astrophysics, LoRA, astronomy, Machine Learning, Artificial Intelligence, Computer Science, parameter-efficient fine-tuning
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
