Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Report
Data sources: ZENODO
addClaim

Scaling Effects of Intermediate-Task Training on Zero-Shot Cross-Lingual Transfer in Small versus Large Multilingual Models for

Authors: Assignee Research;

Scaling Effects of Intermediate-Task Training on Zero-Shot Cross-Lingual Transfer in Small versus Large Multilingual Models for

Abstract

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuning again on the target task---often improves model performance substantially on language understanding tasks in monolingual English settings. We investigate whether English intermediate-task training is still helpful on non-English target tasks. Using nine intermediate language-understanding tasks, we evaluate intermediate-task transfer in a zero-shot cross-lingual setting on the XTREME benchmark. We see large improvements from intermediate training on the BUCC and Tatoeba sentence retrieval tasResearch goal: Does the effectiveness of English intermediate-task training for zero-shot cross-lingual transfer scale differently when using smaller pretrained multilingual models (e.g., 7B parameters) compared to larger models (e.g., 70B parameters) on the XCOPA task?Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 9.3/10.

Powered by OpenAIRE graph
Found an issue? Give us feedback