Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Report
Data sources: ZENODO
addClaim

Zero-shot cross-lingual performance of Llama 2, GPT-4, and Gemini on XTREME-R tasks with English-only vs. multilingual fine-tuning

Authors: Assignee Research;

Zero-shot cross-lingual performance of Llama 2, GPT-4, and Gemini on XTREME-R tasks with English-only vs. multilingual fine-tuning

Abstract

Pre-trained multilingual language models show significant performance gains for zero-shot cross-lingual model transfer on a wide range of natural language understanding (NLU) tasks. Previously, for zero-shot cross-lingual evaluation, pre-trained models are only fine-tuned on English data and tested on a variety of target languages. In this paper, we do cross-lingual evaluation on various NLU tasks (sentence classification, sequence labeling, question answering) using prompt-tuning and compare it with fine-tuning. The results show that prompt tuning achieves much better cross-lingual transfer tResearch goal: How does the zero-shot cross-lingual performance of Llama 2 compare to GPT-4 and Gemini on XTREME-R tasks when fine-tuned on English-only vs. multilingual intermediate tasks, measured by mXGLUE accuracy?Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 8.5/10.

Powered by OpenAIRE graph
Found an issue? Give us feedback