A Ship of Theseus: Curious Cases of Paraphrasing in LLM-Generated Texts

Name: A Ship of Theseus: Curious Cases of Paraphrasing in LLM-Generated Texts
Keywords: FOS: Computer and information sciences, Computer Science - Computation and Language, Computation and Language (cs.CL)

Nafis Irtiza Tripto; Saranya Venkatraman; Dominik Macko; Róbert Móro; Ivan Srba; Adaku Uchendu; Thai Le; Dongwon Lee 0001

Found an issue? Give us feedback

arXiv.org e-Print Ar...arrow_drop_down

arXiv.org e-Print Archive

Preprint . 2023

Data sources: arXiv.org e-Print Archive

https://doi.org/10.18653/v1/20...

Article . 2024 . Peer-reviewed

Data sources: Crossref

https://dx.doi.org/10.48550/ar...

Article . 2023

License: CC BY

Data sources: Datacite

DBLP

Article

Data sources: DBLP

DBLP

Conference object

Data sources: DBLP

http://dx.doi.org/10.48550/arX...

Conference object . 2024

Data sources: European Union Open Data Portal

http://dx.doi.org/10.48550/ARX...

Other literature type . 2023

Data sources: European Union Open Data Portal

http://dx.doi.org/10.18653/v1/...

Conference object

Data sources: Sygma

A Ship of Theseus: Curious Cases of Paraphrasing in LLM-Generated Texts

descriptionPublicationkeyboard_double_arrow_right Article , Preprint , Conference object , Other literature type 01 Jan 2024Embargo end date: 01 Jan 2023Publisher:Association for Computational Linguistics (ACL)Journal:Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)Funded by:NSF | Collaborative Research: C..., EC | AI-CODE, NSF | REU Site: Machine Learnin... +5 projects

Authors: Nafis Irtiza Tripto; Saranya Venkatraman; Dominik Macko; Róbert Móro; Ivan Srba; Adaku Uchendu; Thai Le; +1 Authors

doi: 10.18653/v1/2024.acl-long.357 , 10.48550/arxiv.2311.08374

arXiv: 2311.08374

A Ship of Theseus: Curious Cases of Paraphrasing in LLM-Generated Texts

- Summary
- Subjects
- Metrics

Abstract

In the realm of text manipulation and linguistic transformation, the question of authorship has been a subject of fascination and philosophical inquiry. Much like the Ship of Theseus paradox, which ponders whether a ship remains the same when each of its original planks is replaced, our research delves into an intriguing question: Does a text retain its original authorship when it undergoes numerous paraphrasing iterations? Specifically, since Large Language Models (LLMs) have demonstrated remarkable proficiency in both the generation of original content and the modification of human-authored texts, a pivotal question emerges concerning the determination of authorship in instances where LLMs or similar paraphrasing tools are employed to rephrase the text--i.e., whether authorship should be attributed to the original human author or the AI-powered tool. Therefore, we embark on a philosophical voyage through the seas of language and authorship to unravel this intricate puzzle. Using a computational approach, we discover that the diminishing performance in text classification models, with each successive paraphrasing iteration, is closely associated with the extent of deviation from the original author's style, thus provoking a reconsideration of the current notion of authorship.

To appear in Association for Computational Linguistics (ACL 2024)

Related Organizations

University of California, San Francisco
United States
Kempelen Institute of Intelligent Technologies
Slovakia
KEMPELENOV INSTITUT INTELIGENTNYCH TECHNOLOGII
Slovakia
Pennsylvania State University
United States

Keywords

FOS: Computer and information sciences, Computer Science - Computation and Language, Computation and Language (cs.CL)

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	7
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Top 10%
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Average
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Top 10%

Found an issue? Give us feedback

7

Top 10%

Average

Top 10%

Green

Funded byView all

NSF| Collaborative Research: CISE-MSI: RCBP-RF: SaTC: Building Research Capacity in AI Based Anomaly Detection in Cybersecurity, EC| AI-CODE, NSF| REU Site: Machine Learning in Cybersecurity, NSF| Developing and Evaluating Fraud Informatics Curriculum among Institutions in the Appalachian Region