TCSR‐SQL: Towards Table Content‐Aware Text‐to‐SQL With Self‐Retrieval

Name: TCSR‐SQL: Towards Table Content‐Aware Text‐to‐SQL With Self‐Retrieval
Keywords: FOS: Computer and information sciences, Databases, Databases (cs.DB)

Wenbo Xu; Liang Yan; Chuanyi Liu; Peiyi Han; Haifeng Zhu; Yong Xu; Yingwei Liang; Bob Zhang

Found an issue? Give us feedback

CAAI Transactions on...arrow_drop_down

CAAI Transactions on Intelligence Technology

Article . 2025 . Peer-reviewed

License: CC BY NC

Data sources: Crossref

arXiv.org e-Print Archive

Preprint . 2024

Data sources: arXiv.org e-Print Archive

https://dx.doi.org/10.48550/ar...

Article . 2024

License: CC BY NC ND

Data sources: Datacite

DBLP

Article

Data sources: DBLP

TCSR‐SQL: Towards Table Content‐Aware Text‐to‐SQL With Self‐Retrieval

descriptionPublicationkeyboard_double_arrow_right Article , Preprint 05 Nov 2025Embargo end date: 01 Jan 2024 English Publisher:Institution of Engineering and Technology (IET)Journal:CAAI Transactions on Intelligence Technology (issn: 2468-6557, eissn: 2468-2322,

Copyright policy )

Authors: Wenbo Xu; Liang Yan; Chuanyi Liu; Peiyi Han; Haifeng Zhu; Yong Xu; Yingwei Liang; +1 Authors

doi: 10.1049/cit2.70071 , 10.48550/arxiv.2407.01183

arXiv: 2407.01183

TCSR‐SQL: Towards Table Content‐Aware Text‐to‐SQL With Self‐Retrieval

- Summary
- Subjects
- Metrics

Abstract

ABSTRACT Large language model‐based (LLM‐based) text‐to‐SQL methods have achieved important progress in generating SQL queries for real‐world applications. When confronted with table content‐aware questions in real‐world scenarios, ambiguous data content keywords and nonexistent database schema column names within the question lead to the poor performance of existing methods. To solve this problem, we propose a novel approach towards table content‐aware text‐to‐SQL with self‐retrieval (TCSR‐SQL). It leverages LLM's in‐context learning capability to extract data content keywords within the question and infer possible related database schema, which is used to generate Seed SQL to fuzz search databases. The search results are further used to confirm the encoding knowledge with the designed encoding knowledge table, including column names and exact stored content values used in the SQL. The encoding knowledge is sent to obtain the final Precise SQL following multi‐rounds of generation‐execution‐revision process. To validate our approach, we introduce a table‐content‐aware, question‐related benchmark dataset, containing 2115 question‐SQL pairs. Comprehensive experiments conducted on this benchmark demonstrate the remarkable performance of TCSR‐SQL, achieving an improvement of at least 27.8% in execution accuracy compared to other state‐of‐the‐art methods.

Related Organizations

Harbin Institute of Technology
China (People's Republic of)
Peng Cheng Laboratory
China (People's Republic of)
University of Science and Technology of China
China (People's Republic of)

Keywords

FOS: Computer and information sciences, Databases, Databases (cs.DB)

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	0
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Average
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Average
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Average

Found an issue? Give us feedback

0

Average

Green

gold

Related to Research communities

UArctic