Downloads provided by UsageCounts
Files composing the YADL data lake, for the paper "Benchmarking Data Lake for Join Discovery and Learning with Relational Data" Tabular representation learning is gaining traction as machine learning techniques are increasingly applied to database problems, including data integration across multiple tables. A challenge lies in the split focus among research questions: merging a set of different tables to assemble a larger feature matrix without considering the downstream task (common in database research), or utilizing the assembled feature matrix for model training (prevalent in machine learning). Despite significant work conducted separately, there has been limited emphasis on the end-to-end process – as evidenced by the lack of suitable benchmarks for this issue. This paper introduces the first benchmark data lake on the subject, aiming to encourage reproducible research on learning from data lakes. Using a proof-of-principle complete analytic pipeline, it demonstrates the benefit of studying assembling tables for a supervised-learning goal.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
| views | 31 | |
| downloads | 19 |

Views provided by UsageCounts
Downloads provided by UsageCounts