BigDaMa

Data-cleaning tools from the Big Data Management group (BigDaMa) of Ziawasch Abedjan, TU Berlin, and the dirty/clean benchmarks published with them.

TU Berlin now lists BigDaMa as a former group. Abedjan leads the Data Integration and Data Preparation group (D2IP) at TU Berlin and BIFOLD.

Tools

Raha
Error detection without configuration. Many base detectors flag candidate errors; it then learns from a few user-labelled tuples which candidates are real errors.
Mahdavi, Abedjan, Castro Fernandez, Madden, Ouzzani, Stonebraker, Tang · SIGMOD 2019
Baran
Error correction. Combines several corrector models over one context representation of each value and generalizes from a few user fixes; can use transfer learning.
Mahdavi, Abedjan · PVLDB 13, 2020

Raha and Baran live in one repository, BigDaMa/raha (Apache-2.0), installable as pip install raha. The older BigDaMa/baran repository only points there.

Datasets

Also in BigDaMa/raha — not mirrored here

  • beers2,410 × 11 · 4,362 dirty cells
  • flights2,376 × 7 · 4,920 dirty cells
  • movies_17,390 × 17 · 7,675 dirty cells
  • rayyan1,000 × 11 · 948 dirty cells
  • tax200,000 × 15 · 121,219 dirty cells

Rows × columns and dirty cells (dirty ≠ clean) counted from datasets/ at commit 7be1334 using Raha's own CSV loader. Get the files from the upstream repository.