Skip to contents

Performs fast, scalable probabilistic record linkage and deduplication using the Fellegi-Sunter model. Records lacking a shared unique identifier are compared across configurable dimensions using exact, fuzzy, and distance-based comparisons, with model parameters estimated via unsupervised Expectation-Maximization. Multiple SQL backends are supported through 'DBI', including execution via 'DuckDB'. This package is a translation of the Python 'splink' library by Linacre et al. (2022) doi:10.23889/ijpds.v7i3.1794 into idiomatic R.

Package options

  • irelink.show_sql: If TRUE, print every SQL statement irelink sends to the database as a message. Defaults to FALSE.

Author

Maintainer: Christopher T. Kenny ctkenny@proton.me (ORCID) [copyright holder]

Authors:

Other contributors:

  • Robin Linacre (Lead author of splink, the Python package this is derived from) [copyright holder]

  • Sam Lindsay (Author of splink) [copyright holder]

  • Theodore Manassis (Author of splink) [copyright holder]

  • Tom Hepworth (Author of splink) [copyright holder]

  • Andy Bond (Author of splink) [copyright holder]

  • Ross Kennedy (Author of splink) [copyright holder]

  • UK Ministry of Justice (Copyright holder of splink) [copyright holder]