000 04280nam a22004575i 4500
999 _c387278
_d387278
001 387278
003 ES-MaUEC
005 20230220184153.0
006 a||||fo|||| 00| 0
007 cr nn 008mamaa
008 220601s2021 sz | s |||| 0|eng d
020 _a9783031018787
024 7 _a10.1007/978-3-031-01878-7
_2doi
040 _aES-MaUEC
_bspa
_cES-MaUEC
_dES-MaUEC
050 4 _aQA76.9.D3
_b2021 EB
100 1 _aPapadakis, George
_eautor
_4aut
_4http://id.loc.gov/vocabulary/relators/aut
_9687062
245 1 4 _aThe Four Generations of Entity Resolution
_cby George Papadakis, Ekaterini Ioannou, Emanouil Thanos, Themis Palpanas
250 _a1st edition 2021
264 1 _aCham
_bSpringer International Publishing
_c2021
300 _a1 recurso en línea (XVII, 152 páginas)
336 _atexto
_btxt
_2rdacontent
337 _aelectrónico
_bc
_2rdamedia
338 _arecurso electrónico
_bcr
_2rdacarrier
347 _aarchivo de texto
_bPDF
490 0 _aSynthesis Lectures on Data Management
_x2153-5426
505 0 _aPreface -- Acknowledgments -- Entity Resolution: Past, Present, and Yet-to-Come -- Preliminaries -- Generation 1: Addressing Veracity -- Generation 2: Also Addressing Volume -- Generation 3: Also Addressing Variety -- Generation 4: Also Addressing Velocity -- Leveraging External Knowledge -- Resources for Entity Resolution -- Possible Directions for Future Work -- Bibliography -- Authors' Biographies.
520 _aEntity Resolution (ER) lies at the core of data integration and cleaning and, thus, a bulk of the research examines ways for improving its effectiveness and time efficiency. The initial ER methods primarily target Veracity in the context of structured (relational) data that are described by a schema of well-known quality and meaning. To achieve high effectiveness, they leverage schema, expert, and/or external knowledge. Part of these methods are extended to address Volume, processing large datasets through multi-core or massive parallelization approaches, such as the MapReduce paradigm. However, these early schema-based approaches are inapplicable to Web Data, which abound in voluminous, noisy, semi-structured, and highly heterogeneous information. To address the additional challenge of Variety, recent works on ER adopt a novel, loosely schema-aware functionality that emphasizes scalability and robustness to noise. Another line of present research focuses on the additional challenge of Velocity, aiming to process data collections of a continuously increasing volume. The latest works, though, take advantage of the significant breakthroughs in Deep Learning and Crowdsourcing, incorporating external knowledge to enhance the existing words to a significant extent. This synthesis lecture organizes ER methods into four generations based on the challenges posed by these four Vs. For each generation, we outline the corresponding ER workflow, discuss the state-of-the-art methods per workflow step, and present current research directions. The discussion of these methods takes into account a historical perspective, explaining the evolution of the methods over time along with their similarities and differences. The lecture also discusses the available ER tools and benchmark datasets that allow expert as well as novice users to make use of the available solutions.
988 _aSynthesis Collection of Technology_2021
650 7 _2embne
_9162648
_aData mining
650 7 _2embne
_9150569
_aSistemas de gestión de bases de datos
700 1 _aIoannou, Ekaterini
_eautor
_4aut
_4http://id.loc.gov/vocabulary/relators/aut
_9687063
700 1 _aThanos, Emanouil
_eautor
_4aut
_4http://id.loc.gov/vocabulary/relators/aut
_9687064
700 1 _aPalpanas, Themis,
_eautor
_4aut
_4http://id.loc.gov/vocabulary/relators/aut
_9687052
_d1973-
776 0 8 _iPrinted edition:
_z9783031001055
776 0 8 _iPrinted edition:
_z9783031007507
776 0 8 _iPrinted edition:
_z9783031030062
856 4 0 _uhttps://go.openathens.net/redirector/universidadeuropea.es?url=https://doi.org/10.1007/978-3-031-01878-7
_zAcceso a este recurso digital (usuarios Universidad Europea de Madrid)
942 _2lcc
_cLE
998 _b02/2023
_dz
_esc
_zSI