000 03742nam a22004215i 4500
999 _c387288
_d387288
001 387288
003 ES-MaUEC
005 20230311180718.0
006 a||||fo|||| 00| 0
007 cr nn 008mamaa
008 220601s2013 sz | s |||| 0|eng d
020 _a9783031018978
024 7 _a10.1007/978-3-031-01897-8
_2doi
040 _aES-MaUEC
_bspa
_cES-MaUEC
_dES-MaUEC
050 4 _aQA76.9.D3
_b2013 EB
100 1 _aGanti, Venkatesh
_eautor
_4aut
_4http://id.loc.gov/vocabulary/relators/aut
_9687219
245 1 0 _aData Cleaning
_cby Venkatesh Ganti, Anish Das Sarma
250 _a1st edition 2013
264 1 _aCham
_bSpringer International Publishing
_c2013
300 _a1 recurso en línea (XV, 69 páginas)
336 _atexto
_btxt
_2rdacontent
337 _aelectrónico
_bc
_2rdamedia
338 _arecurso electrónico
_bcr
_2rdacarrier
347 _aarchivo de texto
_bPDF
490 0 _aSynthesis Lectures on Data Management
_x2153-5426
505 0 _aPreface -- Acknowledgments -- Introduction -- Technological Approaches -- Similarity Functions -- Operator: Similarity Join -- Operator: Clustering -- Operator: Parsing -- Task: Record Matching -- Task: Deduplication -- Data Cleaning Scripts -- Conclusion -- Bibliography -- Authors' Biographies.
520 _aData warehouses consolidate various activities of a business and often form the backbone for generating reports that support important business decisions. Errors in data tend to creep in for a variety of reasons. Some of these reasons include errors during input data collection and errors while merging data collected independently across different databases. These errors in data warehouses often result in erroneous upstream reports, and could impact business decisions negatively. Therefore, one of the critical challenges while maintaining large data warehouses is that of ensuring the quality of data in the data warehouse remains high. The process of maintaining high data quality is commonly referred to as data cleaning. In this book, we first discuss the goals of data cleaning. Often, the goals of data cleaning are not well defined and could mean different solutions in different scenarios. Toward clarifying these goals, we abstract out a common set of data cleaning tasks that often need to be addressed. This abstraction allows us to develop solutions for these common data cleaning tasks. We then discuss a few popular approaches for developing such solutions. In particular, we focus on an operator-centric approach for developing a data cleaning platform. The operator-centric approach involves the development of customizable operators that could be used as building blocks for developing common solutions. This is similar to the approach of relational algebra for query processing. The basic set of operators can be put together to build complex queries. Finally, we discuss the development of custom scripts which leverage the basic data cleaning operators along with relational operators to implement effective solutions for data cleaning tasks.
988 _aSynthesis Collection of Technology_2013
650 7 _2embne
_9150569
_aSistemas de gestión de bases de datos
650 7 _2embne
_9141180
_aProceso de datos
700 1 _aDas Sarma, Anish
_eautor
_4aut
_4http://id.loc.gov/vocabulary/relators/aut
_9687220
776 0 8 _iPrinted edition:
_z9783031007699
776 0 8 _iPrinted edition:
_z9783031030253
856 4 0 _uhttps://go.openathens.net/redirector/universidadeuropea.es?url=https://doi.org/10.1007/978-3-031-01897-8
_zAcceso a este recurso digital (usuarios Universidad Europea de Madrid)
942 _2lcc
_cLE
998 _b03/2023
_dz
_esc
_zSI