| 000 | 03742nam a22004215i 4500 | ||
|---|---|---|---|
| 999 |
_c387288 _d387288 |
||
| 001 | 387288 | ||
| 003 | ES-MaUEC | ||
| 005 | 20230311180718.0 | ||
| 006 | a||||fo|||| 00| 0 | ||
| 007 | cr nn 008mamaa | ||
| 008 | 220601s2013 sz | s |||| 0|eng d | ||
| 020 | _a9783031018978 | ||
| 024 | 7 |
_a10.1007/978-3-031-01897-8 _2doi |
|
| 040 |
_aES-MaUEC _bspa _cES-MaUEC _dES-MaUEC |
||
| 050 | 4 |
_aQA76.9.D3 _b2013 EB |
|
| 100 | 1 |
_aGanti, Venkatesh _eautor _4aut _4http://id.loc.gov/vocabulary/relators/aut _9687219 |
|
| 245 | 1 | 0 |
_aData Cleaning _cby Venkatesh Ganti, Anish Das Sarma |
| 250 | _a1st edition 2013 | ||
| 264 | 1 |
_aCham _bSpringer International Publishing _c2013 |
|
| 300 | _a1 recurso en línea (XV, 69 páginas) | ||
| 336 |
_atexto _btxt _2rdacontent |
||
| 337 |
_aelectrónico _bc _2rdamedia |
||
| 338 |
_arecurso electrónico _bcr _2rdacarrier |
||
| 347 |
_aarchivo de texto _bPDF |
||
| 490 | 0 |
_aSynthesis Lectures on Data Management _x2153-5426 |
|
| 505 | 0 | _aPreface -- Acknowledgments -- Introduction -- Technological Approaches -- Similarity Functions -- Operator: Similarity Join -- Operator: Clustering -- Operator: Parsing -- Task: Record Matching -- Task: Deduplication -- Data Cleaning Scripts -- Conclusion -- Bibliography -- Authors' Biographies. | |
| 520 | _aData warehouses consolidate various activities of a business and often form the backbone for generating reports that support important business decisions. Errors in data tend to creep in for a variety of reasons. Some of these reasons include errors during input data collection and errors while merging data collected independently across different databases. These errors in data warehouses often result in erroneous upstream reports, and could impact business decisions negatively. Therefore, one of the critical challenges while maintaining large data warehouses is that of ensuring the quality of data in the data warehouse remains high. The process of maintaining high data quality is commonly referred to as data cleaning. In this book, we first discuss the goals of data cleaning. Often, the goals of data cleaning are not well defined and could mean different solutions in different scenarios. Toward clarifying these goals, we abstract out a common set of data cleaning tasks that often need to be addressed. This abstraction allows us to develop solutions for these common data cleaning tasks. We then discuss a few popular approaches for developing such solutions. In particular, we focus on an operator-centric approach for developing a data cleaning platform. The operator-centric approach involves the development of customizable operators that could be used as building blocks for developing common solutions. This is similar to the approach of relational algebra for query processing. The basic set of operators can be put together to build complex queries. Finally, we discuss the development of custom scripts which leverage the basic data cleaning operators along with relational operators to implement effective solutions for data cleaning tasks. | ||
| 988 | _aSynthesis Collection of Technology_2013 | ||
| 650 | 7 |
_2embne _9150569 _aSistemas de gestión de bases de datos |
|
| 650 | 7 |
_2embne _9141180 _aProceso de datos |
|
| 700 | 1 |
_aDas Sarma, Anish _eautor _4aut _4http://id.loc.gov/vocabulary/relators/aut _9687220 |
|
| 776 | 0 | 8 |
_iPrinted edition: _z9783031007699 |
| 776 | 0 | 8 |
_iPrinted edition: _z9783031030253 |
| 856 | 4 | 0 |
_uhttps://go.openathens.net/redirector/universidadeuropea.es?url=https://doi.org/10.1007/978-3-031-01897-8 _zAcceso a este recurso digital (usuarios Universidad Europea de Madrid) |
| 942 |
_2lcc _cLE |
||
| 998 |
_b03/2023 _dz _esc _zSI |
||