000 03918nam a22004215i 4500
999 _c386857
_d386857
001 386857
003 ES-MaUEC
005 20230130164536.0
006 a||||fo|||| 00| 0
007 cr nn 008mamaa
008 220601s2022 sz | s |||| 0|eng d
020 _a9783031021831
024 7 _a10.1007/978-3-031-02183-1
_2doi
040 _aES-MaUEC
_bspa
_cES-MaUEC
_dES-MaUEC
050 4 _aQ325.5
_b2022 EB
100 1 _aRiezler, Stefan
_eautor
_4aut
_4http://id.loc.gov/vocabulary/relators/aut
_9686343
245 1 0 _aValidity, Reliability, and Significance :
_bEmpirical Methods for NLP and Data Science
_cby Stefan Riezler, Michael Hagmann
250 _a1st edition 2022
264 1 _aCham
_bSpringer International Publishing
_c2022
300 _a1 recurso en línea (XVII, 147 páginas)
336 _atexto
_btxt
_2rdacontent
337 _aelectrónico
_bc
_2rdamedia
338 _arecurso electrónico
_bcr
_2rdacarrier
347 _aarchivo de texto
_bPDF
490 0 _aSynthesis Lectures on Human Language Technologies
_x1947-4059
505 0 _aPreface -- Acknowledgments -- Introduction -- Validity -- Reliability -- Significance -- Bibliography -- Authors' Biographies.
520 _aEmpirical methods are means to answering methodological questions of empirical sciences by statistical techniques. The methodological questions addressed in this book include the problems of validity, reliability, and significance. In the case of machine learning, these correspond to the questions of whether a model predicts what it purports to predict, whether a model's performance is consistent across replications, and whether a performance difference between two models is due to chance, respectively. The goal of this book is to answer these questions by concrete statistical tests that can be applied to assess validity, reliability, and significance of data annotation and machine learning prediction in the fields of NLP and data science. Our focus is on model-based empirical methods where data annotations and model predictions are treated as training data for interpretable probabilistic models from the well-understood families of generalized additive models (GAMs) and linear mixed effects models (LMEMs). Based on the interpretable parameters of the trained GAMs or LMEMs, the book presents model-based statistical tests such as a validity test that allows detecting circular features that circumvent learning. Furthermore, the book discusses a reliability coefficient using variance decomposition based on random effect parameters of LMEMs. Last, a significance test based on the likelihood ratio of nested LMEMs trained on the performance scores of two machine learning models is shown to naturally allow the inclusion of variations in meta-parameter settings into hypothesis testing, and further facilitates a refined system comparison conditional on properties of input data. This book can be used as an introduction to empirical methods for machine learning in general, with a special focus on applications in NLP and data science. The book is self-contained, with an appendix on the mathematical background on GAMs and LMEMs, and with an accompanying webpage including R code to replicate experiments presented in the book.
988 _aSynthesis Collection of Technology_2022
650 7 _2embne
_9166090
_aAprendizaje automático
700 1 _aHagmann, Michael
_eautor
_4aut
_4http://id.loc.gov/vocabulary/relators/aut
_9686342
776 0 8 _iPrinted edition:
_z9783031001949
776 0 8 _iPrinted edition:
_z9783031010552
776 0 8 _iPrinted edition:
_z9783031033117
856 4 0 _uhttps://go.openathens.net/redirector/universidadeuropea.es?url=https://doi.org/10.1007/978-3-031-02183-1
_zAcceso a este recurso digital (usuarios Universidad Europea de Madrid)
942 _2lcc
_cLE
998 _b01/2023
_dz
_esc
_zSI