000 04305nam a22004575i 4500
999 _c387761
_d387761
001 387761
003 ES-MaUEC
005 20230402110104.0
006 a||||fo|||| 00| 0
007 cr nn 008mamaa
008 220601s2020 sz | s |||| 0|eng d
020 _a9783031021749
024 7 _a10.1007/978-3-031-02174-9
_2doi
040 _aES-MaUEC
_bspa
_cES-MaUEC
_dES-MaUEC
050 4 _aQA76.9.N38
_b2020 EB
100 1 _aDror, Rotem
_eautor
_4aut
_4http://id.loc.gov/vocabulary/relators/aut
_9687991
245 1 0 _aStatistical Significance Testing for Natural Language Processing
_cby Rotem Dror, Lotem Peled-Cohen, Segev Shlomov, Roi Reichart
250 _a1st edition 2020
264 1 _aCham
_bSpringer International Publishing
_c2020
300 _a1 recurso en línea (XVII, 98 páginas)
336 _atexto
_btxt
_2rdacontent
337 _aelectrónico
_bc
_2rdamedia
338 _arecurso electrónico
_bcr
_2rdacarrier
347 _aarchivo de texto
_bPDF
490 0 _aSynthesis Lectures on Human Language Technologies
_x1947-4059
505 0 _aPreface -- Acknowledgments -- Introduction -- Statistical Hypothesis Testing -- Statistical Significance Tests -- Statistical Significance in NLP -- Deep Significance -- Replicability Analysis -- Open Questions and Challenges -- Conclusions -- Bibliography -- Authors' Biographies.
520 _aData-driven experimental analysis has become the main evaluation tool of Natural Language Processing (NLP) algorithms. In fact, in the last decade, it has become rare to see an NLP paper, particularly one that proposes a new algorithm, that does not include extensive experimental analysis, and the number of involved tasks, datasets, domains, and languages is constantly growing. This emphasis on empirical results highlights the role of statistical significance testing in NLP research: If we, as a community, rely on empirical evaluation to validate our hypotheses and reveal the correct language processing mechanisms, we better be sure that our results are not coincidental. The goal of this book is to discuss the main aspects of statistical significance testing in NLP. Our guiding assumption throughout the book is that the basic question NLP researchers and engineers deal with is whether or not one algorithm can be considered better than another one. This question drives the field forward as it allows the constant progress of developing better technology for language processing challenges. In practice, researchers and engineers would like to draw the right conclusion from a limited set of experiments, and this conclusion should hold for other experiments with datasets they do not have at their disposal or that they cannot perform due to limited time and resources. The book hence discusses the opportunities and challenges in using statistical significance testing in NLP, from the point of view of experimental comparison between two algorithms. We cover topics such as choosing an appropriate significance test for the major NLP tasks, dealing with the unique aspects of significance testing for non-convex deep neural networks, accounting for a large number of comparisons between two NLP algorithms in a statistically valid manner (multiple hypothesis testing), and, finally, the unique challenges yielded by the nature of the data and practices of the field.
988 _aSynthesis Collection of Technology_2020
650 7 _2embne
_9687996
_aContrastación de hipótesis (Estadística)
650 7 _2embne
_9158738
_aProceso en lenguaje natural (Informática)
700 1 _aPeled-Cohen, Lotem
_eautor
_4aut
_4http://id.loc.gov/vocabulary/relators/aut
_9687992
700 1 _aShlomov, Segev
_eautor
_4aut
_4http://id.loc.gov/vocabulary/relators/aut
_9687993
700 1 _aReichart, Roi,
_d1980-
_eautor
_4aut
_4http://id.loc.gov/vocabulary/relators/aut
_9687994
776 0 8 _iPrinted edition:
_z9783031001857
776 0 8 _iPrinted edition:
_z9783031010460
776 0 8 _iPrinted edition:
_z9783031033025
856 4 0 _uhttps://go.openathens.net/redirector/universidadeuropea.es?url=https://doi.org/10.1007/978-3-031-02174-9
_zAcceso a este recurso digital (usuarios Universidad Europea de Madrid)
942 _2lcc
_cLE
998 _b04/2023
_dz
_esc
_zSI