| 000 | 04051nam a2200433 i 4500 | ||
|---|---|---|---|
| 999 |
_c387628 _d387628 |
||
| 001 | 387628 | ||
| 003 | ES-MaUEC | ||
| 005 | 20230327085123.0 | ||
| 006 | a||||fo|||| 00| 0 | ||
| 007 | cr nn 008mamaa | ||
| 008 | 220601s2022 sz | o |||| 0|eng d | ||
| 020 | _a9783031037634 | ||
| 024 | 7 |
_a10.1007/978-3-031-03763-4 _2doi |
|
| 040 |
_aES-MaUEC _bspa _cES-MaUEC _dES-MaUEC |
||
| 050 | 4 |
_aP98.5 .S83 _b2022 EB |
|
| 100 | 1 |
_aPaun, Silviu _eautor _4aut _4http://id.loc.gov/vocabulary/relators/aut _9687660 |
|
| 245 | 1 | 0 |
_aStatistical Methods for Annotation Analysis _cby Silviu Paun, Ron Artstein, Massimo Poesio |
| 250 | _a1st edition 2022 | ||
| 264 | 1 |
_aCham _bSpringer International Publishing _c2022 |
|
| 300 | _a1 recurso en línea (XIX, 197 páginas) | ||
| 336 |
_atexto _btxt _2rdacontent |
||
| 337 |
_aelectrónico _bc _2rdamedia |
||
| 338 |
_arecurso electrónico _bcr _2rdacarrier |
||
| 347 |
_aarchivo de texto _bPDF |
||
| 490 | 0 |
_aSynthesis Lectures on Human Language Technologies _x1947-4059 |
|
| 505 | 0 | _aPreface -- Acknowledgements -- Introduction -- Coefficients of Agreement -- Using Agreement Measures for CL Annotation Tasks -- Probabilistic Models of Agreement -- Probabilistic Models of Annotation -- Learning from Multi-Annotated Corpora -- Bibliography -- Authors' Biographies. | |
| 520 | _aLabelling data is one of the most fundamental activities in science, and has underpinned practice, particularly in medicine, for decades, as well as research in corpus linguistics since at least the development of the Brown corpus. With the shift towards Machine Learning in Artificial Intelligence (AI), the creation of datasets to be used for training and evaluating AI systems, also known in AI as corpora, has become a central activity in the field as well. Early AI datasets were created on an ad-hoc basis to tackle specific problems. As larger and more reusable datasets were created, requiring greater investment, the need for a more systematic approach to dataset creation arose to ensure increased quality. A range of statistical methods were adopted, often but not exclusively from the medical sciences, to ensure that the labels used were not subjective, or to choose among different labels provided by the coders. A wide variety of such methods is now in regular use. This book is meant to provide a survey of the most widely used among these statistical methods supporting annotation practice. As far as the authors know, this is the first book attempting to cover the two families of methods in wider use. The first family of methods is concerned with the development of labelling schemes and, in particular, ensuring that such schemes are such that sufficient agreement can be observed among the coders. The second family includes methods developed to analyze the output of coders once the scheme has been agreed upon, particularly although not exclusively to identify the most likely label for an item among those provided by the coders. The focus of this book is primarily on Natural Language Processing, the area of AI devoted to the development of models of language interpretation and production, but many if not most of the methods discussed here are also applicable to other areas of AI, or indeed, to other areas of Data Science. | ||
| 988 | _aSynthesis Collection of Technology_2022 | ||
| 650 | 7 |
_2embne _aLingüística computacional _xMétodos estadísticos _9666075 |
|
| 700 | 1 |
_aArtstein, Ron _eautor _4aut _4http://id.loc.gov/vocabulary/relators/aut |
|
| 700 |
_aPoesio, Massimo _eautor _4aut _4http://id.loc.gov/vocabulary/relators/aut |
||
| 776 | 0 | 8 |
_iPrinted edition: _z9783031037733 |
| 776 | 0 | 8 |
_iPrinted edition: _z9783031037535 |
| 776 | 0 | 8 |
_iPrinted edition: _z9783031037832 |
| 856 | 4 | 0 |
_uhttps://go.openathens.net/redirector/universidadeuropea.es?url=https://doi.org/10.1007/978-3-031-03763-4 _zAcceso a este recurso digital (usuarios Universidad Europea de Madrid) |
| 942 |
_2lcc _cLE |
||
| 998 |
_b03/2023 _dz _eb _zSI |
||