000 03734nam a22004815i 4500
999 _c383079
_d383079
_x1
001 383079
003 ES-MaUEC
005 20230102122040.0
006 a||||fo|||| 00| 0
007 cr nn 008mamaa
008 221105s2022 si | s |||| 0|eng d
020 _a9789811909641
024 7 _a10.1007/978-981-19-0964-1
_2doi
040 _aES-MaUEC
_bspa
_cES-MaUEC
_dES-MaUEC
050 4 _aTA1634
_b2022 EB
100 1 _aWu, Qi
_eautor
_0(orcid)0000-0003-3631-256X
_1https://orcid.org/0000-0003-3631-256X
_4aut
_4http://id.loc.gov/vocabulary/relators/aut
_9685208
245 1 0 _aVisual Question Answering :
_bfrom Theory to Application
_cby Qi Wu, Peng Wang, Xin Wang, Xiaodong He, Wenwu Zhu
250 _aFirst edition 2022
264 1 _aSingapore
_bSpringer International Publising
_c2022
300 _a1 recurso en línea (XIII, 238 páginas)
_b104 ilustraciones, 92 ilustraciones a color
336 _atexto
_btxt
_2rdacontent
337 _aelectrónico
_bc
_2rdamedia
338 _arecurso electrónico
_bcr
_2rdacarrier
347 _aarchivo de texto
_bPDF
490 0 _aAdvances in Computer Vision and Pattern Recognition
_x2191-6594
505 0 _a1. Introduction -- 2. Deep Learning Basics -- 3. Question Answering (QA) Basics -- 4. The Classical Visual Question Answering -- 5. Knowledge-based VQA.
520 _aVisual Question Answering (VQA) usually combines visual inputs like image and video with a natural language question concerning the input and generates a natural language answer as the output. This is by nature a multi-disciplinary research problem, involving computer vision (CV), natural language processing (NLP), knowledge representation and reasoning (KR), etc. Further, VQA is an ambitious undertaking, as it must overcome the challenges of general image understanding and the question-answering task, as well as the difficulties entailed by using large-scale databases with mixed-quality inputs. However, with the advent of deep learning (DL) and driven by the existence of advanced techniques in both CV and NLP and the availability of relevant large-scale datasets, we have recently seen enormous strides in VQA, with more systems and promising results emerging. This book provides a comprehensive overview of VQA, covering fundamental theories, models, datasets, and promising future directions. Given its scope, it can be used as a textbook on computer vision and natural language processing, especially for researchers and students in the area of visual question answering. It also highlights the key models used in VQA.
988 _aSpringer_Computer_2022
650 7 _2embne
_9159793
_aVisión por ordenador
650 7 _2embne
_9158738
_aProceso en lenguaje natural (Informática)
700 1 _aWang, Peng
_eautor
_0(orcid)0000-0001-7689-3405
_1https://orcid.org/0000-0001-7689-3405
_4aut
_4http://id.loc.gov/vocabulary/relators/aut
_9685209
700 1 _aWang, Xin
_eautor
_0(orcid)0000-0002-0351-2939
_1https://orcid.org/0000-0002-0351-2939
_4aut
_4http://id.loc.gov/vocabulary/relators/aut
_9681443
700 1 _aHe, Xiaodong,
_eautor
_4aut
_4http://id.loc.gov/vocabulary/relators/aut
_9685210
_d1973-
700 1 _aZhu, Wenwu
_eautor
_4aut
_4http://id.loc.gov/vocabulary/relators/aut
_999421
776 0 8 _iPrinted edition:
_z9789811909634
776 0 8 _iPrinted edition:
_z9789811909658
776 0 8 _iPrinted edition:
_z9789811909665
856 4 0 _uhttps://go.openathens.net/redirector/universidadeuropea.es?url=https://doi.org/10.1007/978-981-19-0964-1
_zAcceso a este recurso digital (usuarios Universidad Europea de Madrid)
942 _2lcc
_cLE
998 _b11/2022
_dz
_eIG
_zSI