{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2023,11,17]],"date-time":"2023-11-17T00:28:42Z","timestamp":1700180922514},"reference-count":21,"publisher":"Wiley","issue":"9","license":[{"start":{"date-parts":[[2005,6,10]],"date-time":"2005-06-10T00:00:00Z","timestamp":1118361600000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/onlinelibrary.wiley.com\/termsAndConditions#vor"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Systems &amp; Computers in Japan"],"published-print":{"date-parts":[[2005,8]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>We present unsupervised speaker indexing, combined with automatic speech recognition (ASR) for speech archives, such as discussions. Our proposed indexing method is based on anchor models, by which we define a feature vector based on the similarity with speakers of a large\u2010scale speech database. We introduce dimensional normalization and reduction on the vectors to improve discriminant ability. These vectors are then clustered and initial speaker labels are obtained. Using the initial labels, speaker models are constructed for respective clusters and the speakers are finally indexed with the speaker models. We perform ASR using the results of this indexing. We achieved a speaker indexing accuracy of 97% and a significant improvement in the ASR for real discussion data. \u00a9 2005 Wiley Periodicals, Inc. Syst Comp Jpn, 36(9): 25\u201333, 2005; Published online in Wiley InterScience (<jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"http:\/\/www.interscience.wiley.com\">www.interscience.wiley.com<\/jats:ext-link>). DOI 10.1002\/scj.20215<\/jats:p>","DOI":"10.1002\/scj.20215","type":"journal-article","created":{"date-parts":[[2005,6,10]],"date-time":"2005-06-10T12:48:31Z","timestamp":1118407711000},"page":"25-33","source":"Crossref","is-referenced-by-count":0,"title":["Unsupervised speaker indexing of discussions using anchor models"],"prefix":"10.1002","volume":"36","author":[{"given":"Yuya","family":"Akita","sequence":"first","affiliation":[]},{"given":"Tatsuya","family":"Kawahara","sequence":"additional","affiliation":[]}],"member":"311","published-online":{"date-parts":[[2005,6,10]]},"reference":[{"key":"e_1_2_1_2_2","unstructured":"AkitaY KawaharaT.Automatic archiving system for meeting speech. IPSJ SIG Tech Rep SLP\u201034\u201011 2000. (in Japanese)"},{"key":"e_1_2_1_3_2","first-page":"601","article-title":"Advances in automatic meeting record creation and access","volume":"1","author":"Metze F","year":"2001","journal-title":"Proc ICASSP"},{"key":"e_1_2_1_4_2","doi-asserted-by":"crossref","first-page":"491","DOI":"10.21437\/Eurospeech.2001-129","article-title":"Towards automatic transcription of spontaneous presentations","volume":"1","author":"Shinozaki T","year":"2001","journal-title":"Proc EUROSPEECH"},{"key":"e_1_2_1_5_2","first-page":"725","article-title":"Speaking\u2010rate dependent decoding and adaptation for spontaneous lecture speech recognition","volume":"1","author":"Nanjo H","year":"2002","journal-title":"Proc ICASSP"},{"key":"e_1_2_1_6_2","unstructured":"BettM GrossR YuH ZhuX PanY YangJ WaibelA.Multimodal meeting tracker. Proc RIAO p1\u201314 2000."},{"key":"e_1_2_1_7_2","first-page":"121","article-title":"Speaker indexing for retrieval of voicemail messages","volume":"1","author":"Charlet D","year":"2002","journal-title":"Proc ICASSP"},{"key":"e_1_2_1_8_2","unstructured":"MimuraM KawaharaT DoshitaS.Automatic indexing of speakers and topics for panel discussion speech. IPSJ SIG Tech Rep SLP\u201011\u20103 1996. (in Japanese)"},{"key":"e_1_2_1_9_2","first-page":"1347","article-title":"Real time speaker indexing based on subspace method\u2014Application to TV News Articles and Debate","volume":"4","author":"Nishida M","year":"1998","journal-title":"Proc ICSLP"},{"key":"e_1_2_1_10_2","first-page":"961","article-title":"On\u2010line incremental speaker adaptation with automatic speaker change detection","volume":"2","author":"Zhang Z\u2010P","year":"2000","journal-title":"Proc ICASSP"},{"key":"e_1_2_1_11_2","first-page":"413","article-title":"Speaker change detection and speaker clustering using VQ distortion for broadcast news speech recognition","volume":"1","author":"Mori K","year":"2001","journal-title":"Proc ICASSP"},{"key":"e_1_2_1_12_2","unstructured":"TagumaR IwanoK FuruiS.Parallel computing\u2010based meeting speech recognition system with incremental on\u2010line speaker adaptation. Proc Spring Meeting of the Acoustical Society of Japan 2\u20105\u201016 2002. (in Japanese)"},{"key":"e_1_2_1_13_2","first-page":"2407","article-title":"Unknown\u2010multiple signal source clustering problem using ergodic HMM and applied to speaker classification","volume":"4","author":"Murakami J","year":"1996","journal-title":"Proc ICSLP"},{"key":"e_1_2_1_14_2","first-page":"573","article-title":"Unknown\u2010multiple speaker clustering using HMM","volume":"1","author":"Ajmera J","year":"2002","journal-title":"Proc ICSLP"},{"key":"e_1_2_1_15_2","first-page":"757","article-title":"Clustering speakers by their voices","volume":"2","author":"Solomonoff A","year":"1998","journal-title":"Proc ICASSP"},{"key":"e_1_2_1_16_2","first-page":"429","article-title":"Speaker indexing in large audio databases using anchor models","volume":"1","author":"Sturim D","year":"2001","journal-title":"Proc ICASSP"},{"key":"e_1_2_1_17_2","first-page":"1333","article-title":"Speaker identification by location in an optimal space of anchor models","volume":"2","author":"Mami Y","year":"2002","journal-title":"Proc ICSLP"},{"key":"e_1_2_1_18_2","volume-title":"Speech recognition system","author":"Shikano K","year":"2001"},{"key":"e_1_2_1_19_2","volume-title":"Pattern information processing","author":"Nakagawa S","year":"1999"},{"key":"e_1_2_1_20_2","doi-asserted-by":"publisher","DOI":"10.1006\/dspr.1999.0361"},{"key":"e_1_2_1_21_2","unstructured":"AkitaY KawaharaT.Adaptation of language and acoustic models for automatic speech recognition of discussions. Proc Spring Meeting of the Acoustical Society of Japan 2\u20104\u20103 2003. (in Japanese)"},{"key":"e_1_2_1_22_2","doi-asserted-by":"crossref","first-page":"1691","DOI":"10.21437\/Eurospeech.2001-396","article-title":"Julius\u2014an open source real\u2010time large vocabulary recognition engine","volume":"3","author":"Lee A","year":"2001","journal-title":"Proc EUROSPEECH"}],"container-title":["Systems and Computers in Japan"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/api.wiley.com\/onlinelibrary\/tdm\/v1\/articles\/10.1002%2Fscj.20215","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1002\/scj.20215","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,11,16]],"date-time":"2023-11-16T22:58:48Z","timestamp":1700175528000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1002\/scj.20215"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2005,6,10]]},"references-count":21,"journal-issue":{"issue":"9","published-print":{"date-parts":[[2005,8]]}},"alternative-id":["10.1002\/scj.20215"],"URL":"https:\/\/doi.org\/10.1002\/scj.20215","archive":["Portico"],"relation":{},"ISSN":["0882-1666","1520-684X"],"issn-type":[{"value":"0882-1666","type":"print"},{"value":"1520-684X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2005,6,10]]}}}