Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ https://zenodo.org/r...arrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
https://zenodo.org/record/1147...
Part of book or chapter of book
License: CC BY
Data sources: UnpayWall
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Part of book or chapter of book . 2006
Data sources: ZENODO
https://doi.org/10.1007/118748...
Part of book or chapter of book . 2006 . Peer-reviewed
Data sources: Crossref
https://dx.doi.org/10.60692/fa...
Other literature type . 2006
Data sources: Datacite
https://dx.doi.org/10.60692/kh...
Other literature type . 2006
Data sources: Datacite
DBLP
Conference object
Data sources: DBLP
versions View all 5 versions
addClaim

Comparing Two Markov Methods for Part-of-Speech Tagging of Portuguese

مقارنة طريقتين لماركوف لوضع علامات على جزء من الكلام باللغة البرتغالية
Authors: Fábio Kepler; Marcelo Finger;

Comparing Two Markov Methods for Part-of-Speech Tagging of Portuguese

Abstract

Il existe une grande variété de méthodes statistiques appliquées au marquage Part-of-Speech (PoS), qui associent des mots dans un texte à leur PoS correspondant. La majorité de ces méthodes analysent un petit voisinage fixe de mots imposant une certaine forme de restriction de Markov. Dans ce travail, nous mettons en œuvre et comparons un modèle de Markov caché (HMM) de longueur fixe avec une chaîne de Markov de longueur variable (VLMC) ; cette dernière est, en principe, capable de détecter les dépendances longue distance. Nous montrons que le modèle VLMC fonctionne mieux en termes de précision et presque également en termes de temps de marquage, et qu'il fonctionne également très bien en temps de formation. Cependant, la méthode VLMC ne parvient pas à capturer les dépendances très longues distances, et nous analysons les raisons de ce comportement.

Existe una amplia variedad de métodos estadísticos aplicados al etiquetado de Part-of-Speech (PoS), que asocian palabras en un texto a su correspondiente PoS. La mayoría de esos métodos analizan un vecindario fijo y pequeño de palabras que imponen alguna forma de restricción de Markov. En este trabajo implementamos y comparamos un modelo oculto de Markov de longitud fija (HMM) con una cadena de Markov de longitud variable (VLMC); esta última es, en principio, capaz de detectar dependencias de larga distancia. Mostramos que el modelo VLMC se desempeña mejor en términos de precisión y casi por igual en términos de tiempo de etiquetado, también lo hace muy bien en el tiempo de entrenamiento. Sin embargo, el método VLMC en realidad no logra capturar dependencias realmente de larga distancia, y analizamos las razones de dicho comportamiento.

There is a wide variety of statistical methods applied to Part-of-Speech (PoS) tagging, that associate words in a text to their corresponding PoS. The majority of those methods analyse a fixed, small neighborhood of words imposing some form of Markov restriction. In this work we implement and compare a fixed length hidden Markov model (HMM) with a variable length Markov chain (VLMC); the latter is, in principle, capable of detecting long distance dependencies. We show that the VLMC model performs better in terms of accuracy and almost equally in terms of tagging time, also doing very well in training time. However, the VLMC method actually fails to capture really long distance dependencies, and we analyse the reasons for such behaviour.

هناك مجموعة واسعة من الأساليب الإحصائية المطبقة على وسم جزء من الكلام (PoS)، والتي تربط الكلمات في النص بنقاط البيع المقابلة لها. تحلل غالبية هذه الأساليب مجموعة صغيرة وثابتة من الكلمات التي تفرض شكلاً من أشكال تقييد ماركوف. في هذا العمل، نقوم بتنفيذ ومقارنة نموذج ماركوف مخفي بطول ثابت (HMM) مع سلسلة ماركوف متغيرة الطول (VLMC) ؛ هذا الأخير، من حيث المبدأ، قادر على اكتشاف تبعيات المسافات الطويلة. نظهر أن نموذج VLMC يحقق أداءً أفضل من حيث الدقة وبشكل متساوٍ تقريبًا من حيث وقت وضع العلامات، كما أنه يحقق أداءً جيدًا جدًا في وقت التدريب. ومع ذلك، فإن طريقة VLMC تفشل في الواقع في التقاط التبعيات لمسافات طويلة حقًا، ونحن نحلل أسباب هذا السلوك.

Keywords

Artificial intelligence, Markov chain, Variety (cybernetics), Speech recognition, Pattern recognition (psychology), Mathematical analysis, Artificial Intelligence, Part-of-Speech Tagging, Part-of-speech tagging, Machine learning, FOS: Mathematics, Variable (mathematics), Hidden semi-Markov model, Hidden Markov Models, Natural Language Processing, Hidden Markov model, Variable-order Markov model, Portuguese, Natural language processing, Linguistics, Statistical Machine Translation and Natural Language Processing, Computer science, Markov model, Language Modeling, FOS: Philosophy, ethics and religion, Speech Recognition Technology, Maximum-entropy Markov model, Philosophy, Part of speech, Dependency Parsing, Computer Science, Physical Sciences, FOS: Languages and literature, Statistical Language Modeling, Mathematics

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    4
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Top 10%
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
    OpenAIRE UsageCounts
    Usage byUsageCounts
    visibility views 4
    download downloads 15
  • 4
    views
    15
    downloads
    Powered byOpenAIRE UsageCounts
Powered by OpenAIRE graph
Found an issue? Give us feedback
visibility
download
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
views
OpenAIRE UsageCountsViews provided by UsageCounts
downloads
OpenAIRE UsageCountsDownloads provided by UsageCounts
4
Average
Top 10%
Average
4
15
Green
Related to Research communities