Linguakit

Kit lingüístico: dependency parser, PoS tagger, NERC, extractor multipalabra...

LINGUAKIT

Developed by ProLNat@GE Group (http://gramatica.usc.es/pln/), CiTIUS, University of Santiago de Compostela, Galiza.

LinguaKit is a Natural Language Processing tool provided with several NLP modules (constantly updated and improved):

  • Dependency parser (DepPattern)
  • PoS tagger
  • NER (named entity recognition)
  • NEC (named entity classification)
  • Coreference resolution of named entities
  • Sentiment analysis
  • Multiword extraction
  • Keyword extraction
  • Relation extraction
  • Language recognition
  • Tokenizer
  • Sentence segmentation
  • Lemmatization
  • Keyword in context
  • Entity linking and semantic annotation
  • Summarizer
  • Verb conjugator
  • Language checker (spelling, lexicon, grammar)

Description

The command linguakit is able to process 4 languages: Portuguese, English, Spanish and Galician. Since February 2018, a new variety/language has been added: historical galician-portuguese (histgz), by Xavier Canosa. The following tools are available. Scroll down for additional documentation and usage examples.

  • Dependency parser (parameter dep): Runs parsers. The parsers are implemented in PERL and stored in the parsers file. The parsers were compiled from formal grammars (more information). There are several parameters to control output: basic triplets (-a), triplets with morphological information (-fa), the same output as the input (-c) for correction purpose, and CoNLL format (-conll). These parameters are further explained in the section Dependency Parser below.
  • PoS tagger (parameter tagger): Provides the PoS tagger CitiusTools. It returns one PoS tag and one lemma per token. This is also known as PoS tagging disambiguation. The module is provided with two submodules: NER (-ner) and NEC (-nec). The NEC module returns semantic tags for named entities: NP0SP00 (Person), NP00G00 (Location), NP00O00 (Organization), NP00V00 (Miscelaneous).
  • COREF (parameter coref) labels the different named entities of the text (identified by the NER and NEC) with a numeric id which represents the discourse entity they refer to (e.g., "Bob Marley NP00SP0 (1)", "Jimi Hendrix NP00SP0 (2)", "Marley NP00SP0 (1)", "Hendrix NP00SP0 (2)", etc.). COREF allows the -crnec option (experimental) which relabels some named entities based on the results of the coreference analysis. Please note that the COREF modules may slow down the execution of the system when analyzing large texts. Also, remember that COREF is performed document by document, so it is not recommended to run it in a large corpus containing several documents.
  • Multiword extraction (parameter mwe): Extracts multiwords from PoS tagged text. There are several optional parameters, each one being a specific lexical association measure for ranking the candidate terms: chi square (-chi, default), loglikelihood (-log), mutual information (-mi), symmetrical conditional probability (-scp), simple co-occurrences (-cooc).
  • Keyword extraction (parameter key): Extracts keywords (lexemes and proper names) from PoS tagged text and ranked them using a reference corpus and chisquare.
  • Sentiment analysis (parameter sent): Returns POSITIVE, NONE (neutral) or NEGATIVE, using a polarity lexicon and a classifier trained from annotated tweets. Given an input text, this module returns a polarity value for each paragraph, namely it returns three columns for each paragraph: the text in the first column, the polarity (pos, neg, none) in the second column, and the polarity score (from 0 to 1) in the third column. The last line of the output returns the overall score computed as the average of all paragraphs.
  • Relation extraction (parameter rel): Returns triples SUBJECT - RELATION - OBJECT using methods based on Open Information Extraction.
  • Language recognition (parameter recog): Returns the language of the input text: en, es, pt, gl, gz (agal galician variety), fr, eu, ca, bn (bengali), ur (urdu), hi (hindi), ta (tamil). This module is also used by other modules to recognize the language before processing (only for the supported languages: pt, en, es, gl).
  • Tokenizer (parameter tok): Returns a tokenized text. Parameter -split splits word contractions and verb clitics. Parameter -sort ranks tokens by frequency.
  • Sentence segmentation (parameter seg): Returns a sentence per line. Sentence segmentation is the problem of dividing a string of written language into its component sentences.
  • Lemmatization (parameter lem): Returns all the lemmas of each token and all the morphological information (PoS tags) associated to each lemma. This is the process running before PoS tagging disambiguation.
  • Keyword in context (parameter kwic): Returns a target word in context (window: 10 tokens). Option -tokens returns tokens as context. This module requires the keyword to be searched as an additional argument.
  • Entity linking (parameter link): Returns a list of terms which represent Wikipedia entities. Besides, the input text is annotated with those terms and their links to Wikipedia. Requires Internet conection since it runs via Web API service. The output can be in two formats: json (default) and xml.
  • Summarizer (parameter sum): Returns an abstract of the input text. You can choose the percentage of the text to be summarized by using as option a number from 1 to 100. The code was developed by Fernando Blanco Dosil when it was working in Cilenis Language Technology.
  • Conjugator (parameter conj): Returns the verb inflection if you enter the infinitive form. Pay attention that the input is not a file but a string, the infinitive verb, and the module should be used like this: ./linguakit conj pt "fazer" -s -pb. The module is working for three languages: Galician, Spanish and Portuguese. In the case of portuguese verbs, you can choose among 4 language varieties: european portuguese after the spelling agreement (-pe), brasilian portuguese after the spelling agreement (-pb), european portuguese before the spelling agreement (-pen), brasilian portuguese before the spelling agreement (-pbn). The output is in json format. This module requires Internet conection since it runs using a Web API.
  • Language checker (parameter aval): Returns the language errors found in the input sentence. Several types of errors are considered: spelling, lexical, and grammatical issues. Suggestions of correction are provided as well as a linguistic explanation for each specific type of error. Requires Internet conection since it runs via Web API service. The output can be in two formats: json (default) and xml. By now, it is only available for Galician language (gl).

Using Make (to be installed in an accessible bin directory):

git clone  https://github.com/citiususc/Linguakit
cd Linguakit
sudo make deps
sudo make install
sudo make test-me

Thanks to José João Almeida (Univ. do Minho) for the Make file.

As all modules are been updated regularly, you'd better use git to install and update the system.

A New more efficient version of LinguaKit was released in 22 June 2017 by César Piñeiro (also for Windows: linguakit.bat command).