initial release

Browse files

Files changed (8) hide show

README.md +44 -0
config.json +0 -0
pytorch_model.bin +3 -0
special_tokens_map.json +9 -0
supar.model +3 -0
tokenizer.json +0 -0
tokenizer_config.json +16 -0
vocab.txt +0 -0

README.md ADDED Viewed

	@@ -0,0 +1,44 @@

+---
+language:
+- "ko"
+tags:
+- "korean"
+- "token-classification"
+- "pos"
+- "dependency-parsing"
+datasets:
+- "universal_dependencies"
+license: "cc-by-sa-4.0"
+pipeline_tag: "token-classification"
+widget:
+- text: "홍시 맛이 나서 홍시라 생각한다."
+---
+# roberta-large-korean-morph-upos
+## Model Description
+This is a RoBERTa model for POS-tagging and dependency-parsing, derived from [klue/roberta-large](https://huggingface.co/klue/roberta-large) and [morphUD-korean](https://github.com/jungyeul/morphUD-korean). Every morpheme (형태소) is tagged by [UPOS](https://universaldependencies.org/u/pos/)(Universal Part-Of-Speech).
+## How to Use
+```py
+from transformers import AutoTokenizer,AutoModelForTokenClassification,TokenClassificationPipeline
+tokenizer=AutoTokenizer.from_pretrained("KoichiYasuoka/roberta-large-korean-morph-upos")
+model=AutoModelForTokenClassification.from_pretrained("KoichiYasuoka/roberta-large-korean-morph-upos")
+pipeline=TokenClassificationPipeline(tokenizer=tokenizer,model=model,aggregation_strategy="simple")
+nlp=lambda x:[(x[t["start"]:t["end"]],t["entity_group"]) for t in pipeline(x)]
+print(nlp("홍시 맛이 나서 홍시라 생각한다."))
+```
+or
+```py
+import esupar
+nlp=esupar.load("KoichiYasuoka/roberta-large-korean-morph-upos")
+print(nlp("홍시 맛이 나서 홍시라 생각한다."))
+```
+## See Also
+[esupar](https://github.com/KoichiYasuoka/esupar): Tokenizer POS-tagger and Dependency-parser with BERT/RoBERTa/DeBERTa models

config.json ADDED Viewed

The diff for this file is too large to render. See raw diff

pytorch_model.bin ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:eebadebe77ac215e7e3d092db1d4d302f28d15495bb423aeaa0805881094f601
+size 1346220849

special_tokens_map.json ADDED Viewed

	@@ -0,0 +1,9 @@

+{
+  "bos_token": "[CLS]",
+  "cls_token": "[CLS]",
+  "eos_token": "[SEP]",
+  "mask_token": "[MASK]",
+  "pad_token": "[PAD]",
+  "sep_token": "[SEP]",
+  "unk_token": "[UNK]"
+}

supar.model ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:b6709f1bbce1c7babe3c9e5adad4b36aaec899028f20a25e0210caf94bf229d6
+size 1394978533

tokenizer.json ADDED Viewed

The diff for this file is too large to render. See raw diff

tokenizer_config.json ADDED Viewed

	@@ -0,0 +1,16 @@

+{
+  "bos_token": "[CLS]",
+  "cls_token": "[CLS]",
+  "do_basic_tokenize": true,
+  "do_lower_case": false,
+  "eos_token": "[SEP]",
+  "mask_token": "[MASK]",
+  "model_max_length": 512,
+  "never_split": null,
+  "pad_token": "[PAD]",
+  "sep_token": "[SEP]",
+  "strip_accents": null,
+  "tokenize_chinese_chars": true,
+  "tokenizer_class": "BertTokenizerFast",
+  "unk_token": "[UNK]"
+}

vocab.txt ADDED Viewed

The diff for this file is too large to render. See raw diff