spaCy/spacy/tests/lang/ko
Adriane Boyd 2a558a7cdc
Switch to mecab-ko as default Korean tokenizer (#11294)
* Switch to mecab-ko as default Korean tokenizer

Switch to the (confusingly-named) mecab-ko python module for default Korean
tokenization.

Maintain the previous `natto-py` tokenizer as
`spacy.KoreanNattoTokenizer.v1`.

* Temporarily run tests with mecab-ko tokenizer

* Fix types

* Fix duplicate test names

* Update requirements test

* Revert "Temporarily run tests with mecab-ko tokenizer"

This reverts commit d2083e7044.

* Add mecab_args setting, fix pickle for KoreanNattoTokenizer

* Fix length check

* Update docs

* Formatting

* Update natto-py error message

Co-authored-by: Paul O'Leary McCann <polm@dampfkraft.com>

Co-authored-by: Paul O'Leary McCann <polm@dampfkraft.com>
2022-08-26 10:11:18 +02:00
..
__init__.py Revert #4334 2019-09-29 17:32:12 +02:00
test_lemmatization.py Switch to mecab-ko as default Korean tokenizer (#11294) 2022-08-26 10:11:18 +02:00
test_serialize.py Switch to mecab-ko as default Korean tokenizer (#11294) 2022-08-26 10:11:18 +02:00
test_tokenizer.py Switch to mecab-ko as default Korean tokenizer (#11294) 2022-08-26 10:11:18 +02:00