spaCy/spacy/lang/tl/tokenizer_exceptions.py

from ...symbols import NORM, ORTH
from ...util import update_exc
from ..tokenizer_exceptions import BASE_EXCEPTIONS

_exc = {
    "tayo'y": [{ORTH: "tayo"}, {ORTH: "'y", NORM: "ay"}],
    "isa'y": [{ORTH: "isa"}, {ORTH: "'y", NORM: "ay"}],
    "baya'y": [{ORTH: "baya"}, {ORTH: "'y", NORM: "ay"}],
    "sa'yo": [{ORTH: "sa"}, {ORTH: "'yo", NORM: "iyo"}],
    "ano'ng": [{ORTH: "ano"}, {ORTH: "'ng", NORM: "ang"}],
    "siya'y": [{ORTH: "siya"}, {ORTH: "'y", NORM: "ay"}],
    "nawa'y": [{ORTH: "nawa"}, {ORTH: "'y", NORM: "ay"}],
    "papa'no": [{ORTH: "papa'no", NORM: "papaano"}],
    "'di": [{ORTH: "'di", NORM: "hindi"}],
}


TOKENIZER_EXCEPTIONS = update_exc(BASE_EXCEPTIONS, _exc)
Configure isort to use the Black profile, recursively isort the `spacy` module (#12721) * Use isort with Black profile * isort all the things * Fix import cycles as a result of import sorting * Add DOCBIN_ALL_ATTRS type definition * Add isort to requirements * Remove isort from build dependencies check * Typo 2023-06-14 18:48:41 +03:00			`from ...symbols import NORM, ORTH`
Tidy up and move noun_chunks, token_match, url_match 2020-07-22 23:18:46 +03:00			`from ...util import update_exc`
Configure isort to use the Black profile, recursively isort the `spacy` module (#12721) * Use isort with Black profile * isort all the things * Fix import cycles as a result of import sorting * Add DOCBIN_ALL_ATTRS type definition * Add isort to requirements * Remove isort from build dependencies check * Typo 2023-06-14 18:48:41 +03:00			`from ..tokenizer_exceptions import BASE_EXCEPTIONS`
Added alpha support for Tagalog language (#3062) I have added alpha support for the Tagalog language from the Philippines. It is the basis for the country's national language Filipino. I have heavily based the format to the EN and ES languages. I have provided several words in the lemmatizer lookup table, added stop words from a source, translated numeric words to its Tagalog counterpart, added some tokenizer exceptions, and kept the tag map the same as the English language. While the alpha language passed the preliminary testing that you provided, I think it needs more data to be useful for most cases. * Added alpha support for Tagalog language * Edited contributor template * Included SCA; Reverted templates * Fixed SCA template * Fixed changes in SCA template 2018-12-18 15:08:38 +03:00
			`_exc = {`
Remove POS, TAG and LEMMA from tokenizer exceptions 2020-07-23 00:09:01 +03:00			`"tayo'y": [{ORTH: "tayo"}, {ORTH: "'y", NORM: "ay"}],`
			`"isa'y": [{ORTH: "isa"}, {ORTH: "'y", NORM: "ay"}],`
			`"baya'y": [{ORTH: "baya"}, {ORTH: "'y", NORM: "ay"}],`
			`"sa'yo": [{ORTH: "sa"}, {ORTH: "'yo", NORM: "iyo"}],`
			`"ano'ng": [{ORTH: "ano"}, {ORTH: "'ng", NORM: "ang"}],`
			`"siya'y": [{ORTH: "siya"}, {ORTH: "'y", NORM: "ay"}],`
			`"nawa'y": [{ORTH: "nawa"}, {ORTH: "'y", NORM: "ay"}],`
			`"papa'no": [{ORTH: "papa'no", NORM: "papaano"}],`
			`"'di": [{ORTH: "'di", NORM: "hindi"}],`
Added alpha support for Tagalog language (#3062) I have added alpha support for the Tagalog language from the Philippines. It is the basis for the country's national language Filipino. I have heavily based the format to the EN and ES languages. I have provided several words in the lemmatizer lookup table, added stop words from a source, translated numeric words to its Tagalog counterpart, added some tokenizer exceptions, and kept the tag map the same as the English language. While the alpha language passed the preliminary testing that you provided, I think it needs more data to be useful for most cases. * Added alpha support for Tagalog language * Edited contributor template * Included SCA; Reverted templates * Fixed SCA template * Fixed changes in SCA template 2018-12-18 15:08:38 +03:00			`}`


Tidy up and move noun_chunks, token_match, url_match 2020-07-22 23:18:46 +03:00			`TOKENIZER_EXCEPTIONS = update_exc(BASE_EXCEPTIONS, _exc)`