spaCy/spacy/lang/it/__init__.py

from typing import Optional
from thinc.api import Model

from .stop_words import STOP_WORDS
from .tokenizer_exceptions import TOKENIZER_EXCEPTIONS
from .punctuation import TOKENIZER_PREFIXES, TOKENIZER_INFIXES
from ...language import Language
from .lemmatizer import ItalianLemmatizer


class ItalianDefaults(Language.Defaults):
    tokenizer_exceptions = TOKENIZER_EXCEPTIONS
    stop_words = STOP_WORDS
    prefixes = TOKENIZER_PREFIXES
    infixes = TOKENIZER_INFIXES


class Italian(Language):
    lang = "it"
    Defaults = ItalianDefaults


@Italian.factory(
    "lemmatizer",
    assigns=["token.lemma"],
    default_config={"model": None, "mode": "pos_lookup", "overwrite": False},
    default_score_weights={"lemma_acc": 1.0},
)
def make_lemmatizer(
    nlp: Language, model: Optional[Model], name: str, mode: str, overwrite: bool
):
    return ItalianLemmatizer(nlp.vocab, model, name, mode=mode, overwrite=overwrite)


__all__ = ["Italian"]
Added Italian POS-aware lemmatizer. (#8079) * Added Italian POS-aware lemmatizer. Also added the code used to build the lookup tables by POS. * Create gtoffoli.md * Add imports and format * Remove helper script * Use lemma_lookup instead of lemma_lookup_legacy Co-authored-by: Adriane Boyd <adrianeboyd@gmail.com> 2021-06-16 12:14:45 +03:00			`from typing import Optional`
			`from thinc.api import Model`

Reorganise Italian language data 2017-05-08 16:50:17 +03:00			`from .stop_words import STOP_WORDS`
Clean up of char classes, few tokenizer fixes and faster default French tokenizer (#3293) * splitting up latin unicode interval * removing hyphen as infix for French * adding failing test for issue 1235 * test for issue #3002 which now works * partial fix for issue #2070 * keep the hyphen as infix for French (as it was) * restore french expressions with hyphen as infix (as it was) * added succeeding unit test for Issue #2656 * Fix issue #2822 with custom Italian exception * Fix issue #2926 by allowing numbers right before infix / * splitting up latin unicode interval * removing hyphen as infix for French * adding failing test for issue 1235 * test for issue #3002 which now works * partial fix for issue #2070 * keep the hyphen as infix for French (as it was) * restore french expressions with hyphen as infix (as it was) * added succeeding unit test for Issue #2656 * Fix issue #2822 with custom Italian exception * Fix issue #2926 by allowing numbers right before infix / * remove duplicate * remove xfail for Issue #2179 fixed by Matt * adjust documentation and remove reference to regex lib 2019-02-21 00:10:13 +03:00			`from .tokenizer_exceptions import TOKENIZER_EXCEPTIONS`
Improve Italian tokenization (#5204) Improve Italian tokenization for UD_Italian-ISDT. 2020-03-25 13:28:02 +03:00			`from .punctuation import TOKENIZER_PREFIXES, TOKENIZER_INFIXES`
Fix relative imports 2017-05-08 23:29:04 +03:00			`from ...language import Language`
Added Italian POS-aware lemmatizer. (#8079) * Added Italian POS-aware lemmatizer. Also added the code used to build the lookup tables by POS. * Create gtoffoli.md * Add imports and format * Remove helper script * Use lemma_lookup instead of lemma_lookup_legacy Co-authored-by: Adriane Boyd <adrianeboyd@gmail.com> 2021-06-16 12:14:45 +03:00			`from .lemmatizer import ItalianLemmatizer`
Improve Italian & Urdu tokenization accuracy (#3228) ## Description 1. Added the same infix rule as in French (`d'une`, `j'ai`) for Italian (`c'è`, `l'ha`), bringing F-score on `it_isdt-ud-train.txt` from 96% to 99%. Added unit test to check this behaviour. 2. Added specific Urdu punctuation character as suffix, improving F-score on `ur_udtb-ud-train.txt` from 94% to 100%. Added unit test to check this behaviour. ### Types of change Enhancement of Italian & Urdu tokenization ## Checklist - [x] I have submitted the spaCy Contributor Agreement. - [x] I ran the tests, and all new and existing tests passed. - [x] My changes don't require a change to the documentation, or if they do, I've added all required information. 2019-02-05 00:39:25 +03:00
* Add spacy.it 2015-09-06 23:10:37 +03:00
Move Defaults subclass to module scope (necessary for pickling) 2017-05-20 20:02:27 +03:00			`class ItalianDefaults(Language.Defaults):`
Tidy up and move noun_chunks, token_match, url_match 2020-07-22 23:18:46 +03:00			`tokenizer_exceptions = TOKENIZER_EXCEPTIONS`
Don't make copies of language data components 2017-10-11 16:34:55 +03:00			`stop_words = STOP_WORDS`
Improve Italian tokenization (#5204) Improve Italian tokenization for UD_Italian-ISDT. 2020-03-25 13:28:02 +03:00			`prefixes = TOKENIZER_PREFIXES`
Improve Italian & Urdu tokenization accuracy (#3228) ## Description 1. Added the same infix rule as in French (`d'une`, `j'ai`) for Italian (`c'è`, `l'ha`), bringing F-score on `it_isdt-ud-train.txt` from 96% to 99%. Added unit test to check this behaviour. 2. Added specific Urdu punctuation character as suffix, improving F-score on `ur_udtb-ud-train.txt` from 94% to 100%. Added unit test to check this behaviour. ### Types of change Enhancement of Italian & Urdu tokenization ## Checklist - [x] I have submitted the spaCy Contributor Agreement. - [x] I ran the tests, and all new and existing tests passed. - [x] My changes don't require a change to the documentation, or if they do, I've added all required information. 2019-02-05 00:39:25 +03:00			`infixes = TOKENIZER_INFIXES`
Work on draft Italian tokenizer 2016-11-02 21:56:32 +03:00
Lazy imports language 2017-05-03 12:01:42 +03:00
Move Defaults subclass to module scope (necessary for pickling) 2017-05-20 20:02:27 +03:00			`class Italian(Language):`
💫 Tidy up and auto-format .py files (#2983) <!--- Provide a general summary of your changes in the title. --> ## Description - [x] Use [`black`](https://github.com/ambv/black) to auto-format all `.py` files. - [x] Update flake8 config to exclude very large files (lemmatization tables etc.) - [x] Update code to be compatible with flake8 rules - [x] Fix various small bugs, inconsistencies and messy stuff in the language data - [x] Update docs to explain new code style (`black`, `flake8`, when to use `# fmt: off` and `# fmt: on` and what `# noqa` means) Once #2932 is merged, which auto-formats and tidies up the CLI, we'll be able to run `flake8 spacy` actually get meaningful results. At the moment, the code style and linting isn't applied automatically, but I'm hoping that the new [GitHub Actions](https://github.com/features/actions) will let us auto-format pull requests and post comments with relevant linting information. ### Types of change enhancement, code style ## Checklist <!--- Before you submit the PR, go over this checklist and make sure you can tick off all the boxes. [] -> [x] --> - [x] I have submitted the spaCy Contributor Agreement. - [x] I ran the tests, and all new and existing tests passed. - [x] My changes don't require a change to the documentation, or if they do, I've added all required information. 2018-11-30 19:03:03 +03:00			`lang = "it"`
Move Defaults subclass to module scope (necessary for pickling) 2017-05-20 20:02:27 +03:00			`Defaults = ItalianDefaults`
Adding method lemmatizer for every class 2017-05-03 13:14:42 +03:00
Lazy imports language 2017-05-03 12:01:42 +03:00
Added Italian POS-aware lemmatizer. (#8079) * Added Italian POS-aware lemmatizer. Also added the code used to build the lookup tables by POS. * Create gtoffoli.md * Add imports and format * Remove helper script * Use lemma_lookup instead of lemma_lookup_legacy Co-authored-by: Adriane Boyd <adrianeboyd@gmail.com> 2021-06-16 12:14:45 +03:00			`@Italian.factory(`
			`"lemmatizer",`
			`assigns=["token.lemma"],`
			`default_config={"model": None, "mode": "pos_lookup", "overwrite": False},`
			`default_score_weights={"lemma_acc": 1.0},`
			`)`
			`def make_lemmatizer(`
			`nlp: Language, model: Optional[Model], name: str, mode: str, overwrite: bool`
			`):`
			`return ItalianLemmatizer(nlp.vocab, model, name, mode=mode, overwrite=overwrite)`


💫 Tidy up and auto-format .py files (#2983) <!--- Provide a general summary of your changes in the title. --> ## Description - [x] Use [`black`](https://github.com/ambv/black) to auto-format all `.py` files. - [x] Update flake8 config to exclude very large files (lemmatization tables etc.) - [x] Update code to be compatible with flake8 rules - [x] Fix various small bugs, inconsistencies and messy stuff in the language data - [x] Update docs to explain new code style (`black`, `flake8`, when to use `# fmt: off` and `# fmt: on` and what `# noqa` means) Once #2932 is merged, which auto-formats and tidies up the CLI, we'll be able to run `flake8 spacy` actually get meaningful results. At the moment, the code style and linting isn't applied automatically, but I'm hoping that the new [GitHub Actions](https://github.com/features/actions) will let us auto-format pull requests and post comments with relevant linting information. ### Types of change enhancement, code style ## Checklist <!--- Before you submit the PR, go over this checklist and make sure you can tick off all the boxes. [] -> [x] --> - [x] I have submitted the spaCy Contributor Agreement. - [x] I ran the tests, and all new and existing tests passed. - [x] My changes don't require a change to the documentation, or if they do, I've added all required information. 2018-11-30 19:03:03 +03:00			`__all__ = ["Italian"]`