spaCy

mirror of https://github.com/explosion/spaCy.git synced 2026-01-27 02:33:56 +03:00

Author	SHA1	Message	Date
Adriane Boyd	5861308910	Generalize handling of tokenizer special cases Handle tokenizer special cases more generally by using the Matcher internally to match special cases after the affix/token_match tokenization is complete. Instead of only matching special cases while processing balanced or nearly balanced prefixes and suffixes, this recognizes special cases in a wider range of contexts: * Allows arbitrary numbers of prefixes/affixes around special cases * Allows special cases separated by infixes Existing tests/settings that couldn't be preserved as before: * The emoticon '")' is no longer a supported special case * The emoticon ':)' in "example:)" is a false positive again When merged with #4258 (or the relevant cache bugfix), the affix and token_match properties should be modified to flush and reload all special cases to use the updated internal tokenization with the Matcher.	2019-09-08 20:35:16 +02:00
Sofie Van Landeghem	53a9ca45c9	Docs: bufsize instead of buffsize (#4247 )	2019-09-06 11:11:54 +02:00
Sofie Van Landeghem	6b012cebff	Make pos/tag distinction more clear in docs (#4246 ) * make distinction between tag and pos more prominent in docs * out of the 101	2019-09-06 10:31:21 +02:00
Bae Yong-Ju	a55f5a744f	Fix ValueError exception on empty Korean text. (#4245 )	2019-09-06 10:29:40 +02:00
Ines Montani	232a029de6	Send referrer for internal links [ci skip]	2019-09-05 10:41:46 +02:00
Matthew Honnibal	b94c34ec8f	Merge pull request #4239 from adrianeboyd/bugfix/tokenizer-cache-test-1061 Add regression test for #1061 back to test suite	2019-09-04 23:10:12 +02:00
Adriane Boyd	0f28418446	Add regression test for #1061 back to test suite	2019-09-04 20:42:24 +02:00
Ines Montani	2f31f96fce	Update languages.json [ci skip]	2019-09-04 18:15:42 +02:00
Ines Montani	2245e95e2d	Update languages.json [ci skip]	2019-09-04 17:11:40 +02:00
Matthew Honnibal	17c039406b	Merge pull request #4232 from adrianeboyd/bugfix/entityruler-ner-4229 Fix handling of preset entities in NER	2019-09-04 15:02:31 +02:00
Adriane Boyd	6b0fec76fd	Fix handling of preset entities in NER * Fix check of valid ent_type for B * Add valid L as preset-I followed by not-I	2019-09-04 13:42:42 +02:00
Ines Montani	419ae59c79	Make flaky test test_issue_1971_4 more explicit	2019-08-31 14:08:05 +02:00
Ines Montani	dad5621166	Tidy up and auto-format [ci skip]	2019-08-31 13:39:31 +02:00
Ines Montani	cd90752193	Tidy up and auto-format [ci skip]	2019-08-31 13:39:06 +02:00
Ines Montani	bcd1b12f43	Add contributor agreement [ci skip]	2019-08-30 17:02:43 +02:00
Matthew Honnibal	516650f58f	Merge pull request #4207 from svlandeg/bugfix/serialize-tok-exc Bugfix for serializing tokenizer rules/exceptions	2019-08-30 11:04:58 +02:00
adrianeboyd	82159b5c19	Updates/bugfixes for NER/IOB converters (#4186 ) * Updates/bugfixes for NER/IOB converters * Converter formats `ner` and `iob` use autodetect to choose a converter if possible * `iob2json` is reverted to handle sentence-per-line data like `word1\|pos1\|ent1 word2\|pos2\|ent2` * Fix bug in `merge_sentences()` so the second sentence in each batch isn't skipped * `conll_ner2json` is made more general so it can handle more formats with whitespace-separated columns * Supports all formats where the first column is the token and the final column is the IOB tag; if present, the second column is the POS tag * As in CoNLL 2003 NER, blank lines separate sentences, `-DOCSTART- -X- O O` separates documents * Add option for segmenting sentences (new flag `-s`) * Parser-based sentence segmentation with a provided model, otherwise with sentencizer (new option `-b` to specify model) * Can group sentences into documents with `n_sents` as long as sentence segmentation is available * Only applies automatic segmentation when there are no existing delimiters in the data * Provide info about settings applied during conversion with warnings and suggestions if settings conflict or might not be not optimal. * Add tests for common formats * Add '(default)' back to docs for -c auto * Add document count back to output * Revert changes to converter output message * Use explicit tabs in convert CLI test data * Adjust/add messages for n_sents=1 default * Add sample NER data to training examples * Update README * Add links in docs to example NER data * Define msg within converters	2019-08-29 12:04:01 +02:00
adrianeboyd	5feb342f5e	Add more token attributes to token pattern schema (#4210 ) Add token attributes with tests to token pattern schema.	2019-08-29 12:02:26 +02:00
svlandeg	c54aabc3cd	fix loading custom tokenizer rules/exceptions from file	2019-08-28 14:17:44 +02:00
svlandeg	7bec0ebbcb	failing unit test for Issue 4190	2019-08-28 14:16:34 +02:00
Ines Montani	b91425f803	Update universe.json [ci skip]	2019-08-28 13:45:06 +02:00
Ines Montani	aedae8b4c5	Update universe.json [ci skip]	2019-08-28 11:59:06 +02:00
Björn Böing	bae0455f91	Fix visualizer options linking for displaCy. (#4202 )	2019-08-27 14:04:28 +02:00
Ines Montani	8114933f01	Fix universe.json [ci skip]	2019-08-27 12:13:42 +02:00
Ines Montani	48385552c6	Update languages.json [ci skip]	2019-08-27 11:52:51 +02:00
Ines Montani	f4012ba054	Update README.md [ci skip]	2019-08-26 12:32:52 +02:00
yanaiela	5d7bc26735	new universe project - the numeric fused-head (#4192 ) * new universe project * Update website/meta/universe.json Co-Authored-By: Ines Montani <ines@ines.io> * Update website/meta/universe.json Co-Authored-By: Ines Montani <ines@ines.io>	2019-08-25 17:25:28 +02:00
Wannaphong Phatthiyaphaibun	d53c3fcbc1	Add Thai Language tokenizers (#4191 ) Add th (pythainlp)	2019-08-25 11:35:21 +02:00
Christos Aridas	61f5c007a0	DOC Fix pipeline functions examples (#4189 )	2019-08-23 19:15:32 +02:00
Matthew Honnibal	bb911e5f4e	Fix #3830 : 'subtok' label being added even if learn_tokens=False (#4188 ) * Prevent subtok label if not learning tokens The parser introduces the subtok label to mark tokens that should be merged during post-processing. Previously this happened even if we did not have the --learn-tokens flag set. This patch passes the config through to the parser, to prevent the problem. * Make merge_subtokens a parser post-process if learn_subtokens * Fix train script * Add test for 3830: subtok problem * Fix handlign of non-subtok in parser training	2019-08-23 17:54:00 +02:00
Sofie Van Landeghem	c417c380e3	Matcher ID fixes (#4179 ) * allow phrasematcher to link one match to multiple original patterns * small fix for defining ent_id in the matcher (anti-ghost prevention) * cleanup * formatting	2019-08-22 17:17:07 +02:00
Ines Montani	f5d3afb1a3	Fix typo in docstrings [ci skip]	2019-08-22 16:24:15 +02:00
Ines Montani	5ca7dd0f94	💫 WIP: Basic lookup class scaffolding and JSON for all lemmati… (#4167 ) * Improve load_language_data helper * WIP: Add Lookups implementation * Start moving lemma data over to JSON * WIP: move data over for more languages * Convert more languages * Fix lemmatizer fixtures in tests * Finish conversion * Auto-format JSON files * Fix test for now * Make sure tables are stored on instance	2019-08-22 14:21:32 +02:00
Sofie Van Landeghem	73b38c33e4	Small retokenizer fix (#4174 )	2019-08-22 12:23:54 +02:00
Ines Montani	a8752a569d	Auto-format [ci skip]	2019-08-22 11:44:39 +02:00
Pavle Vidanović	60e10a9f93	Serbian language improvement (#4169 ) * Serbian stopwords added. (cyrillic alphabet) * spaCy Contribution agreement included. * Test initialize updated * Serbian language code update. --bugfix * Tokenizer exceptions added. Init file updated. * Norm exceptions and lexical attributes added. * Examples added. * Tests added. * sr_lang examples update. * Tokenizer exceptions updated. (Serbian)	2019-08-22 11:43:07 +02:00
Sofie Van Landeghem	de272f8b82	adding double match for optional operator at the end (#4166 )	2019-08-21 22:46:56 +02:00
Sofie Van Landeghem	01c5980187	Serialize POS attribute when doc.is_tagged (#4092 ) * fix and unit test for issue 3959 * additional unit test for manifestation of the same (resolved) bug	2019-08-21 21:59:30 +02:00
Sofie Van Landeghem	7539a4f3a8	use states[q] in while retry loop (#4162 )	2019-08-21 21:58:04 +02:00
Ines Montani	b072c13017	Update universe with videos [ci skip]	2019-08-21 21:35:37 +02:00
adrianeboyd	2d17b047e2	Check for is_tagged/is_parsed for Matcher attrs (#4163 ) Check for relevant components in the pipeline when Matcher is called, similar to the checks for PhraseMatcher in #4105. * keep track of attributes seen in patterns * when Matcher is called on a Doc, check for is_tagged for LEMMA, TAG, POS and for is_parsed for DEP	2019-08-21 20:52:36 +02:00
Pavle Vidanović	4fe9329bfb	Serbian language code update "rs" -> "sr" (#4159 ) * Serbian stopwords added. (cyrillic alphabet) * spaCy Contribution agreement included. * Test initialize updated * Serbian language code update. --bugfix	2019-08-21 19:57:37 +02:00
adrianeboyd	8fe7bdd0fa	Improve token pattern checking without validation (#4105 ) * Fix typo in rule-based matching docs * Improve token pattern checking without validation Add more detailed token pattern checks without full JSON pattern validation and provide more detailed error messages. Addresses #4070 (also related: #4063, #4100). * Check whether top-level attributes in patterns and attr for PhraseMatcher are in token pattern schema * Check whether attribute value types are supported in general (as opposed to per attribute with full validation) * Report various internal error types (OverflowError, AttributeError, KeyError) as ValueError with standard error messages * Check for tagger/parser in PhraseMatcher pipeline for attributes TAG, POS, LEMMA, and DEP * Add error messages with relevant details on how to use validate=True or nlp() instead of nlp.make_doc() * Support attr=TEXT for PhraseMatcher * Add NORM to schema * Expand tests for pattern validation, Matcher, PhraseMatcher, and EntityRuler * Remove unnecessary .keys() * Rephrase error messages * Add another type check to Matcher Add another type check to Matcher for more understandable error messages in some rare cases. * Support phrase_matcher_attr=TEXT for EntityRuler * Don't use spacy.errors in examples and bin scripts * Fix error code * Auto-format Also try get Azure pipelines to finally start a build :( * Update errors.py Co-authored-by: Ines Montani <ines@ines.io> Co-authored-by: Matthew Honnibal <honnibal+gh@gmail.com>	2019-08-21 14:00:37 +02:00
Ines Montani	3134a9b6e0	Add section on expanding regex match to token boundaries (see #4158 ) [ci skip]	2019-08-21 12:53:31 +02:00
Ines Montani	f580302673	Tidy up and auto-format	2019-08-20 17:36:34 +02:00
Ines Montani	364aaf5bc2	Simplify test	2019-08-20 16:41:58 +02:00
Sofie Van Landeghem	68ee0384fd	Unit test for Issue 3879 (#4153 ) * failing unit test for Issue #3879 * mark test as failing	2019-08-20 16:40:25 +02:00
Ines Montani	86cd7f0efd	Add regression test for #4120	2019-08-20 16:33:09 +02:00
Ines Montani	104125edd2	Tidy up errors	2019-08-20 16:03:45 +02:00
Ines Montani	cc76a26fe8	Raise error for negative arc indices (closes #3917 )	2019-08-20 15:51:37 +02:00

1 2 3 4 5 ...

10500 Commits