spaCy

mirror of https://github.com/explosion/spaCy.git synced 2025-09-14 16:12:39 +03:00

Author	SHA1	Message	Date
Matthew Honnibal	449b889454	Fix KeyError in Vectors.most_similar. Fixes #2648	2018-12-10 16:19:18 +01:00
Matthew Honnibal	90aec6d2f6	Fix vectors for reserved words. Closes #2871	2018-12-10 16:09:49 +01:00
Matthew Honnibal	16fd8dce1d	Add get_string_id helper to spacy.strings	2018-12-10 16:09:26 +01:00
Matthew Honnibal	cc1ea03004	Add test for issue #2871 -- vectors for reserved words	2018-12-10 16:09:10 +01:00
Matthew Honnibal	375f0dc529	💫 Make TextCategorizer default to a simpler, GPU-friendly model (#3038 ) Currently the TextCategorizer defaults to a fairly complicated model, designed partly around the active learning requirements of Prodigy. The model's a bit slow, and not very GPU-friendly. This patch implements a straightforward CNN model that still performs pretty well. The replacement model also makes it easy to use the LMAO pretraining, since most of the parameters are in the CNN. The replacement model has a flag to specify whether labels are mutually exclusive, which defaults to True. This has been a common problem with the text classifier. We'll also now be able to support adding labels to pretrained models again. Resolves #2934, #2756, #1798, #1748.	2018-12-10 14:37:39 +01:00
Matthew Honnibal	b1c8731b4d	Make spacy train respect LOG_FRIENDLY	2018-12-10 09:46:53 +01:00
Matthew Honnibal	6936ca1664	Merge branch 'develop' of https://github.com/explosion/spaCy into develop	2018-12-10 09:44:07 +01:00
Matthew Honnibal	4405b5c875	Fix resizing edge-case for NER	2018-12-10 06:25:17 +00:00
Matthew Honnibal	0994dc50d8	Merge branch 'develop' of https://github.com/explosion/spaCy into develop	2018-12-10 05:35:01 +00:00
Matthew Honnibal	24f2e9bc07	Tweak training params	2018-12-09 17:08:58 +00:00
Matthew Honnibal	16c5861d29	Fix NER space constraints Allow entities to end on spaces, to avoid stumping the oracle when we're inside an entity, and there's a space just before a correct entity.	2018-12-09 08:06:45 +01:00
Matthew Honnibal	1b1a1af193	Fix printing in spacy train	2018-12-09 06:03:49 +01:00
Matthew Honnibal	d2ac618af1	Set cbb_maxout_pieces=3	2018-12-08 23:27:29 +01:00
Matthew Honnibal	cb16b78b0d	Set dropout rate to 0.2	2018-12-08 19:59:11 +01:00
Matthew Honnibal	e5685d98a2	Fix averaging in textcat example (closes #2745 ) (#3032 ) [ci skip]	2018-12-08 13:27:05 +01:00
Matthew Honnibal	2c2db0c492	💫 Allow Span to take text label (#3031 ) Fixes #3027. * Allow Span.__init__ to take unicode values for the `label` argument. * Allow `Span.label_` to be writeable. - [x] I have submitted the spaCy Contributor Agreement. - [x] I ran the tests, and all new and existing tests passed. - [ ] My changes don't require a change to the documentation, or if they do, I've added all required information.	2018-12-08 13:08:41 +01:00
Matthew Honnibal	11a29af751	Set cupy.random seed in fix_random_seed helper	2018-12-08 12:37:38 +01:00
Ines Montani	8c0f0f50bc	Use nlp.make_doc instead of nlp for patterns [ci skip]	2018-12-08 11:56:01 +01:00
Ines Montani	ffdd5e964f	Small CLI improvements (#3030 ) * Add todo * Auto-format * Update wasabi pin * Format training results with wasabi * Remove loading animation from model saving Currently behaves weirdly * Inline messages * Remove unnecessary path2str Already taken care of by printer * Inline messages in CLI * Remove unused function * Move loading indicator into loading function * Check for invalid whitespace entities	2018-12-08 11:49:43 +01:00
Matthew Honnibal	8aa7882762	Make NORM a token attribute (#3029 ) See #3028. The solution in this patch is pretty debateable. What we do is give the TokenC struct a .norm field, by repurposing the previously idle .sense attribute. It's nice to repurpose a previous field because it means the TokenC doesn't change size, so even if someone's using the internals very deeply, nothing will break. The weird thing here is that the TokenC and the LexemeC both have an attribute named NORM. This arguably assists in backwards compatibility. On the other hand, maybe it's really bad! We're changing the semantics of the attribute subtly, so maybe it's better if someone calling lex.norm gets a breakage, and instead is told to write lex.default_norm? Overall I believe this patch makes the NORM feature work the way we sort of expected it to work. Certainly it's much more like how the docs describe it, and more in line with how we've been directing people to use the norm attribute. We'll also be able to use token.norm to do stuff like spelling correction, which is pretty cool.	2018-12-08 10:49:10 +01:00
Matthew Honnibal	a338c6f8f6	Fix JSON segmentation bug that affected French Fix a bug in the JSON streaming code that GoldCorpus uses. Escaped slashes were being handled incorrectly. This bug caused low scores for French in the early v2.1.0 alphas, because most of the data was not being read in. Fittingly, the document that triggered the bug was a Wikipedia article about Perl. Parsing perl remains difficult!	2018-12-08 10:41:24 +01:00
Paul O'Leary McCann	7dd21b66d5	Extras require mecab (#3024 ) * Add note that Unidic is required for Japanese This addresses #3001. -POLM * Add extras_require for mecab with old version Related to issue #3018. * mecab → ja Co-Authored-By: polm <polm@dampfkraft.com>	2018-12-08 06:34:49 +01:00
Matthew Honnibal	6f36b6bc4e	Pin pex version	2018-12-07 23:42:48 +01:00
Matthew Honnibal	b2bfd1e1c8	Move dropout and batch sizes out of global scope in train cmd	2018-12-07 20:54:35 +01:00
Aki Ariga	7fcd6419ff	Upadate the document for Unidic link with latest version URL (#3022 ) * Upadate Unidic link for latest version in document This patch improves #3017 . The link for Unidic was old version one, so will the lates version. * Add contributor agreement * Use more specific link for unidic-cwj	2018-12-07 17:24:48 +01:00
Matthew Honnibal	f8c4ee34fe	Update wasabi pin	2018-12-07 01:43:07 +01:00
Matthew Honnibal	40e0da9cc1	Merge branch 'develop' of https://github.com/explosion/spaCy into develop	2018-12-07 00:12:22 +00:00
Matthew Honnibal	1e6725e9b7	Try to prevent spaces from being tagged as entities	2018-12-07 00:12:12 +00:00
Matthew Honnibal	427c0693c8	Fix missing comma in init-model command	2018-12-06 22:48:31 +01:00
Amandine Périnet	0b44ea23bd	Lemmatization of Nouns - French : adding rules and vocabulary (#2992 ) * modifying FR lemmatization for nouns * modifying FR lemmatization for nouns * adding contributor agreement for amperinet * adding rules for words with inclusive parentheses wrongly tokenized * adding contributor agreement for amperinet * adding a missing comma	2018-12-06 22:42:18 +01:00
Matthew Honnibal	d896fbca62	Fix batch size in parser.pipe	2018-12-06 21:45:56 +01:00
Matthew Honnibal	bb3304a4f1	Fix pickle tests	2018-12-06 20:46:36 +01:00
Matthew Honnibal	e619f45287	Fix pickle tests	2018-12-06 20:43:47 +01:00
Matthew Honnibal	0a60726215	Remove cytoolz usage in CLI	2018-12-06 20:37:00 +01:00
Matthew Honnibal	c0af627f32	Fix dill usage in vocab	2018-12-06 18:53:16 +01:00
Matthew Honnibal	9520489225	Fix removabl of dill (for srsly)	2018-12-06 18:46:09 +01:00
Ines Montani	27905a7b14	Remove reference to cuda10 in docs (closes #2894 ) [ci skip]	2018-12-06 16:05:37 +01:00
Matthew Honnibal	711f108532	Fix cytoolz import cytoolz	2018-12-06 16:04:12 +01:00
Gavriel Loria	9c8c4287bf	Accept iob2 and allow generic whitespace (#2999 ) * accept non-pipe whitespace as delimiter; allow iob2 filename * added small documentation note for IOB2 allowance * added contributor agreement	2018-12-06 15:50:25 +01:00
Amandine Périnet	2457318b7a	Lemmatization of Verbs - French : adding rules and vocabulary (#3006 ) * updating rules and vocabulary for French lemmatization of verbs * updating the file with French auxiliary verb * updating rules and vocabulary for French lemmatization of verbs * adding contributor agreement for amperinet * adding rules for words with inclusive parentheses wrongly tokenized	2018-12-06 15:49:28 +01:00
Beate Sildnes	f0d7e206ec	Updated wordforms for Norwegian lemmatizer (#3007 ) * Updated wordforms for Norwegian lemmatizer Upload of updated lists of wordforms for the Norwegian lemmatizer (nouns, verbs, adverbs, adjectives and lookup). * Add spaCy contributor agreement for user beatesi * Updated wordforms for Norwegian lemmatizer	2018-12-06 15:46:18 +01:00
Paul O'Leary McCann	b36f6eabfb	Add note that Unidic is required for Japanese (#3017 ) This addresses #3001. -POLM	2018-12-06 15:14:10 +01:00
Matthew Honnibal	cabaadd793	Fix build error from bad import Thinc v7.0.0.dev6 moved FeatureExtracter around and didn't add a compatibility import.	2018-12-06 15:12:39 +01:00
Matthew Honnibal	8f6555df4e	Update requirements	2018-12-04 00:07:28 +01:00
Matthew Honnibal	378ca4b46d	Fix OSX build problem	2018-12-04 00:06:42 +01:00
Matthew Honnibal	ea00dbaaa4	Remove usage of itertools.islice	2018-12-03 02:43:03 +01:00
Matthew Honnibal	3df26d820f	Sort requirements	2018-12-03 02:41:05 +01:00
Matthew Honnibal	5ed19fbee2	Remove cytoolz dependency	2018-12-03 02:37:22 +01:00
Matthew Honnibal	db75c70550	Remove dill dependency	2018-12-03 02:31:19 +01:00
Matthew Honnibal	c7b33b24f1	Fix conflict	2018-12-03 02:20:20 +01:00

... 3 4 5 6 7 ...

9496 Commits