spaCy

mirror of https://github.com/explosion/spaCy.git synced 2025-03-14 23:22:19 +03:00

Author	SHA1	Message	Date
Sofie Van Landeghem	75a202ce65	TextCat updates and fixes (#6263 ) * small fix in example imports * throw error when train_corpus or dev_corpus is not a string * small fix in custom logger example * limit macro_auc to labels with 2 annotations * fix typo * also create parents of output_dir if need be * update documentation of textcat scores * refactor TextCatEnsemble * fix tests for new AUC definition * bump to 3.0.0a42 * update docs * rename to spacy.TextCatEnsemble.v2 * spacy.TextCatEnsemble.v1 in legacy * cleanup * small fix * update to 3.0.0rc2 * fix import that got lost in merge * cursed IDE * fix two typos	2020-10-18 14:50:41 +02:00
Ines Montani	5a6ed01ce0	Merge pull request #6262 from adrianeboyd/bugfix/template-en-vectors	2020-10-16 15:38:08 +02:00
Adriane Boyd	c8d04b79e2	Sort and add vectors for langs without transformers	2020-10-16 08:25:16 +02:00
Adriane Boyd	2fbd43c603	Use core lg models as vectors models in quickstart	2020-10-16 08:17:53 +02:00
Jan Margeta	1ad2213349	Fix TokenPatternSchema pattern field validation Empty pattern field should be considered invalid This is fixed by replacing minItems with min_items as described in Pydantic docs: https://pydantic-docs.helpmanual.io/usage/schema/	2020-10-16 00:41:21 +02:00
Borijan Georgievski	2311192ba1	Include Macedonian language (#6230 ) * Include Macedonian language * Fix indentation at char_classes.py * Fix indentation at char_classes.py * Add Macedonian tests, update lex_attrs and char_classes * Import unicode literals for python 2	2020-10-15 15:55:01 +02:00
Ines Montani	ff4267d181	Fix success message [ci skip]	2020-10-15 14:42:08 +02:00
Ines Montani	10611bf56a	Increment version [ci skip]	2020-10-15 13:30:11 +02:00
Ines Montani	4e17ddf75e	Merge pull request #6256 from adrianeboyd/bugfix/docs-to-json-raw	2020-10-15 10:35:01 +02:00
Ines Montani	b1d568a4df	Tidy up tests	2020-10-15 10:20:21 +02:00
Ines Montani	d165af26be	Auto-format [ci skip]	2020-10-15 10:08:53 +02:00
Adriane Boyd	a93d42861d	Use null raw for has_unknown_spaces in docs_to_json	2020-10-15 09:57:54 +02:00
Ines Montani	5665a21517	Tidy up	2020-10-15 09:30:32 +02:00
Ines Montani	5d62499266	Fix tests	2020-10-15 09:29:15 +02:00
Ines Montani	178760855f	Merge branch 'develop' into master-tmp	2020-10-15 09:06:03 +02:00
Ines Montani	bc85b12e6d	Merge pull request #6249 from svlandeg/feature/batch-tests	2020-10-15 08:57:56 +02:00
svlandeg	0796401c19	call NumpyOps instead of get_current_ops()	2020-10-14 16:55:00 +02:00
svlandeg	44e14ccae8	one more losses fix	2020-10-14 15:11:34 +02:00
svlandeg	0aa8851878	always return losses	2020-10-14 15:00:49 +02:00
svlandeg	e94a21638e	adding tests for trained models to ensure predict reproducibility	2020-10-13 21:07:13 +02:00
svlandeg	ede979d42f	formattting	2020-10-13 18:53:17 +02:00
svlandeg	ff83bfae3f	naming	2020-10-13 18:52:37 +02:00
svlandeg	6ccacff54e	add tests for individual spacy layers	2020-10-13 18:50:07 +02:00
svlandeg	c23041ae60	component tests single or multiple prediction	2020-10-13 16:26:53 +02:00
Ines Montani	1f49300862	Update transformer recommendations [ci skip]	2020-10-13 15:41:17 +02:00
Sofie Van Landeghem	f8a1c1afd6	avoid dropout at runtime (#6247 )	2020-10-13 14:39:59 +02:00
Ines Montani	86d648740f	Fix morph representation in Doc.to_json	2020-10-13 11:39:03 +02:00
Ines Montani	7f92a5ee6a	Update spacy/lang/ta/examples.py	2020-10-13 11:03:35 +02:00
Ines Montani	a0e12c136b	Increment version [ci skip]	2020-10-13 10:00:53 +02:00
Ines Montani	f090f39f17	Merge pull request #6245 from svlandeg/bugfix/else bugfix in _pipe	2020-10-13 09:59:06 +02:00
svlandeg	1f465bea18	if-else	2020-10-13 09:27:19 +02:00
svlandeg	40276fd3be	update NEL docs after latest refactor	2020-10-12 11:41:27 +02:00
Ines Montani	4fa967ea84	Increment version [ci skip]	2020-10-11 13:10:58 +02:00
Ines Montani	ab890a35f9	Make console logger table more compact	2020-10-11 12:55:46 +02:00
Ines Montani	99606e46fe	Relax meta.json schema [ci skip]	2020-10-11 12:30:57 +02:00
svlandeg	3a505e7e14	small edit to ensure the new word was indeed new	2020-10-10 21:05:28 +02:00
svlandeg	68d79796c6	add test for vocab after serializing KB	2020-10-10 20:59:48 +02:00
Ines Montani	539b0c10da	Tidy up and auto-format	2020-10-10 19:14:48 +02:00
Ines Montani	bfa3931c9d	Revert added_strings change (#6236 )	2020-10-10 18:55:07 +02:00
Ines Montani	796f8b9424	Increment version	2020-10-09 18:00:27 +02:00
Ines Montani	525f798841	Fix typo in test	2020-10-09 18:00:21 +02:00
Ines Montani	8ac5f22253	Adjust error message	2020-10-09 18:00:16 +02:00
svlandeg	08cb085f6c	Merge remote-tracking branch 'upstream/develop' into fix/various	2020-10-09 17:01:27 +02:00
Ines Montani	b7cb9d95e4	Merge pull request #6229 from svlandeg/bugfix/disabled	2020-10-09 16:05:11 +02:00
svlandeg	e972ecba72	add utf8 encoding for opening file	2020-10-09 16:03:14 +02:00
Ines Montani	9fb3244672	Merge pull request #6231 from adrianeboyd/feature/include-static-vectors	2020-10-09 15:54:52 +02:00
svlandeg	040c7c0541	fix get_dim calls in build_simple_cnn_text_classifier	2020-10-09 15:40:58 +02:00
Adriane Boyd	727370c633	Remove Span._recalculate_indices Remove `Span._recalculate_indices`, which is a remnant from the deprecated `Span.merge`.	2020-10-09 14:42:51 +02:00
svlandeg	853edace37	fix MultiHashEmbed example in documentation	2020-10-09 14:11:06 +02:00
svlandeg	06b9d213fd	formatting	2020-10-09 12:19:47 +02:00
svlandeg	2cafba5f50	shorten error message for clarity	2020-10-09 12:17:35 +02:00
Ines Montani	4771a10503	Make test more explicit [ci skip]	2020-10-09 12:15:26 +02:00
Ines Montani	cc3646b06c	Add xfailing test for peculiar spans failure [ci skip]	2020-10-09 12:10:25 +02:00
svlandeg	8316bc7d4a	bugfix DisabledPipes	2020-10-09 12:06:20 +02:00
svlandeg	18dfb27985	Add custom error when evaluation throws a KeyError	2020-10-09 12:05:33 +02:00
Adriane Boyd	39aabf50ab	Also rename to include_static_vectors in CharEmbed	2020-10-09 11:54:48 +02:00
Florijan Stamenković	18f5c309dc	Fix Issue 6207 (#6208 ) * Regression test for issue 6207 * Fix issue 6207 * Sign contributor agreement * Minor adjustments to test Co-authored-by: Adriane Boyd <adrianeboyd@gmail.com>	2020-10-09 10:14:40 +02:00
Duygu Altinok	80fb1bffc9	Ordinal numbers for Turkish (#6142 ) * minor ordinal number addition * fixed typo * added corresponding lexical test	2020-10-09 10:13:15 +02:00
Duygu Altinok	2fad279a44	Turkish language syntax iterators (#6191 ) * added tr_vocab to config * basic test * added syntax iterator to Turkish lang class * first version for Turkish syntax iter, without flat * added simple tests with nmod, amod, det * more tests to amod and nmod * separated noun chunks and parser test * rearrangement after nchunk parser separation * added recursive NPs * tests with complicated recursive NPs * tests with conjed NPs * additional tests for conj NP * small modification for shaving off conj from NP * added tests with flat * more tests with flat * added examples with flats conjed * added inner func for flat trick * corrected parse Co-authored-by: Adriane Boyd <adrianeboyd@gmail.com>	2020-10-09 10:10:22 +02:00
Sofie Van Landeghem	d093d6343b	TrainablePipe (#6213 ) * rename Pipe to TrainablePipe * split functionality between Pipe and TrainablePipe * remove unnecessary methods from certain components * cleanup * hasattr(component, "pipe") should be sufficient again * remove serialization and vocab/cfg from Pipe * unify _ensure_examples and validate_examples * small fixes * hasattr checks for self.cfg and self.vocab * make is_resizable and is_trainable properties * serialize strings.json instead of vocab * fix KB IO + tests * fix typos * more typos * _added_strings as a set * few more tests specifically for _added_strings field * bump to 3.0.0a36	2020-10-08 21:33:49 +02:00
Ines Montani	8ff73f04db	Fix morph in Doc.to_json	2020-10-08 14:44:35 +02:00
Ines Montani	064575d79d	Merge pull request #6216 from svlandeg/feature/nel-initialize	2020-10-08 11:14:12 +02:00
svlandeg	3e2e1fd323	cleanup	2020-10-08 10:37:32 +02:00
svlandeg	eaf5c265cb	set_kb method for entity_linker	2020-10-08 10:34:01 +02:00
Ines Montani	010956d493	Clear rule-based components on initialize	2020-10-08 09:51:31 +02:00
Baranitharan	d6037c1860	added sentence	2020-10-08 08:22:58 +05:30
Baranitharan	81afe9b19d	Update examples.py	2020-10-08 08:17:25 +05:30
Sofie Van Landeghem	241cd112f5	add reenabled pipe names back to the meta before serializing (#6219 )	2020-10-08 00:44:16 +02:00
Sofie Van Landeghem	2998131416	Reproducibility for TextCat and Tok2Vec (#6218 ) * ensure fixed seed in HashEmbed layers * forgot about the joys of python 2	2020-10-08 00:43:46 +02:00
svlandeg	efedccea8d	fix tests	2020-10-07 15:29:52 +02:00
svlandeg	6b8bdb2d39	add init_config to nlp.create_pipe	2020-10-07 14:58:16 +02:00
svlandeg	33c2d4af16	move kb_loader to initialize for NEL instead of constructor	2020-10-07 14:56:00 +02:00
Wannaphong Phatthiyaphaibun	9fc8392b38	Add Thai tag map (LST20 Corpus) (#6163 ) * Add Thai tag map (LST20 Corpus) By @korakot * Update tag_map.py * Update tag_map.py * Update tag_map.py	2020-10-07 11:12:01 +02:00
Duygu Altinok	7e821c2776	Turkish language syntax iterators (#6191 ) * added tr_vocab to config * basic test * added syntax iterator to Turkish lang class * first version for Turkish syntax iter, without flat * added simple tests with nmod, amod, det * more tests to amod and nmod * separated noun chunks and parser test * rearrangement after nchunk parser separation * added recursive NPs * tests with complicated recursive NPs * tests with conjed NPs * additional tests for conj NP * small modification for shaving off conj from NP * added tests with flat * more tests with flat * added examples with flats conjed * added inner func for flat trick * corrected parse Co-authored-by: Adriane Boyd <adrianeboyd@gmail.com>	2020-10-07 11:07:52 +02:00
Duygu Altinok	2ce6fc2611	Turkish tag map and morph rules addition (#6141 ) * feat: added turkish tag map * feat: morph rules cconj and sconj * feat: more conjuncts * feat: added popular postpositions * feat: added adverbs * feat: added personal pronouns * feat: added reflexive pronouns * minor: corrected case capital * minor: fixed comma typo * feat: added indef pronouns * feat: added dict iter * fixed comma typo * updated language class with tag map and morph * use default tag map instead * removed tag map	2020-10-07 10:27:36 +02:00
Duygu Altinok	b95a11dd95	Ordinal numbers for Turkish (#6142 ) * minor ordinal number addition * fixed typo * added corresponding lexical test	2020-10-07 10:25:37 +02:00
Rahul Gupta	1a00bff06d	Hindi: Adds tests for lexical attributes (norm and like_num) (#5829 ) * Hindi: Adds tests for lexical attributes (norm and like_num) * Signs and sdds the contributor agreement * Add ordinal numbers to be tagged as like_num * Adds alternate pronunciation for 31 and 39	2020-10-07 10:23:32 +02:00
Nuccy90	c809b2c8e7	Update morph_rules.py (#6102 ) * Update morph_rules.py Added "dig" and "dej" ("you" in accusative form) * Create Nuccy90.md * Update Nuccy90.md	2020-10-06 15:14:47 +02:00
Matthew Honnibal	1a500f9717	Set version to v3.0.0a35	2020-10-06 14:19:07 +02:00
Sofie Van Landeghem	fff3f8ccfa	Fix packaging pin (#6212 ) * pin packaging to >=20.0 * ignore spacy-pkuseg in requirements unit test	2020-10-06 14:16:05 +02:00
Matthew Honnibal	cfb9770a94	Fix empty input into StaticVectors layer (#6211 ) * Add test for empty doc(s) * Fix empty check in staticvectors * Remove xfail * Update spacy/ml/staticvectors.py	2020-10-06 14:15:41 +02:00
Florijan Stamenković	9db670b996	Fix Issue 6207 (#6208 ) * Regression test for issue 6207 * Fix issue 6207 * Sign contributor agreement * Minor adjustments to test Co-authored-by: Adriane Boyd <adrianeboyd@gmail.com>	2020-10-06 11:17:37 +02:00
Ines Montani	568e12215d	Merge pull request #6206 from svlandeg/fix/patterns-init	2020-10-06 10:27:23 +02:00
svlandeg	9b4cf7b0b6	update output of debug config command	2020-10-06 09:47:23 +02:00
svlandeg	ff9ac39c88	read entity_ruler patterns with srsly.read_jsonl.v1	2020-10-05 22:50:14 +02:00
Ines Montani	126268ce50	Auto-format [ci skip]	2020-10-05 21:58:18 +02:00
Ines Montani	1a554bdcb1	Update docs and docstring [ci skip]	2020-10-05 21:55:27 +02:00
Ines Montani	9614e53b02	Tidy up and auto-format	2020-10-05 21:55:18 +02:00
Ines Montani	181039bd17	Merge pull request #6205 from explosion/feature/embed-features	2020-10-05 21:49:10 +02:00
Ines Montani	5ba418b08c	Merge branch 'develop' of https://github.com/explosion/spaCy into develop	2020-10-05 21:44:01 +02:00
Ines Montani	568617af58	Merge pull request #6202 from explosion/feature/project-spacy-version	2020-10-05 21:40:52 +02:00
Ines Montani	2d0c0134bc	Adjust message [ci skip]	2020-10-05 21:38:23 +02:00
Ines Montani	6abfc2911d	Merge pull request #6203 from adrianeboyd/feature/zh-spacy-pkuseg	2020-10-05 21:35:57 +02:00
Matthew Honnibal	b7e01d2024	Fix quickstart	2020-10-05 21:21:30 +02:00
Matthew Honnibal	ff8b980775	Upd quickstart template	2020-10-05 21:19:41 +02:00
Matthew Honnibal	91d0fbb588	Fix test	2020-10-05 21:13:53 +02:00
Ines Montani	9ca283a899	Merge branch 'develop' into feature/project-spacy-version	2020-10-05 21:06:07 +02:00
Ines Montani	0135f6ed95	Enable commit check via env var	2020-10-05 20:51:15 +02:00
Matthew Honnibal	b392d48e76	Fix test	2020-10-05 20:17:07 +02:00
Ines Montani	be99f1e4de	Remove output dirs before training (#6204 ) * Remove output dirs before training * Re-raise error if cleaning fails	2020-10-05 20:11:16 +02:00
Matthew Honnibal	e50047f1c5	Check lengths match	2020-10-05 20:02:45 +02:00
Ines Montani	582701519e	Remove __release__ flag	2020-10-05 20:00:49 +02:00
Ines Montani	d58fb42707	Add spacy_version option and validation for project.yml	2020-10-05 20:00:42 +02:00
Matthew Honnibal	db84d175c3	Fix test	2020-10-05 19:59:30 +02:00
Matthew Honnibal	cdd2b79b6d	Remove deprecated MultiHashEmbed	2020-10-05 19:58:18 +02:00
Matthew Honnibal	6dcc4a0ba6	Simplify MultiHashEmbed signature	2020-10-05 19:57:45 +02:00
svlandeg	193e0d5a98	add docs for entity_ruler.initialize	2020-10-05 18:04:08 +02:00
svlandeg	3ac3447eee	cleanup	2020-10-05 17:50:37 +02:00
svlandeg	9eb813a35d	Merge remote-tracking branch 'upstream/develop' into fix/patterns-init	2020-10-05 17:49:44 +02:00
Adriane Boyd	f102ef6b54	Read features.msgpack instead of features.pkl	2020-10-05 17:47:39 +02:00
svlandeg	4e3ace4b8c	is_trainable method	2020-10-05 17:43:42 +02:00
Ines Montani	84fedcebab	Make args keyword-only [ci skip] Co-authored-by: Matthew Honnibal <honnibal+gh@gmail.com>	2020-10-05 17:07:35 +02:00
Matthew Honnibal	71e73ed0a6	Merge branch 'develop' into feature/embed-features	2020-10-05 17:00:05 +02:00
Matthew Honnibal	3ee3649b52	Fix augment	2020-10-05 16:59:49 +02:00
Matthew Honnibal	22937d25a9	Merge branch 'develop' into feature/embed-features	2020-10-05 16:42:17 +02:00
Matthew Honnibal	8deed614e9	Fix augment	2020-10-05 16:41:45 +02:00
Matthew Honnibal	4ed3e037df	Fix augment	2020-10-05 16:40:55 +02:00
Matthew Honnibal	9f1bc3f24c	Fix augment	2020-10-05 16:40:23 +02:00
svlandeg	dc06912c76	prevent loss keyerror for non-trainable components	2020-10-05 16:33:28 +02:00
Adriane Boyd	187234648c	Revert back to "default" as default for pkuseg_user_dict	2020-10-05 16:24:28 +02:00
svlandeg	65abd77779	add finish_update to Pipe	2020-10-05 16:23:33 +02:00
Matthew Honnibal	90040aacec	Fix merge	2020-10-05 16:12:01 +02:00
Matthew Honnibal	93a98e8c3e	Merge branch 'develop' into feature/embed-features	2020-10-05 15:51:31 +02:00
Matthew Honnibal	eb9ba61517	Format	2020-10-05 15:29:49 +02:00
Matthew Honnibal	7d93575f35	spacy/tests/	2020-10-05 15:28:12 +02:00
Matthew Honnibal	f4ca9a39cb	spacy/tests/	2020-10-05 15:27:06 +02:00
Matthew Honnibal	f2f1deca66	spacy/tests/	2020-10-05 15:24:33 +02:00
Matthew Honnibal	8ec79ad3fa	Allow configuration of MultiHashEmbed features Update arguments to MultiHashEmbed layer so that the attributes can be controlled. A kind of tricky scheme is used to allow optional specification of the rows. I think it's an okay balance between flexibility and convenience.	2020-10-05 15:22:00 +02:00
Ines Montani	7946fd84bb	Merge pull request #6200 from adrianeboyd/bugfix/vocab-disk-lookups-vectors Always serialize lookups and vectors to disk	2020-10-05 15:15:25 +02:00
Ines Montani	8171e28b20	Remove logging [ci skip] This would be fired on each example, which is wrong	2020-10-05 15:09:52 +02:00
svlandeg	251b3eb4e5	add initialize method for entity_ruler	2020-10-05 14:59:13 +02:00
Sofie Van Landeghem	f4f49f5877	update blis (#6198 ) * allow higher blis version * fix typo * bump to 3.0.0a34 * fix pins in other files	2020-10-05 14:58:56 +02:00
Adriane Boyd	5d19dfc9d3	Update Chinese tokenizer for spacy-pkuseg fork	2020-10-05 14:21:53 +02:00
Matthew Honnibal	6a9d14e35a	Merge branch 'develop' of https://github.com/explosion/spaCy into develop	2020-10-05 14:17:41 +02:00
Matthew Honnibal	d2b9aafb8c	Fix augmenter	2020-10-05 14:14:49 +02:00
Ines Montani	6260fa3c10	Merge pull request #6201 from svlandeg/fix/error_nr	2020-10-05 14:00:57 +02:00
Ines Montani	6958510bda	Include spaCy version check in project CLI	2020-10-05 13:53:07 +02:00
Ines Montani	20f2a17a09	Merge test_misc and test_util	2020-10-05 13:45:57 +02:00
svlandeg	fd2d48556c	fix E902 and E903 numbering	2020-10-05 13:43:32 +02:00
Ines Montani	1c641e41c3	Remove unused import [ci skip]	2020-10-05 11:50:11 +02:00
Adriane Boyd	03cfb2d2f4	Always serialize lookups and vectors to disk	2020-10-05 09:40:20 +02:00
Adriane Boyd	b0b93854cb	Update ru/uk lemmatizers for new nlp.initialize	2020-10-05 09:27:16 +02:00
Ines Montani	549758f67d	Adjust test for now	2020-10-04 23:16:09 +02:00
Ines Montani	4b15ff7504	Increment version [ci skip]	2020-10-04 22:47:04 +02:00
Ines Montani	f1d1f78636	Make warning debug log [ci skip]	2020-10-04 22:44:21 +02:00
Ines Montani	3c36a57e84	Update data augmenters (#6196 ) * Draft lower-case augmenter * Make warning a debug log * Update lowercase augmenter, docs and tests Co-authored-by: Matthew Honnibal <honnibal+gh@gmail.com>	2020-10-04 17:46:29 +02:00
Ines Montani	d38dc466c5	Adjust error [ci skip]	2020-10-04 15:26:01 +02:00
Ines Montani	496228771d	Merge pull request #6194 from explosion/master-tmp	2020-10-04 15:25:41 +02:00
Ines Montani	0307a228c8	Merge pull request #6193 from explosion/fix/adjust-pipe-init Adjust [initialize.components] on Language.remove_pipe and Language.rename_pipe	2020-10-04 15:20:54 +02:00
Ines Montani	59deeb7da6	Merge branch 'develop' into master-tmp	2020-10-04 14:52:20 +02:00
Ines Montani	43d7652635	Merge pull request #6192 from explosion/feature/init-attr-ruler	2020-10-04 14:46:37 +02:00
Ines Montani	8f018e47f8	Adjust [initialize.components] on Language.remove_pipe and Language.rename_pipe	2020-10-04 14:43:45 +02:00
Matthew Honnibal	84ae197dd6	Fix logger	2020-10-04 14:16:53 +02:00
Ines Montani	11347f34da	Tidy up, tests and docs	2020-10-04 13:54:05 +02:00
Matthew Honnibal	96b636c2d3	Update attribute ruler	2020-10-04 13:08:21 +02:00
Ines Montani	bcd52e5486	Tidy up errors and warnings	2020-10-04 11:16:31 +02:00
Ines Montani	ff914f4e6f	Lazy-load xx	2020-10-04 11:10:26 +02:00
Ines Montani	d3b3663942	Adjust error message and add test	2020-10-04 10:11:27 +02:00
Ines Montani	2110e8f86d	Auto-format	2020-10-04 10:06:49 +02:00
Ines Montani	cc08c88a89	Merge pull request #6187 from svlandeg/fix/begin_training_pipe	2020-10-04 10:01:02 +02:00
svlandeg	3f657ed3a1	implement warning in __init_subclass__ instead	2020-10-03 22:34:10 +02:00
Matthew Honnibal	3b2a78720c	Upd morphologizer	2020-10-03 19:35:19 +02:00
Matthew Honnibal	835070cedc	Upd test	2020-10-03 19:35:10 +02:00
Matthew Honnibal	70b9de8e58	Set version to v3.0.0a32	2020-10-03 19:26:52 +02:00
Matthew Honnibal	85ede32680	Format	2020-10-03 19:26:23 +02:00
Matthew Honnibal	b305f2ff5a	Fix loggers	2020-10-03 19:26:10 +02:00
Matthew Honnibal	4fccd2ceaf	Merge branch 'develop' of https://github.com/explosion/spaCy into develop	2020-10-03 19:13:55 +02:00
Matthew Honnibal	8ea8b7d940	Support loading labels in morphologizer	2020-10-03 19:13:42 +02:00
Ines Montani	c2401fca41	Add tests for Pipe.label_data	2020-10-03 19:12:46 +02:00
Ines Montani	80603f0fa5	Make SentenceRecognizer.label_data return None Overwrite the method from the base class (Tagger) but don't export anything in "init labels"	2020-10-03 18:54:09 +02:00
Ines Montani	d6c967401f	Increment version	2020-10-03 17:20:47 +02:00
Ines Montani	3bc3c05fcc	Tidy up and auto-format	2020-10-03 17:20:18 +02:00
Ines Montani	7c4ab7e82c	Fix Lemmatizer.get_lookups_config	2020-10-03 17:16:10 +02:00
Ines Montani	dd542ec6a4	Fix label initialization of textcat component (#6190 )	2020-10-03 17:07:38 +02:00
Ines Montani	989a96308f	Tidy up, auto-format, types	2020-10-03 16:31:58 +02:00
Matthew Honnibal	7b127f307e	Set version to v3.0.0a30	2020-10-03 16:06:42 +02:00
Matthew Honnibal	db419f6b2f	Improve control of training progress and logging (#6184 ) * Make logging and progress easier to control * Update docs * Cleanup errors * Fix ConfigValidationError * Pass stdout/stderr, not wasabi.Printer * Fix type * Upd logging example * Fix logger example * Fix type	2020-10-03 14:57:46 +02:00
Ines Montani	ae15c9de79	Raise error from caught KeyError to preserve traceback	2020-10-03 11:43:56 +02:00
Ines Montani	f758804401	Save one line of code	2020-10-03 11:41:28 +02:00
Stanislav Schmidt	3589a64d44	Change type of texts argument in pipe to iterable (#6186 ) * Change type of texts argument in pipe to iterable * Add contributor agreement	2020-10-02 21:00:11 +02:00
svlandeg	02247cccaf	Merge remote-tracking branch 'upstream/develop' into feature/small-fixes	2020-10-02 20:48:11 +02:00
svlandeg	fb48de349c	bwd compat for pipe.begin_training	2020-10-02 20:31:14 +02:00
Matthew Honnibal	6965cdf16d	Fix comment	2020-10-02 17:26:21 +02:00
Ines Montani	3cf10a0729	Merge pull request #6183 from adrianeboyd/feature/quickstart-morphologizer Add morphologizer to quickstart template	2020-10-02 17:08:01 +02:00
Adriane Boyd	62ccd5c4df	Relax model meta performance schema (#6185 ) Allow more embedded per_x in `ModelMetaSchema`	2020-10-02 16:37:21 +02:00
Sofie Van Landeghem	09dcb75076	small UX fix for DocBin (#6167 ) * add informative warning when messing up store_user_data DocBin flags * add informative warning when messing up store_user_data DocBin flags * cleanup test * rename to patterns_path	2020-10-02 15:43:32 +02:00
Ines Montani	f0b30aedad	Make lemmatizers use initialize logic (#6182 ) * Make lemmatizer use initialize logic and tidy up * Fix typo * Raise for uninitialized tables	2020-10-02 15:42:36 +02:00
Adriane Boyd	22158dc24a	Add morphologizer to quickstart template	2020-10-02 15:06:16 +02:00
Ines Montani	d2aa662ab2	Merge pull request #6179 from adrianeboyd/feature/token-morph-refactor-2 [ci skip]	2020-10-02 12:10:27 +02:00
Ines Montani	c41a4332e4	Add test for custom data augmentation	2020-10-02 11:37:56 +02:00
svlandeg	acc391c2a8	remove redundant str() call	2020-10-02 11:05:59 +02:00
Ines Montani	3856048437	Merge pull request #6178 from explosion/feature/file-readers Integrate file readers via srsly, update orth_variants loading	2020-10-02 10:26:09 +02:00
Adriane Boyd	f83dfe62da	Fix test	2020-10-02 10:17:26 +02:00
Adriane Boyd	65dfaa4f4b	Also accept MorphAnalysis in set_morph	2020-10-02 08:33:43 +02:00
Adriane Boyd	77e08c398f	Switch reset value for set_morph to None	2020-10-02 08:25:15 +02:00
Ines Montani	568768643e	Increment version [ci skip]	2020-10-02 01:50:13 +02:00
Ines Montani	01c1538c72	Integrate file readers	2020-10-02 01:36:06 +02:00
Ines Montani	af282ae732	Fix import	2020-10-02 01:12:34 +02:00
Ines Montani	e59ecb12c0	Auto-format	2020-10-02 01:12:30 +02:00
Matthew Honnibal	75a1569908	Merge	2020-10-01 23:07:53 +02:00
Matthew Honnibal	300e5a9928	Avoid relying on NORM in default v3 models (#6176 ) * Allow CharacterEmbed to specify feature * Default to LOWER in character embed * Update tok2vec * Use LOWER, not NORM	2020-10-01 23:05:55 +02:00
Ines Montani	5762876dcc	Update default config [ci skip]	2020-10-01 22:27:37 +02:00
Adriane Boyd	86c3ec9c2b	Refactor Token morph setting (#6175 ) * Refactor Token morph setting * Remove `Token.morph_` * Add `Token.set_morph()` * `0` resets `token.c.morph` to unset * Any other values are passed to `Morphology.add` * Add token.morph setter to set from MorphAnalysis	2020-10-01 22:21:46 +02:00
Matthew Honnibal	b854bca15c	Default to LOWER in character embed	2020-10-01 22:17:58 +02:00
Matthew Honnibal	684a77870b	Allow CharacterEmbed to specify feature	2020-10-01 22:17:26 +02:00
Ines Montani	da30701cd1	Increment version [ci skip]	2020-10-01 21:58:11 +02:00
Ines Montani	d48ddd6c9a	Remove default initialize lookups	2020-10-01 21:54:33 +02:00
Ines Montani	1700c8541e	Increment version [ci skip]	2020-10-01 17:57:16 +02:00
Ines Montani	f2627157c8	Update docs [ci skip]	2020-10-01 17:38:17 +02:00
Ines Montani	7f68f4bd92	Hide jsonl_loc on init vectors and tidy up [ci skip]	2020-10-01 16:44:17 +02:00
Adriane Boyd	27cbffff1b	Minor edit to CoNLL-U converter (#6172 ) This doesn't make a difference given how the `merged_morph` values override the `morph` values for all the final docs, but could have led to unexpected bugs in the future if the converter is modified.	2020-10-01 16:23:42 +02:00
Sofie Van Landeghem	a22215f427	Add FeatureExtractor from Thinc (#6170 ) * move featureextractor from Thinc * Update website/docs/api/architectures.md Co-authored-by: Ines Montani <ines@ines.io> * Update website/docs/api/architectures.md Co-authored-by: Ines Montani <ines@ines.io> Co-authored-by: Ines Montani <ines@ines.io>	2020-10-01 16:22:48 +02:00
Adriane Boyd	73538782a0	Switch Doc.__init__(ents=) to IOB tags (#6173 ) * Switch Doc.__init__(ents=) to IOB tags * Fix check for "-" * Allow "" or None as missing IOB tag	2020-10-01 16:22:18 +02:00
Adriane Boyd	df98d3ef9f	Update import from collections.abc (#6174 )	2020-10-01 16:21:49 +02:00
Yohei Tamura	3243ddac8f	Fix/span.sent (#6083 ) * add fail test * fix test * fix span.sent * Remove incorrect implicit check Co-authored-by: Adriane Boyd <adrianeboyd@gmail.com>	2020-10-01 14:01:52 +02:00
Ines Montani	0a8a124a6e	Update docs [ci skip]	2020-10-01 12:15:53 +02:00
Ines Montani	44160cd52f	Tidy up [ci skip]	2020-10-01 10:41:19 +02:00
Ines Montani	381258b75b	Merge pull request #6165 from explosion/feature/update-tokenizers-initialize	2020-10-01 09:49:47 +02:00
svlandeg	6787e56315	print debugging warning before raising error if model not properly initialized	2020-10-01 09:21:00 +02:00
svlandeg	5121972930	add types of Tok2Vec embedding layers	2020-10-01 09:20:09 +02:00
Ines Montani	4b6afd3611	Remove English [initialize] default block for now to get tests to pass	2020-09-30 23:49:29 +02:00
Ines Montani	6f29f68f69	Update errors and make Tokenizer.initialize args less strict	2020-09-30 23:48:47 +02:00
Ines Montani	a103ab5f1a	Update augmenter lookups and docs	2020-09-30 23:03:47 +02:00
Matthew Honnibal	5128298964	Add missing augmenter	2020-09-30 20:18:45 +02:00
Matthew Honnibal	59294e91aa	Restore the 'jsonl' arg for init vectors The lexemes.jsonl file is still used in our English vectors, and it may be required by users as well. I think it's worth supporting the option.	2020-09-30 19:06:50 +02:00
Matthew Honnibal	c379a4274a	Merge branch 'develop' of https://github.com/explosion/spaCy into develop	2020-09-30 16:52:42 +02:00
Matthew Honnibal	e58dca3028	Add read_labels	2020-09-30 16:52:27 +02:00
Ines Montani	23c63eefaf	Tidy up env vars [ci skip]	2020-09-30 15:15:11 +02:00
Elijah Rippeth	4cbb954281	reorder so tagmap is replaced only if a custom file is provided. (#6164 ) * reorder so tagmap is replaced only if a custom file is provided. * Remove unneeded variable initialization Co-authored-by: Adriane Boyd <adrianeboyd@gmail.com>	2020-09-30 13:26:06 +02:00
Adriane Boyd	6b7bb32834	Refactor Chinese initialization	2020-09-30 11:46:45 +02:00
Ines Montani	34f9c26c62	Add lexeme norm defaults	2020-09-30 10:20:14 +02:00
Ines Montani	a5debb356d	Tidy up and adjust logging [ci skip]	2020-09-30 01:22:08 +02:00
Ines Montani	56a2f778c4	Add logging [ci skip]	2020-09-30 01:08:55 +02:00
Ines Montani	fe3f111c37	Merge pull request #6168 from explosion/fix/default-corpus-values	2020-09-30 00:24:02 +02:00
Ines Montani	b799af16de	Don't raise in Pipe.initialize if not implemented	2020-09-30 00:05:27 +02:00
Matthew Honnibal	bc61691f6f	Merge branch 'develop' of https://github.com/explosion/spaCy into develop	2020-09-29 23:41:04 +02:00
Matthew Honnibal	f52249fe2e	Fix data augmentation	2020-09-29 23:40:54 +02:00
Matthew Honnibal	14c4da547f	Try to fix augmentation	2020-09-29 23:08:56 +02:00
Ines Montani	ae51843468	Remove augmenter from jinja template [ci skip]	2020-09-29 23:08:50 +02:00
Ines Montani	9bb958fd0a	Fix debug data [ci skip]	2020-09-29 23:07:11 +02:00
Matthew Honnibal	a2aa1f6882	Disable the OVL augmentation by default	2020-09-29 23:02:40 +02:00
Ines Montani	df8dd91b6f	Merge branch 'develop' into fix/default-corpus-values	2020-09-29 22:55:39 +02:00
Ines Montani	0a1ee109db	Remove init form path	2020-09-29 22:53:18 +02:00
Ines Montani	ad6d40d028	Add logging	2020-09-29 22:53:14 +02:00
Ines Montani	c334a7d45f	Remove	2020-09-29 22:38:39 +02:00
Ines Montani	1aeef3bfbb	Make corpus paths default to None and improve errors	2020-09-29 22:33:46 +02:00
Ines Montani	0250bcf6a3	Show validation error during init	2020-09-29 22:29:09 +02:00
Ines Montani	da30bae8a6	Use __pyx_vtable__ instead of __reduce_cython__	2020-09-29 22:04:17 +02:00
Ines Montani	43c92ec8c9	Resolve dir for better output [ci skip]	2020-09-29 22:01:04 +02:00
Ines Montani	fa47f87924	Tidy up and auto-format	2020-09-29 21:39:28 +02:00
Ines Montani	604be54a5c	Support --code in evaluate CLI [ci skip]	2020-09-29 21:20:56 +02:00
Ines Montani	6467a560e3	WIP: Test updating Chinese tokenizer	2020-09-29 21:10:22 +02:00
Ines Montani	4f3102d09c	Auto-format	2020-09-29 21:09:10 +02:00
Ines Montani	798040bc1d	Fix language detection	2020-09-29 21:08:13 +02:00
Ines Montani	78021089f9	Merge pull request #6160 from explosion/feature/prepare	2020-09-29 20:55:13 +02:00
Ines Montani	c3f8c09d7d	Merge pull request #6154 from adrianeboyd/bugfix/chinese-tokenizer-pickle	2020-09-29 20:54:59 +02:00
Ines Montani	d3c63b7965	Merge branch 'develop' into feature/prepare	2020-09-29 20:53:05 +02:00
Ines Montani	2be80379ec	Fix small issues, resolve_dot_names and debug model	2020-09-29 20:38:35 +02:00
Matthew Honnibal	a4da3120b4	Fix multitasks	2020-09-29 18:33:16 +02:00
Matthew Honnibal	0b5c72fce2	Fix incorrect docstrings	2020-09-29 18:30:38 +02:00
Ines Montani	7851020653	Update tests	2020-09-29 18:14:15 +02:00
Ines Montani	71a0ee274a	Move init labels to init pipeline module	2020-09-29 18:09:33 +02:00
Ines Montani	dba26186ef	Handle None default args in Cython methods	2020-09-29 18:08:02 +02:00
Ines Montani	9353a82076	Auto-format	2020-09-29 18:07:48 +02:00
Ines Montani	534e1ef498	Fix template	2020-09-29 17:02:55 +02:00
Ines Montani	f2352eb701	Test with default value	2020-09-29 17:00:40 +02:00
Matthew Honnibal	8ce9f44433	Merge branch 'feature/prepare' of https://github.com/explosion/spaCy into feature/prepare	2020-09-29 16:57:38 +02:00
Matthew Honnibal	e4f535a964	Fix Pipe.labels	2020-09-29 16:55:07 +02:00
Matthew Honnibal	4ad26f4a2f	Move reader	2020-09-29 16:54:53 +02:00
Ines Montani	30c76dbd67	Merge branch 'feature/prepare' of https://github.com/explosion/spaCy into feature/prepare	2020-09-29 16:53:48 +02:00
Matthew Honnibal	43fc7a316d	Add registry function for reading jsonl	2020-09-29 16:49:09 +02:00
Matthew Honnibal	1fd002180e	Allow more components to use labels	2020-09-29 16:48:56 +02:00
Matthew Honnibal	99bff78617	Use labels in tagger	2020-09-29 16:48:44 +02:00
Matthew Honnibal	ca72608059	Fix language	2020-09-29 16:48:33 +02:00
Matthew Honnibal	10847c7f4e	Fix arg	2020-09-29 16:48:07 +02:00
Ines Montani	fd594cfb9b	Tighten up format	2020-09-29 16:47:55 +02:00
Matthew Honnibal	e70a00fa76	Remove unnecessary warning from train	2020-09-29 16:47:54 +02:00
Matthew Honnibal	3f0d61232d	Remove outdated arg from train	2020-09-29 16:47:44 +02:00
Matthew Honnibal	e957d66b92	Merge branch 'feature/prepare' of https://github.com/explosion/spaCy into feature/prepare	2020-09-29 16:22:53 +02:00
Ines Montani	978ab54a84	Fix logging	2020-09-29 16:22:41 +02:00
Matthew Honnibal	45daf5c9fe	Add init labels command	2020-09-29 16:22:37 +02:00
Matthew Honnibal	58c8d4b414	Add label_data property to pipeline	2020-09-29 16:22:13 +02:00
Ines Montani	aa2a6882d0	Fix logging	2020-09-29 16:08:39 +02:00
Ines Montani	63d1598137	Simplify config use in Language.initialize	2020-09-29 16:05:48 +02:00
Ines Montani	56f8bc73ef	Add more tests	2020-09-29 15:23:34 +02:00
Sofie Van Landeghem	6a04e5adea	encoding UTF8 (#6161 )	2020-09-29 14:49:55 +02:00
Ines Montani	591038b1a4	Add test	2020-09-29 12:54:52 +02:00
Ines Montani	adca08a12f	Pass nlp forward	2020-09-29 12:21:52 +02:00
Ines Montani	f171903139	Clean up sgd and pipeline -> nlp	2020-09-29 12:20:26 +02:00
Ines Montani	612bbf85ab	Update initialize.py	2020-09-29 12:14:47 +02:00
Ines Montani	42f0e4c946	Clean up	2020-09-29 12:14:08 +02:00
Matthew Honnibal	9c8b2524fe	Upd initialize args	2020-09-29 12:08:37 +02:00
Matthew Honnibal	e1fdf2b7c5	Upd tests	2020-09-29 12:05:38 +02:00
Ines Montani	50410c17ac	Update schemas.py	2020-09-29 12:05:38 +02:00
Matthew Honnibal	f2d1b7feb5	Clean up sgd	2020-09-29 12:00:08 +02:00
Ines Montani	78396d137f	Integrate initialize settings	2020-09-29 11:57:08 +02:00
Ines Montani	dec984a9c1	Update Language.initialize and support components/tokenizer settings	2020-09-29 11:52:45 +02:00
Matthew Honnibal	b3b6868639	Remove 'sgd' arg from component initialize	2020-09-29 11:42:35 +02:00
Matthew Honnibal	5276db6f3f	Remove 'device' argument from Language, clean up 'sgd' arg	2020-09-29 11:42:19 +02:00
Ines Montani	4925ad760a	Add init vectors	2020-09-29 10:58:50 +02:00
svlandeg	64d90039a1	encoding UTF8	2020-09-29 10:54:42 +02:00
Ines Montani	ff9a63bfbd	begin_training -> initialize	2020-09-28 21:35:09 +02:00
Ines Montani	046f655d86	Fix error	2020-09-28 21:17:45 +02:00
Ines Montani	a139fe672b	Fix typos and refactor CLI logging	2020-09-28 21:17:10 +02:00
Ines Montani	2e9c9e74af	Fix config resolution and interpolation TODO: auto-interpolate in Thinc if config is dict (i.e. likely subsection)	2020-09-28 15:34:00 +02:00
Ines Montani	02838a1d47	Fix resolve_dot_names	2020-09-28 15:27:10 +02:00
Ines Montani	822ea4ef61	Refactor CLI	2020-09-28 15:09:59 +02:00
Ines Montani	a89e0ff7cb	Fix typo	2020-09-28 12:55:21 +02:00
Ines Montani	a62337b3f3	Tidy up vocab init	2020-09-28 12:53:06 +02:00
Ines Montani	c22ecc66bb	Don't support init path for now	2020-09-28 12:46:28 +02:00
Ines Montani	f49288ab81	Update default_config_pretraining.cfg	2020-09-28 12:31:54 +02:00
Ines Montani	a5f2cc0509	Tidy up and remove raw text (rehearsal) for now	2020-09-28 12:30:13 +02:00
Ines Montani	1590de11b1	Update config	2020-09-28 12:05:23 +02:00
Matthew Honnibal	9f6ad06452	Upd default config	2020-09-28 12:00:23 +02:00
Ines Montani	e44a7519cd	Update CLI and add [initialize] block	2020-09-28 11:56:14 +02:00
Ines Montani	d5155376fd	Update vocab init	2020-09-28 11:30:18 +02:00
Ines Montani	8b74fd19df	init pipeline -> init nlp	2020-09-28 11:13:38 +02:00
Ines Montani	2fdb7285a0	Update CLI	2020-09-28 11:06:07 +02:00
Ines Montani	553bfea641	Fix commands	2020-09-28 10:53:17 +02:00
Matthew Honnibal	44bad1474c	Add init_pipeline file	2020-09-28 09:47:34 +02:00
Matthew Honnibal	65448b2e34	Remove schema=None until Optional	2020-09-28 03:42:58 +02:00
Matthew Honnibal	b886f53c31	init-pipeline runs (maybe doesnt work)	2020-09-28 03:42:47 +02:00
Matthew Honnibal	ed2aff2db3	Remove unused train code	2020-09-28 03:12:31 +02:00
Matthew Honnibal	3a0a3b8db6	Dont hard-code for 'corpora' name	2020-09-28 03:06:33 +02:00
Matthew Honnibal	a023cf3ecc	Add (untested) resolve_dot_names util	2020-09-28 03:06:12 +02:00
Matthew Honnibal	a976da168c	Support data augmentation in Corpus (#6155 ) * Support data augmentation in Corpus * Note initial docs for data augmentation * Add augmenter to quickstart * Fix flake8 * Format * Fix test * Update spacy/tests/training/test_training.py * Improve data augmentation arguments * Update templates * Move randomization out into caller * Refactor * Update spacy/training/augment.py * Update spacy/tests/training/test_training.py * Fix augment * Fix test	2020-09-28 03:03:27 +02:00
Matthew Honnibal	13b1605ee6	Add init script	2020-09-28 01:08:49 +02:00
Matthew Honnibal	a3e1791c9c	Upd train	2020-09-28 01:08:30 +02:00
Matthew Honnibal	b5556093e2	Start updating train script	2020-09-27 23:59:44 +02:00
Ines Montani	9016d23cc5	Fix exclude and add test	2020-09-27 23:34:03 +02:00
Ines Montani	658fad428a	Fix base schema integration	2020-09-27 22:50:36 +02:00
Ines Montani	e04bd16f7f	Merge branch 'develop' into feature/new-thinc-config-resolution	2020-09-27 22:34:46 +02:00
Ines Montani	d7ad65a9bb	Fix handling of error description [ci skip]	2020-09-27 22:31:57 +02:00
Ines Montani	7e938ed63e	Update config resolution to use new Thinc	2020-09-27 22:21:31 +02:00
Adriane Boyd	013b66de05	Add tokenizer scoring to ja / ko / zh (#6152 )	2020-09-27 22:20:45 +02:00
Adriane Boyd	a6548ead17	Add _ as a symbol (#6153 ) * Add _ to StringStore in Morphology * Add _ as a symbol Add `_` as a symbol instead of adding to the `StringStore`.	2020-09-27 22:20:14 +02:00
Matthew Honnibal	39b178999c	Tmp notes	2020-09-27 20:13:38 +02:00
Adriane Boyd	8393dbedad	Minor fixes * Put `cfg` back in serialization * Add `pickle5` to pytest conf	2020-09-27 15:15:53 +02:00
Adriane Boyd	54fe871935	Fix formatting, refactor pickle5 exceptions	2020-09-27 14:37:28 +02:00
Adriane Boyd	11e195d3ed	Update ChineseTokenizer * Allow `pkuseg_model` to be set to `None` on initialization * Don't save config within tokenizer * Force convert pkuseg_model to use pickle protocol 4 by reencoding with `pickle5` on serialization * Update pkuseg serialization test	2020-09-27 14:00:18 +02:00
Ines Montani	b4486d747d	Merge branch 'develop' into fix/train-config-interpolation	2020-09-26 15:32:14 +02:00
Ines Montani	8fea06d55e	Merge pull request #6149 from adrianeboyd/feature/attributeruler-match-ids Simplify string match IDs for AttributeRuler	2020-09-26 15:31:30 +02:00
Ines Montani	b2d07de786	Construct nlp from uninterpolated config before training	2020-09-26 15:16:59 +02:00
Ines Montani	ca3c997062	Improve CLI config validation with latest Thinc	2020-09-26 13:13:57 +02:00
Adriane Boyd	6c25e60089	Simplify string match IDs for AttributeRuler	2020-09-26 11:12:39 +02:00
Matthew Honnibal	702edf52a0	Fix attributeruler	2020-09-26 00:30:48 +02:00
Matthew Honnibal	821f37254c	Fix attributeruler	2020-09-26 00:19:53 +02:00
Matthew Honnibal	98327f66a9	Fix attributeruler key	2020-09-25 23:20:50 +02:00
Matthew Honnibal	092ce4648e	Make DocBin output stable data (set iteration)	2020-09-25 22:20:44 +02:00
Matthew Honnibal	26afd3bd90	Fix iteration order	2020-09-25 21:47:22 +02:00

... 5 6 7 8 9 ...

8520 Commits