spaCy

mirror of https://github.com/explosion/spaCy.git synced 2024-12-29 03:16:31 +03:00

Author	SHA1	Message	Date
Ines Montani	4a1029a9b6	Add infobox [ci skip]	2021-01-19 19:18:39 +11:00
Sofie Van Landeghem	fed8f48965	raise NotImplementedError when noun_chunks iterator is not implemented (#6711 ) * raise NotImplementedError when noun_chunks iterator is not implemented * bring back, fix and document span.noun_chunks * formatting Co-authored-by: Matthew Honnibal <honnibal+gh@gmail.com>	2021-01-17 19:56:05 +08:00
Adriane Boyd	bf0cdae8d4	Add token_splitter component (#6726 ) * Add long_token_splitter component Add a `long_token_splitter` component for use with transformer pipelines. This component splits up long tokens like URLs into smaller tokens. This is particularly relevant for pretrained pipelines with `strided_spans`, since the user can't change the length of the span `window` and may not wish to preprocess the input texts. The `long_token_splitter` splits tokens that are at least `long_token_length` tokens long into smaller tokens of `split_length` size. Notes: * Since this is intended for use as the first component in a pipeline, the token splitter does not try to preserve any token annotation. * API docs to come when the API is stable. * Adjust API, add test * Fix name in factory	2021-01-17 19:54:41 +08:00
Matthew Honnibal	f277bfdf0f	Add SpanGroup and Graph container types to represent arbitrary annotations (#6696 ) * Draft out initial Spans data structure * Initial span group commit * Basic span group support on Doc * Basic test for span group * Compile span_group.pyx * Draft addition of SpanGroup to DocBin * Add deserialization for SpanGroup * Add tests for serializing SpanGroup * Fix serialization of SpanGroup * Add EdgeC and GraphC structs * Add draft Graph data structure * Compile graph * More work on Graph * Update GraphC * Upd graph * Fix walk functions * Let Graph take nodes and edges on construction * Fix walking and getting * Add graph tests * Fix import * Add module with the SpanGroups dict thingy * Update test * Rename 'span_groups' attribute * Try to fix c++11 compilation * Fix test * Update DocBin * Try to fix compilation * Try to fix graph * Improve SpanGroup docstrings * Add doc.spans to documentation * Fix serialization * Tidy up and add docs * Update docs [ci skip] * Add SpanGroup.has_overlap * WIP updated Graph API * Start testing new Graph API * Update Graph tests * Update Graph * Add docstring Co-authored-by: Ines Montani <ines@ines.io>	2021-01-14 17:30:41 +11:00
Adriane Boyd	a45d89f09a	Add initialize.before_init and after_init callbacks Add `initialize.before_init` and `initialize.after_init` callbacks to the config. The `initialize.before_init` callback is a place to implement one-time tokenizer customizations that are then saved with the model.	2021-01-12 13:07:44 +01:00
Sofie Van Landeghem	a612a5ba3f	fix small typos (#6698 )	2021-01-08 09:39:47 +01:00
Sofie Van Landeghem	75d9019343	Fix types of Tok2Vec encoding architectures (#6442 ) * fix TorchBiLSTMEncoder documentation * ensure the types of the encoding Tok2vec layers are correct * update references from v1 to v2 for the new architectures	2021-01-07 16:39:27 +11:00
Sofie Van Landeghem	82ae95267a	Docs for pretrain architectures (#6605 ) * document pretraining architectures * formatting * bit more info * small fixes	2021-01-06 16:12:30 +11:00
Sofie Van Landeghem	afc5714d32	multi-label textcat component (#6474 ) * multi-label textcat component * formatting * fix comment * cleanup * fix from #6481 * random edit to push the tests * add explicit error when textcat is called with multi-label gold data * fix error nr * small fix	2021-01-06 13:07:14 +11:00
Ines Montani	85ca8c2bdd	Merge branch 'master' into develop	2020-12-11 13:44:41 +11:00
Ines Montani	fb43a30a71	Merge pull request #6545 from svlandeg/feature/discussions [ci skip]	2020-12-11 10:20:35 +11:00
svlandeg	5afa567767	replace gitter with discussions in 101	2020-12-10 20:17:36 +01:00
Adriane Boyd	27bb75e2a0	Docs and extras updates for v2.3.5 * Update install instructions for updated packages * Add `cuda110` and `cuda111` extras, remove upper `cupy` pins (only compatible with `thinc>=7.4.4`)	2020-12-10 15:34:34 +01:00
Ines Montani	513c4e332a	Include custom code via spacy package command (#6531 )	2020-12-10 20:36:46 +08:00
Ines Montani	1980203229	Merge branch 'master' into pr/6444	2020-12-09 11:09:40 +11:00
Ines Montani	05a2812ae0	Merge branch 'develop' into pr/6444	2020-12-09 11:04:03 +11:00
Ines Montani	8921364579	Merge pull request #6521 from explosion/feature/config-stdin Allow reading config from stdin in spacy train	2020-12-08 22:07:43 +11:00
Ines Montani	94a5a9814f	Update argument handling and documentation	2020-12-08 20:41:18 +11:00
Ines Montani	ef59ce783b	Adjust install instructions [ci skip]	2020-12-08 18:06:50 +11:00
Ines Montani	d8e01ca931	Merge pull request #6391 from adrianeboyd/docs/install-guide	2020-12-08 07:42:16 +01:00
Ines Montani	c2b196c2c1	Merge pull request #6419 from svlandeg/feature/rel-docs	2020-12-08 06:30:41 +01:00
Adriane Boyd	1442d2f213	Improve simple training example in v3 migration (#6438 ) * Create the examples once * Use the examples in the initialization * Provide the batch size * Fix `begin_training` migration example	2020-11-30 09:39:45 +08:00
Adriane Boyd	03ae77e603	Add SPACY as a Matcher attribute (#6463 )	2020-11-30 09:34:50 +08:00
Adriane Boyd	724831b066	Merge remote-tracking branch 'upstream/master' into chore/update-develop-from-master * Update Macedonian for v3 * Update Turkish for v3	2020-11-25 11:49:34 +01:00
Jacob Bortell	fe9009911a	Update rule-based-matching.md (#6421 ) * Update rule-based-matching.md Clarified case-sensititivy of dictionary-referencing attributes (POS/TAG/DEP/etc). Clarified "Type" column header to "Value Type" * Update rule-based-matching.md Improved clarity of wording	2020-11-24 16:20:19 +01:00
Adriane Boyd	6f133877aa	Update source install instructions * Don't recommend an editable install in the default source instructions. * Use `pip install --no-build-isolation` for editable installs. * Remove reference to `virtualenv`.	2020-11-24 14:44:13 +01:00
svlandeg	218abaa69a	typo	2020-11-20 22:36:49 +01:00
svlandeg	e861e928df	more small corrections	2020-11-20 22:29:58 +01:00
svlandeg	5ac0867427	final fixes	2020-11-20 22:18:53 +01:00
svlandeg	331ec83493	edits and updates to implementing REL component docs	2020-11-20 21:41:52 +01:00
svlandeg	4a3e611abc	small fixes and formatting	2020-11-20 15:55:05 +01:00
svlandeg	124f49feb6	update REL model code	2020-11-20 15:25:20 +01:00
Adriane Boyd	96726ec1f6	Fix DocBin init in training example (#6396 )	2020-11-17 14:36:44 +01:00
Adriane Boyd	ed32fa80cd	Update source install instructions * Use `pip install` instead of `python setup.py install` * For developers recommend: * `python setup.py build_ext --inplace -j N` * `python setup.py develop`	2020-11-16 10:13:51 +01:00
svlandeg	99d0412b6e	add link to REL project	2020-11-15 18:35:56 +01:00
Ines Montani	de6453940e	Merge pull request #6305 from svlandeg/feature/score-docs [ci skip]	2020-11-10 02:52:11 +01:00
Ines Montani	d7950c5ada	Merge pull request #6297 from adrianeboyd/docs/nightly-conda-install [ci skip]	2020-11-10 02:45:52 +01:00
Ines Montani	363ac73c72	Update docs [ci skip]	2020-11-09 12:43:26 +08:00
Ines Montani	019a1dd5e8	Fix v3 overview [ci skip]	2020-11-03 18:10:06 +01:00
Adriane Boyd	dc816bba9d	Fix node name typo in dependency matcher example (#6311 )	2020-10-28 16:32:46 +01:00
svlandeg	77688b0072	fix config	2020-10-26 11:14:34 +01:00
svlandeg	5878ff6bcd	cleanup	2020-10-26 11:13:02 +01:00
svlandeg	e95d9caa87	small edits	2020-10-26 11:09:25 +01:00
svlandeg	a664994a81	adding score method to explanation of new component	2020-10-26 10:52:47 +01:00
Adriane Boyd	c0b76f4c19	Add install step to "Compile from source"	2020-10-23 11:36:36 +02:00
Ines Montani	b6b1c1e23c	Merge pull request #6271 from walterhenry/develop-proof [ci skip]	2020-10-19 16:31:43 +02:00
walterhenry	db24dc5614	Proofread remarks I think these may the last remarks for the nightly docs. Only two minor things actually.	2020-10-19 11:11:32 +02:00
Sofie Van Landeghem	75a202ce65	TextCat updates and fixes (#6263 ) * small fix in example imports * throw error when train_corpus or dev_corpus is not a string * small fix in custom logger example * limit macro_auc to labels with 2 annotations * fix typo * also create parents of output_dir if need be * update documentation of textcat scores * refactor TextCatEnsemble * fix tests for new AUC definition * bump to 3.0.0a42 * update docs * rename to spacy.TextCatEnsemble.v2 * spacy.TextCatEnsemble.v1 in legacy * cleanup * small fix * update to 3.0.0rc2 * fix import that got lost in merge * cursed IDE * fix two typos	2020-10-18 14:50:41 +02:00
Ines Montani	c655742b8b	Remove docs references to starters for now (see #6262 ) [ci skip]	2020-10-16 15:46:34 +02:00
Ines Montani	c968d1560f	Fix docs example [ci skip]	2020-10-16 11:33:20 +02:00
Ines Montani	ba1e004049	Fix typo [ci skip]	2020-10-15 23:39:04 +02:00
Ines Montani	20f80587d6	Merge pull request #6257 from walterhenry/develop-proof A few tiny typo fixes to push through with release of nightly	2020-10-15 18:17:30 +02:00
walterhenry	75b7f86383	Three small typos Some little typos since v3.0 is out.	2020-10-15 18:06:37 +02:00
Ines Montani	09dbbe75d7	Update docs [ci skip]	2020-10-15 17:27:24 +02:00
Ines Montani	7f05ccc170	Update docs [ci skip]	2020-10-15 12:35:30 +02:00
Ines Montani	4fa869e6f7	Update docs [ci skip]	2020-10-15 11:16:06 +02:00
Ines Montani	178760855f	Merge branch 'develop' into master-tmp	2020-10-15 09:06:03 +02:00
Ines Montani	abeafcbc08	Update docs [ci skip]	2020-10-15 08:58:30 +02:00
Ines Montani	a2d4aaee70	Apply suggestions from code review	2020-10-14 19:51:36 +02:00
Ines Montani	d94e241fce	Merge branch 'develop' into pr/6253	2020-10-14 16:55:46 +02:00
Ines Montani	cb47f25cda	Merge pull request #6252 from svlandeg/fix/docs	2020-10-14 16:43:12 +02:00
walterhenry	6af585dba5	New batch of proofs Just tiny fixes to the docs as a proofreader	2020-10-14 16:37:57 +02:00
svlandeg	478a14a619	fix few typos	2020-10-14 15:01:19 +02:00
Ines Montani	1aa8e8f2af	Update docs [ci skip]	2020-10-14 14:58:45 +02:00
svlandeg	08cb085f6c	Merge remote-tracking branch 'upstream/develop' into fix/various	2020-10-09 17:01:27 +02:00
Ines Montani	97ff090e49	Fix docs example [ci skip]	2020-10-09 16:03:57 +02:00
Ines Montani	9fb3244672	Merge pull request #6231 from adrianeboyd/feature/include-static-vectors	2020-10-09 15:54:52 +02:00
Adriane Boyd	2dd79454af	Update docs	2020-10-09 14:42:07 +02:00
svlandeg	853edace37	fix MultiHashEmbed example in documentation	2020-10-09 14:11:06 +02:00
Ines Montani	e50dc2c1c9	Update docs [ci skip]	2020-10-09 12:04:52 +02:00
Ines Montani	7c52def5da	Merge pull request #6227 from adrianeboyd/chore/update-3.0.0a36-from-master	2020-10-09 10:49:20 +02:00
Ines Montani	329b61ee7b	Update docs [ci skip]	2020-10-09 10:36:06 +02:00
delzac	668507be1b	Reflect on usage doc that IS_SENT_START attribute exist (#6114 ) * Reflect on usage doc that IS_SENT_START attribute exist * Create delzac.md	2020-10-09 10:14:40 +02:00
Sofie Van Landeghem	d093d6343b	TrainablePipe (#6213 ) * rename Pipe to TrainablePipe * split functionality between Pipe and TrainablePipe * remove unnecessary methods from certain components * cleanup * hasattr(component, "pipe") should be sufficient again * remove serialization and vocab/cfg from Pipe * unify _ensure_examples and validate_examples * small fixes * hasattr checks for self.cfg and self.vocab * make is_resizable and is_trainable properties * serialize strings.json instead of vocab * fix KB IO + tests * fix typos * more typos * _added_strings as a set * few more tests specifically for _added_strings field * bump to 3.0.0a36	2020-10-08 21:33:49 +02:00
Ines Montani	5ebd1fc2cf	Update docs [ci skip]	2020-10-08 16:23:12 +02:00
Ines Montani	d1602e1ece	Update docs [ci skip]	2020-10-08 11:56:50 +02:00
Ines Montani	064575d79d	Merge pull request #6216 from svlandeg/feature/nel-initialize	2020-10-08 11:14:12 +02:00
Ines Montani	43e59bb22a	Update docs and install extras [ci skip]	2020-10-08 10:58:50 +02:00
svlandeg	bcaad28eda	fix typos	2020-10-07 13:05:37 +02:00
delzac	15ea401b39	Reflect on usage doc that IS_SENT_START attribute exist (#6114 ) * Reflect on usage doc that IS_SENT_START attribute exist * Create delzac.md	2020-10-06 15:11:01 +02:00
Ines Montani	ce14520789	Update docs [ci skip]	2020-10-06 14:35:17 +02:00
Ines Montani	2a17566da3	Update docs [ci skip]	2020-10-06 14:15:08 +02:00
Ines Montani	967377287a	Merge pull request #6210 from adrianeboyd/docs/various-v3-3 [ci skip]	2020-10-06 11:28:45 +02:00
Adriane Boyd	aa9c9f3bf0	Update Chinese usage for spacy-pkuseg	2020-10-06 11:21:17 +02:00
Ines Montani	2e961817cb	Update docs [ci skip]	2020-10-06 10:23:01 +02:00
svlandeg	fd0f60e2bc	updates to data format for training and pretraining	2020-10-06 09:28:53 +02:00
Ines Montani	706b7f6973	Update docs	2020-10-05 20:51:22 +02:00
Ines Montani	e3acad6264	Update docs [ci skip]	2020-10-05 13:06:20 +02:00
Ines Montani	0f64556c04	Merge pull request #6197 from svlandeg/feature/pipe-docs [ci skip]	2020-10-05 11:55:40 +02:00
svlandeg	9a6c9b133b	various small fixes	2020-10-05 01:05:37 +02:00
svlandeg	52b660e9dc	initialize and update explanation	2020-10-05 00:39:36 +02:00
svlandeg	b0463fbf75	set_annotations explanation	2020-10-04 14:56:48 +02:00
Ines Montani	43d7652635	Merge pull request #6192 from explosion/feature/init-attr-ruler	2020-10-04 14:46:37 +02:00
Ines Montani	9b3a934361	Update docs [ci skip]	2020-10-04 14:14:55 +02:00
svlandeg	9f40d963fd	highlight the two steps: the model and the pipeline component	2020-10-04 14:11:53 +02:00
Ines Montani	11347f34da	Tidy up, tests and docs	2020-10-04 13:54:05 +02:00
svlandeg	452b8309f9	slight rewrite to hide some thinc implementation details	2020-10-04 13:26:46 +02:00
svlandeg	08ad349a18	tok2vec layer	2020-10-04 00:08:02 +02:00
svlandeg	2c4b2ee5e9	REL intro and get_candidates function	2020-10-03 23:27:05 +02:00
Ines Montani	3b8f352eda	Merge branch 'develop' of https://github.com/explosion/spaCy into develop	2020-10-03 16:08:27 +02:00

1 2 3 4 5 ...

932 Commits