spaCy

mirror of https://github.com/explosion/spaCy.git synced 2026-02-05 23:09:48 +03:00

Author	SHA1	Message	Date
Koichi Yasuoka	0afb54ac93	JapaneseTokenizer.pipe added (#6515 ) * JapaneseTokenizer.pipe added For [spacymoji](https://spacy.io/universe/project/spacymoji) with `Japanese()`. * DummyTokenizer.pipe added instead	2020-12-08 20:02:23 +01:00
Adriane Boyd	df4891bed1	Remove blis python version constraints (#6522 ) * Remove blis version constraints After updating the blis sdist in v0.7.4, remove python version constraints for blis build and install dependencies. * Install sdist with --prefer-binary for python 3.5 * Fix duplicate sdist install steps * Fix sdist install step types * Fix blis pins in requirements.txt * Remove wheel hack for python 3.5 from CI	2020-12-08 15:25:19 +01:00
Ines Montani	4e77349106	Merge pull request #6524 from adrianeboyd/bugfix/entity-ruler-subsequent Fix subsequent pipe detection in EntityRuler	2020-12-08 22:17:28 +11:00
Adriane Boyd	6c221d4841	Fix subsequent pipe detection in EntityRuler Fix subsequent pipe detection to detect the position of the current object by comparing the component itself rather than from the factory name.	2020-12-08 10:01:30 +01:00
Ines Montani	b87793a89a	Merge pull request #6523 from adrianeboyd/bugfix/remove-use-chars Remove non-working --use-chars from train CLI	2020-12-08 09:30:48 +01:00
Adriane Boyd	5ceac425ee	Remove non-working --use-chars from train CLI Remove the non-working `--use-chars` option from the train CLI. The implementation of the option across component types and the CLI settings could be fixed, but the `CharacterEmbed` model does not work on GPU in v2 so it's better to remove it.	2020-12-08 08:30:00 +01:00
Adriane Boyd	dcecc75270	Improve blis and numpy build dependencies (#6455 ) * Fix blis build dependencies * Add blis with python_version constraints to pyproject.toml * Add blis to setup_requires * Remove --only-binary from CI * Reduce number of builds to speed up CI * Add hack to install wheel for python 3.5 in linux * Remove os spec from CI * Remove detailed numpy build constraints * Remove detailed numpy build constraints from `pyproject.toml` because it is too difficult to maintain for many architectures * These constraints are more a reflection of what is available on pypi as binary wheels rather than any real build requirements that it is necessary for users to follow when building from source * Users building their own binary packages will need to enforce the constraints that make sense in their environments, e.g., the `conda` compatible numpy pins * Keep the build constraints in `build-constraints.txt` for use with our builds * Our builds with wheelwright are built against the earliest compatible binary versions of numpy on pypi * These constraints are documented within the distribution * Revert "Remove os spec from CI" This reverts commit `7489476688`.	2020-12-08 14:29:34 +08:00
Adriane Boyd	e931d3f72b	Move max_length to nlp.make_doc() (#6512 ) Move max_length check to `nlp.make_doc()` so that's it's also checked for `nlp.pipe()`.	2020-12-08 14:24:02 +08:00
Sofie Van Landeghem	52fa46dd58	tested EL scripts with 2.3.4 (#6517 )	2020-12-07 20:46:38 +01:00
Adriane Boyd	53c0fb7431	Only set NORM on Token in retokenizer (#6464 ) * Only set NORM on Token in retokenizer Instead of setting `NORM` on both the token and lexeme, set `NORM` only on the token. The retokenizer tries to set all possible attributes with `Token/Lexeme.set_struct_attr` so that it doesn't have to enumerate which attributes are available for each. `NORM` is the only attribute that's stored on both and for most cases it doesn't make sense to set the global norms based on a individual retokenization. For lexeme-only attributes like `IS_STOP` there's no way to avoid the global side effects, but I think that `NORM` would be better only on the token. * Fix test	2020-11-30 09:35:42 +08:00
Adriane Boyd	03ae77e603	Add SPACY as a Matcher attribute (#6463 )	2020-11-30 09:34:50 +08:00
Adriane Boyd	3a5cc5f8b4	Set version to v2.3.4	2020-11-26 08:48:52 +01:00
Adriane Boyd	e0f5646a4a	Restore cleanup_beam method (#6446 )	2020-11-25 13:21:48 +01:00
Jacob Bortell	fe9009911a	Update rule-based-matching.md (#6421 ) * Update rule-based-matching.md Clarified case-sensititivy of dictionary-referencing attributes (POS/TAG/DEP/etc). Clarified "Type" column header to "Value Type" * Update rule-based-matching.md Improved clarity of wording	2020-11-24 16:20:19 +01:00
Jacob Bortell	992723dfac	Add jabortell to the contributors (#6422 ) * Add jabortell to the contributors * Update jabortell.md Added tick to applicable statement	2020-11-24 16:15:31 +01:00
Adriane Boyd	afd744bc05	Update Travis CI pip install steps (#6440 )	2020-11-24 14:10:16 +01:00
Adriane Boyd	573f5c863f	Fix tag map clobbering in spacy train (#6437 ) Fix bug from #5768 where the tag map is clobbered if a custom tag map isn't provided.	2020-11-24 13:13:16 +01:00
Adriane Boyd	ce18fc6588	Set version to v2.3.3	2020-11-24 10:03:45 +01:00
Adriane Boyd	cd61d264ef	Set version to v2.3.3.dev0	2020-11-23 13:51:59 +01:00
Sofie Van Landeghem	2af31a8c8d	Bugfix textcat reproducibility on GPU (#6411 ) * add seed argument to ParametricAttention layer * bump thinc to 7.4.3 * set thinc version range Co-authored-by: Adriane Boyd <adrianeboyd@gmail.com>	2020-11-23 12:29:35 +01:00
Adriane Boyd	cdca44ac11	Dynamically include numpy headers (#6418 ) * Dynamically include numpy headers * Add `build-constraints.txt` with numpy version pins for building wheels with `pip` and `wheelwright` * Update `setup.py` to add current numpy include directory * Assume `cython` and `numpy` are installed for `setup.py` * Remove included numpy headers * Fix typo in requirements.txt * Use script in CI	2020-11-23 11:15:11 +01:00
Adriane Boyd	3f61f5eb54	Use int8_t instead of char in Matcher (#6413 ) * Use signed char instead of char in Matcher Remove unused char* utf8_t typedef * Use int8_t instead of signed char	2020-11-23 10:26:47 +01:00
Adriane Boyd	4284605683	Remove Beam cleanup (#6414 ) Beam cleanup is handled through the Beam finalization method.	2020-11-23 10:01:46 +01:00
Adriane Boyd	a8c2dad466	Add all vectors to vocab before pruning (#6408 ) Add all vectors to the vocab before pruning to correct the selection of vectors to prioritize.	2020-11-23 10:00:59 +01:00
Adriane Boyd	13f0676f04	Updates for python 3.9 (#6338 ) * Update blis and thinc version ranges * Update thinc version range * Update setup.cfg for python 3.9 * Adjust blis and thinc ranges * Add python 3.9 classifier * Update CI for python 3.9 * Add --prefer-binary to CI sdist install * Update CI python 3.7 mac image * Add --prefer-binary to Travis CI * Update install instructions in README * Specify blis versions separately for < / >= 3.6 * Update --prefer-binary in README * Test cleaner sdist install * Also upgrade pip (This is kind of unnecessary given --prefer-binary but may avoid other issues related to sdist installs in the future.) * Compile with -j 2 * Remove wheel from setup_requires * Update to have separate CI uninstall step * Remove wheel from pyproject.toml * Recommend upgrading setuptools in addition to pip	2020-11-23 09:45:18 +01:00
Yusuke Mori	e3ac90b035	Avoid a SyntaxError in self-attentive-parser (#6428 ) * Avoid a SyntaxError in self-attentive-parser Fix a usage of quotation marks in the example of spaCy Universe self-attentive-parser * Create forest1988.md Fill in the spaCy contributor agreement	2020-11-22 21:59:37 +01:00
M. Revuelta Espinosa	51232ffb9e	Update universe.json (include PatternOmatic) (#6399 ) Request to include PatternOmatic in spaCy Universe Adds @revuel to contributors	2020-11-19 13:15:50 +01:00
Adriane Boyd	3cf6479467	Fix JSON in #6395	2020-11-17 15:25:41 +01:00
Sam Edwardes	78913a4f95	Added spaCyTextBlob to universe.json (#6395 )	2020-11-17 14:38:34 +01:00
Adriane Boyd	320a8b1481	Add ent_id_ to strings serialized with Doc (#6353 )	2020-11-10 20:16:07 +08:00
Ines Montani	d490428089	Update README.md [ci skip]	2020-11-10 09:51:20 +08:00
Ines Montani	4d337eedf2	Merge pull request #6322 from medspacy/master	2020-11-10 02:47:29 +01:00
Adriane Boyd	90550552a0	CI updates for python 3.5 (#6354 ) * Update pip in CI * Use --prefer-binary * Use `--prefer-binary` * Delete all installed packages before testing source install * sdist install with --only-binary :all:	2020-11-06 13:35:51 +01:00
Daniel Vasic	20d72de986	Added Multext-East V5 tagset for Croatian language (#6248 ) * Added Multext-East V5 tagset for Croatian language * Create danielvasic.md * Update danielvasic.md * Update danielvasic.md * Add tag map to CroatianDefaults Co-authored-by: Adriane Boyd <adrianeboyd@gmail.com>	2020-11-05 12:19:22 +01:00
Robert Šípek	6069efe57d	Add tag map to cs language (#6284 )	2020-11-05 10:13:11 +01:00
Adriane Boyd	8644ee3e3f	Update TIGER link and tag description (#6344 )	2020-11-05 09:33:00 +01:00
Vu Ha	6d465ec52c	add oprd to the list of accepted deps for noun chunking (#6302 ) * add oprd to the list of accepted deps for noun chunking * add SCA	2020-11-05 09:17:35 +01:00
Adriane Boyd	31de700b0f	Fix on_match callback and remove empty patterns (#6312 ) For the `DependencyMatcher`: * Fix on_match callback so that it is called once per matched pattern * Fix results so that patterns with empty match lists are not returned	2020-11-05 09:16:26 +01:00
Alec Chapman	204c7c8a00	fix thumbnail link to be github raw url	2020-11-01 07:53:48 -07:00
Adriane Boyd	45c9a68828	Identify final Matcher pattern node by quantifier (#6317 ) Modify the internal pattern representation in `Matcher` patterns to identify the final ID state using a unique quantifier rather than a combination of other attributes. It was insufficient to identify the final ID node based on an uninitialized `quantifier` (coincidentally being the same as the `ZERO`) with `nr_attr` as 0. (In addition, it was potentially bug-prone that `nr_attr` was set to 0 even though attrs were allocated.) In the case of `{"OP": "!"}` (a valid, if pointless, pattern), `nr_attr` is 0 and the quantifier is ZERO, so the previous methods for incrementing to the ID node at the end of the pattern weren't able to distinguish the final ID node from the `{"OP": "!"}` pattern.	2020-10-31 12:18:48 +01:00
Alec Chapman	73d22d96ff	add medspacy to universe and fix example w/ cov-bsv	2020-10-29 07:53:56 -06:00
Duygu Altinok	0e55f806dd	Turkish tokenization improvements (#6268 ) * added single and paired orth variants * added token match * added long text tokenization test * inverted init * normalized lemmas to lowercase * more abbrevs * tests for ordinals and abbrevs * separated period abbvrevs to another list * fiex typo * added ordinal and abbrev tests * added number tests for dates * minor refinement * added inflected abbrevs regex * added percentage and inflection * cosmetics * added token match * added url inflection tests * excluded url tokens from custom pattern * removed url match import	2020-10-29 09:43:17 +01:00
Adriane Boyd	8cc5ed6771	Add Macedonian to website languages	2020-10-29 08:49:56 +01:00
Ines Montani	1e4d7e059f	Revert "Test FUNDING.yml [ci skip]" This reverts commit `287be48ad0`.	2020-10-28 17:42:23 +01:00
Ines Montani	287be48ad0	Test FUNDING.yml [ci skip]	2020-10-28 17:36:25 +01:00
Adriane Boyd	4dd86306e9	Add Nepali to supported languages on website (#6315 )	2020-10-28 16:32:07 +01:00
Robert Šípek	260c29794a	Fill contributor agreement by robertsipek (#6285 ) * Fill contributor agreement by robertsipek * Fill contributor agreement by robertsipek	2020-10-22 22:13:17 +02:00
Kunal Sharma	01aec7a313	Adding MindMeld to Universe JSON (#6275 ) * Adding Mindmeld to Universe JSON Mindmeld is a conversational AI platform for deep-domain voice interfaces and chatbots. https://www.mindmeld.com/ * Signing contribution agreement. Co-authored-by: kunshar2 <kunshar2@cisco.com>	2020-10-21 18:42:11 +02:00
Ines Montani	d7a4e8454b	Merge pull request #6274 from walterhenry/master User contributor agreement	2020-10-19 16:30:58 +02:00
walterhenry	ff82644746	User contributor agreement Here it is!	2020-10-19 16:25:09 +02:00

1 2 3 4 5 ...

11687 Commits