spaCy

mirror of https://github.com/explosion/spaCy.git synced 2025-07-24 23:19:45 +03:00

Author	SHA1	Message	Date
Matthew Honnibal	2a5a61683e	Add function to get train format from Doc objects Our JSON training format is annoying to work with, and we've wanted to retire it for some time. In the meantime, we can at least add some missing functions to make it easier to live with. This patch adds a function that generates the JSON format from a list of Doc objects, one per paragraph. This should be a convenient way to handle a lot of data conversions: whatever format you have the source information in, you can use it to setup a Doc object. This approach should offer better future-proofing as well. Hopefully, we can steadily rewrite code that is sensitive to the current data-format, so that it instead goes through this function. Then when we change the data format, we won't have such a problem.	2018-08-14 13:13:10 +02:00
Matthew Honnibal	4336397ecb	Update develop from master	2018-08-14 03:04:28 +02:00
Matthew Honnibal	13fa550b36	Merge branch 'master' of https://github.com/explosion/spaCy	2018-08-14 02:32:01 +02:00
Ioannis Daras	fe94e696d3	Optimize Greek language support (#2658 )	2018-08-14 02:31:32 +02:00
Wojciech Łukasiewicz	3953e967a0	User correct variable name in the examples (#2664 ) * correct naming * add contributor agreement	2018-08-13 22:21:24 +02:00
Matthew Honnibal	85000ea13b	Increment version to 2.0.13.dev2	2018-08-10 00:43:55 +02:00
Matthew Honnibal	c4ac981e6d	Try again to filter warnings	2018-08-10 00:42:54 +02:00
Matthew Honnibal	ae7fc42a41	Increment version to v2.0.13.dev1	2018-08-10 00:14:31 +02:00
Matthew Honnibal	7be9118be3	Require numpy>=1.15.0 to avoid the RuntimeWarning	2018-08-10 00:14:13 +02:00
Matthew Honnibal	19f5046934	Undoing warning suppression, as doesnt really work	2018-08-10 00:13:34 +02:00
Matthew Honnibal	3fb828352d	Set version to 2.0.13.dev0	2018-08-09 23:49:34 +02:00
Matthew Honnibal	1c0614ecd2	Catch numpy warning	2018-08-09 23:49:24 +02:00
Aashish Gangwani	6eebfc7bf4	Added numbers to ../lang/hi/lex_attrs.py (#2629 ) I have added numbers in hindi lex_attrs.py file according to Indian numbering system(https://en.wikipedia.org/wiki/Indian_numbering_system) and here are there english translations: 'शून्य' => zero 'एक' => one 'दो' => two 'तीन' => three 'चार' => four 'पांच' => five 'छह' => six 'सात'=>seven 'आठ' => eight 'नौ' => nine 'दस' => ten 'ग्यारह' => eleven 'बारह' => twelve 'तेरह' => thirteen 'चौदह' => fourteen 'पंद्रह' => fifteen 'सोलह'=> sixteen 'सत्रह' => seventeen 'अठारह' => eighteen 'उन्नीस' => nineteen 'बीस' => twenty 'तीस' => thirty 'चालीस' => forty 'पचास' => fifty 'साठ' => sixty 'सत्तर' => seventy 'अस्सी' => eighty 'नब्बे' => ninety 'सौ' => hundred 'हज़ार' => thousand 'लाख' => hundred thousand 'करोड़' => ten million 'अरब' => billion 'खरब' => hundred billion <!--- Provide a general summary of your changes in the title. --> ## Description <!--- Use this section to describe your changes. If your changes required testing, include information about the testing environment and the tests you ran. If your test fixes a bug reported in an issue, don't forget to include the issue number. If your PR is still a work in progress, that's totally fine – just include a note to let us know. --> ### Types of change <!-- What type of change does your PR cover? Is it a bug fix, an enhancement or new feature, or a change to the documentation? --> ## Checklist <!--- Before you submit the PR, go over this checklist and make sure you can tick off all the boxes. [] -> [x] --> - [x] I have submitted the spaCy Contributor Agreement. - [x] I ran the tests, and all new and existing tests passed. - [x] My changes don't require a change to the documentation, or if they do, I've added all required information.	2018-08-08 16:06:11 +02:00
Ines Montani	71723cece1	Add note on visualizing long texts ans sentences (see #2636 ) [ci skip]	2018-08-08 15:28:21 +02:00
Ines Montani	6147bd3eb4	Fix link target (closes #2645 ) [ci skip]	2018-08-08 15:03:52 +02:00
Ines Montani	8c47da1f19	Update Language serialization docs (see #2628 ) [ci skip] Add note on using from_disk and from_bytes via subclasses and add example	2018-08-07 14:17:57 +02:00
Emil Stenström	3834f4146d	Add abbreviations from UD_Swedish-Talbanken (#2613 ) * Add abbreviations from UD_Swedish-Talbanken * Add contributor agreement.	2018-08-07 13:53:17 +02:00
Ole Henrik Skogstrøm	0473add369	Feature/span ents (#2599 ) * Created Span.ents property * Add tests for span.ents * Add tests for start and end of sentence	2018-08-07 13:52:32 +02:00
Xiaoquan Kong	87fa847e6e	Fix Chinese language related bugs (#2634 )	2018-08-07 11:26:31 +02:00
Matthew Honnibal	664cfc29bc	Merge branch 'master' of https://github.com/explosion/spaCy	2018-08-07 10:49:39 +02:00
Matthew Honnibal	2278c9734e	Fix spelling error #2640	2018-08-07 10:49:21 +02:00
Xiaoquan Kong	f0c9652ed1	New Feature: display more detail when Error E067 (#2639 ) * Fix off-by-one error * Add verbose option * Update verbose option * Update documents for verbose option	2018-08-07 10:45:29 +02:00
Emil Stenström	1914c488d3	Swedish: Exceptions for single letter words ending sentence (#2615 ) * Exceptions for single letter words ending sentence Sentences ending in "i." (as in "... peka i."), "m." (as in "...än 2000 m."), should be tokenized as two separate tokens. * Add test	2018-08-05 14:14:30 +02:00
Matthew Honnibal	860f5bd91f	Add test for issue 2626	2018-08-05 13:46:57 +02:00
Matthew Honnibal	f762d52b24	Add example for Issue #2627	2018-08-05 13:33:52 +02:00
Ines Montani	6a4360e425	Update universe [ci skip]	2018-08-02 17:33:08 +02:00
Sami	dbc993f5b3	Updating description and code snippet spacy-lefff (#2623 ) * updating description and code snippet spacy-lefff * contributors agreement	2018-08-02 17:25:27 +02:00
Vikas Kumar Yadav	23876dbc70	Create vikaskyadav.md (#2621 )	2018-08-02 14:03:44 +02:00
Vikas Kumar Yadav	d3e21aad64	Update _benchmarks.jade (#2618 )	2018-08-02 00:28:28 +02:00
Brian Phillips	8227de0099	Update language.jade (#2616 )	2018-07-31 12:34:42 +02:00
Ioannis Daras	055cc0de44	Bug fix to pseudocode for tokenizer customization (#2604 )	2018-07-27 11:04:12 +02:00
Kaisa (Katarzyna) Korsak	e531a827db	Changed conllu2json to be able to extract NER tags (#2594 ) * extract ner tags from conllu file if available * fixed a bug in regex	2018-07-25 22:21:31 +02:00
Dmitry Bruhanov	07d0cc9de7	Update examples.py (#2597 )	2018-07-25 22:20:24 +02:00
Andriy Mulyar	e9ef51137d	Fixed typo (#2596 ) Changed 'The index of the first character after the span.' to The index of the last character after the span' in description of doc.char_span	2018-07-25 22:17:15 +02:00
Matthew Honnibal	66983d8412	Port BenDerPan's Chinese changes to v2 (finally) (#2591 ) * add template files for Chinese * add template files for Chinese, and test directory .	2018-07-25 02:47:23 +02:00
ines	f2e3e039b7	Update French stop words (resolves #2540 )	2018-07-24 23:41:51 +02:00
Ines Montani	75f3234404	💫 Refactor test suite (#2568 ) ## Description Related issues: #2379 (should be fixed by separating model tests) * total execution time down from > 300 seconds to under 60 seconds 🎉 * removed all model-specific tests that could only really be run manually anyway – those will now live in a separate test suite in the [`spacy-models`](https://github.com/explosion/spacy-models) repository and are already integrated into our new model training infrastructure * changed all relative imports to absolute imports to prepare for moving the test suite from `/spacy/tests` to `/tests` (it'll now always test against the installed version) * merged old regression tests into collections, e.g. `test_issue1001-1500.py` (about 90% of the regression tests are very short anyways) * tidied up and rewrote existing tests wherever possible ### Todo - [ ] move tests to `/tests` and adjust CI commands accordingly - [x] move model test suite from internal repo to `spacy-models` - [x] ~~investigate why `pipeline/test_textcat.py` is flakey~~ - [x] review old regression tests (leftover files) and see if they can be merged, simplified or deleted - [ ] update documentation on how to run tests ### Types of change enhancement, tests ## Checklist <!--- Before you submit the PR, go over this checklist and make sure you can tick off all the boxes. [] -> [x] --> - [x] I have submitted the spaCy Contributor Agreement. - [x] I ran the tests, and all new and existing tests passed. - [ ] My changes don't require a change to the documentation, or if they do, I've added all required information.	2018-07-24 23:38:44 +02:00
Matthew Honnibal	82277f63a3	💫 Small efficiency fixes to tokenizer (#2587 ) This patch improves tokenizer speed by about 10%, and reduces memory usage in the `Vocab` by removing a redundant index. The `vocab._by_orth` and `vocab._by_hash` indexed on different data in v1, but in v2 the orth and the hash are identical. The patch also fixes an uninitialized variable in the tokenizer, the `has_special` flag. This checks whether a chunk we're tokenizing triggers a special-case rule. If it does, then we avoid caching within the chunk. This check led to incorrectly rejecting some chunks from the cache. With the `en_core_web_md` model, we now tokenize the IMDB train data at 503,104k words per second. Prior to this patch, we had 465,764k words per second. Before switching to the regex library and supporting more languages, we had 1.3m words per second for the tokenizer. In order to recover the missing speed, we need to: * Fix the variable-length lookarounds in the suffix, infix and `token_match` rules * Improve the performance of the `token_match` regex * Switch back from the `regex` library to the `re` library. ## Checklist <!--- Before you submit the PR, go over this checklist and make sure you can tick off all the boxes. [] -> [x] --> - [x] I have submitted the spaCy Contributor Agreement. - [x] I ran the tests, and all new and existing tests passed. - [x] My changes don't require a change to the documentation, or if they do, I've added all required information.	2018-07-24 23:35:54 +02:00
kororo	b1ec827ee0	Fix typo (#2579 ) Update slogan, desc and code snippet to latest version	2018-07-24 22:47:33 +02:00
ines	cd687091fb	Remove nl examples from widget for now [ci skip] Restore for next spaCy version when path to example sentences is fixed	2018-07-24 22:41:20 +02:00
ines	2d8ffb8bcd	Fix formatting	2018-07-24 22:40:49 +02:00
ines	1b3da8d2ae	Update website for v2.0.12 [ci skip]	2018-07-24 21:04:22 +02:00
Matthew Honnibal	e05bebce8e	Try setting appveyor to Python2 64	2018-07-24 20:47:03 +02:00
Matthew Honnibal	6303ce3d0e	Try to fix memory error by moving fr_tokenizer to module scope	2018-07-24 20:09:06 +02:00
Matthew Honnibal	afe3fa4449	Merge branch 'master' of https://github.com/explosion/spaCy	2018-07-24 19:44:31 +02:00
Matthew Honnibal	b2e9e958b9	Add session scoping to tokenizers to try to fix oom on Appveyor	2018-07-24 19:44:18 +02:00
Ines Montani	a43ad114c2	Fix typo [ci skip]	2018-07-24 18:45:40 +02:00
Dmitry Bruhanov	27160b1516	added some widespread written jargon & dialectizms (#2584 ) This jargon is not offencive but emotionally colored as funny due to its deviation from the norm for various reasons: immitating a dialect, deliberately wrong spelling emphasizing its low colloquial nature, obsolete form, foreign borrowing with native flections, etc. Dmitry Briukhanov, Linguist & Pythonist	2018-07-24 18:44:29 +02:00
Dmitry Bruhanov	4ad7de6ca9	DimaBryuhanov.md (#2590 ) # spaCy contributor agreement This spaCy Contributor Agreement ("SCA") is based on the [Oracle Contributor Agreement](http://www.oracle.com/technetwork/oca-405177.pdf). The SCA applies to any contribution that you make to any product or project managed by us (the "project"), and sets out the intellectual property rights you grant to us in the contributed materials. The term "us" shall mean [ExplosionAI UG (haftungsbeschränkt)](https://explosion.ai/legal). The term "you" shall mean the person or entity identified below. If you agree to be bound by these terms, fill in the information requested below and include the filled-in version with your first pull request, under the folder [`.github/contributors/`](/.github/contributors/). The name of the file should be your GitHub username, with the extension `.md`. For example, the user example_user would create the file `.github/contributors/example_user.md`. Read this agreement carefully before signing. These terms and conditions constitute a binding legal agreement. ## Contributor Agreement 1. The term "contribution" or "contributed materials" means any source code, object code, patch, tool, sample, graphic, specification, manual, documentation, or any other material posted or submitted by you to the project. 2. With respect to any worldwide copyrights, or copyright applications and registrations, in your contribution: * you hereby assign to us joint ownership, and to the extent that such assignment is or becomes invalid, ineffective or unenforceable, you hereby grant to us a perpetual, irrevocable, non-exclusive, worldwide, no-charge, royalty-free, unrestricted license to exercise all rights under those copyrights. This includes, at our option, the right to sublicense these same rights to third parties through multiple levels of sublicensees or other licensing arrangements; * you agree that each of us can do all things in relation to your contribution as if each of us were the sole owners, and if one of us makes a derivative work of your contribution, the one who makes the derivative work (or has it made will be the sole owner of that derivative work; * you agree that you will not assert any moral rights in your contribution against us, our licensees or transferees; * you agree that we may register a copyright in your contribution and exercise all ownership rights associated with it; and * you agree that neither of us has any duty to consult with, obtain the consent of, pay or render an accounting to the other for any use or distribution of your contribution. 3. With respect to any patents you own, or that you can license without payment to any third party, you hereby grant to us a perpetual, irrevocable, non-exclusive, worldwide, no-charge, royalty-free license to: * make, have made, use, sell, offer to sell, import, and otherwise transfer your contribution in whole or in part, alone or in combination with or included in any product, work or materials arising out of the project to which your contribution was submitted, and * at our option, to sublicense these same rights to third parties through multiple levels of sublicensees or other licensing arrangements. 4. Except as set out above, you keep all right, title, and interest in your contribution. The rights that you grant to us under these terms are effective on the date you first submitted a contribution to us, even if your submission took place before the date you sign these terms. 5. You covenant, represent, warrant and agree that: * Each contribution that you submit is and shall be an original work of authorship and you can legally grant the rights set out in this SCA; * to the best of your knowledge, each contribution will not violate any third party's copyrights, trademarks, patents, or other intellectual property rights; and * each contribution shall be in compliance with U.S. export control laws and other applicable export and import laws. You agree to notify us if you become aware of any circumstance which would make any of the foregoing representations inaccurate in any respect. We may publicly disclose your participation in the project, including the fact that you have signed the SCA. 6. This SCA is governed by the laws of the State of California and applicable U.S. Federal law. Any choice of law rules will not apply. 7. Please place an “x” on one of the applicable statement below. Please do NOT mark both statements: * [X] I am signing on behalf of myself as an individual and no other person or entity, including my employer, has or will have rights with respect to my contributions. * [ ] I am signing on behalf of my employer or a legal entity and I have the actual authority to contractually bind that entity. ## Contributor Details \| Field \| Entry \| \|------------------------------- \| -------------------- \| \| Name \| Dmitry Briukhanov \| \| Company name (if applicable) \| - \| \| Title or role (if applicable) \| - \| \| Date \| 7/24/2018 \| \| GitHub username \| DimaBryuhanov \| \| Website (optional) \| \|	2018-07-24 18:43:27 +02:00
Matthew Honnibal	1a16162da9	Merge branch 'master' of https://github.com/explosion/spaCy	2018-07-21 15:57:18 +02:00

1 2 3 4 5 ...

9001 Commits