spaCy

mirror of https://github.com/explosion/spaCy.git synced 2024-12-26 18:06:29 +03:00

Author	SHA1	Message	Date
Matthew Honnibal	4dc57d9e15	Update train_new_entity_type example	2019-02-24 16:41:03 +01:00
Matthew Honnibal	7ac0f9626c	Update rehearsal example	2019-02-24 16:17:41 +01:00
Matthew Honnibal	981cb89194	Fix f-score calculation if zero	2019-02-23 12:45:41 +01:00
Matthew Honnibal	5063d999e5	Set architecture in textcat example	2019-02-23 11:57:59 +01:00
Ines Montani	9696cf16c1	Merge branch 'master' into develop	2019-02-20 21:31:27 +01:00
Michael Liberman	386cec1979	- Json fix in comment (#3294 )	2019-02-19 18:01:35 +01:00
Ines Montani	5d0b60999d	Merge branch 'master' into develop	2019-02-07 20:54:07 +01:00
Laura Baakman	04aa041c9e	Update Example input JSON file to adhere to specification. (#3243 ) * Example file does not adhere to json input spec. According to the [json input spec ](https://spacy.io/api/annotation#json-input) the `id ` needs to be an `int` not a string. Using a string as `id` results in a `TypeError` when calling `spacy.gold.read_json_file()`. * Add spaCy Contributor Agreement.	2019-02-07 16:18:01 +01:00
Hunter Kelly	f28a1c7271	Update call to `mkdir()` to create the parents (#3139 ) * Update call to `mkdir()` to create the parents - Update the call to `output_dir.mkdir()` to also create the parents if needed * don't automatically create parents but fail fast if cannot create directory * add signed contributors agreement for retnuh	2019-01-11 03:02:18 +01:00
Ines Montani	61d09c481b	Merge branch 'master' into develop	2018-12-18 13:48:10 +01:00
Ines Montani	e3405f8af3	Don't call begin_training if updating new model (see #3059 ) [ci skip]	2018-12-17 13:45:49 +01:00
Ines Montani	c9a89bba50	Don't call begin_training if updating new model (see #3059 ) [ci skip]	2018-12-17 13:45:28 +01:00
Ines Montani	6f1438b5d9	Auto-format example	2018-12-17 13:44:38 +01:00
Matthew Honnibal	83ac227bd3	💫 Better support for semi-supervised learning (#3035 ) The new spacy pretrain command implemented BERT/ULMFit/etc-like transfer learning, using our Language Modelling with Approximate Outputs version of BERT's cloze task. Pretraining is convenient, but in some ways it's a bit of a strange solution. All we're doing is initialising the weights. At the same time, we're putting a lot of work into our optimisation so that it's less sensitive to initial conditions, and more likely to find good optima. I discuss this a bit in the pseudo-rehearsal blog post: https://explosion.ai/blog/pseudo-rehearsal-catastrophic-forgetting Support semi-supervised learning in spacy train One obvious way to improve these pretraining methods is to do multi-task learning, instead of just transfer learning. This has been shown to work very well: https://arxiv.org/pdf/1809.08370.pdf . This patch makes it easy to do this sort of thing. Add a new argument to spacy train, --raw-text. This takes a jsonl file with unlabelled data that can be used in arbitrary ways to do semi-supervised learning. Add a new method to the Language class and to pipeline components, .rehearse(). This is like .update(), but doesn't expect GoldParse objects. It takes a batch of Doc objects, and performs an update on some semi-supervised objective. Move the BERT-LMAO objective out from spacy/cli/pretrain.py into spacy/_ml.py, so we can create a new pipeline component, ClozeMultitask. This can be specified as a parser or NER multitask in the spacy train command. Example usage: python -m spacy train en ./tmp ~/data/en-core-web/train/nw.json ~/data/en-core-web/dev/nw.json --pipeline parser --raw-textt ~/data/unlabelled/reddit-100k.jsonl --vectors en_vectors_web_lg --parser-multitasks cloze Implement rehearsal methods for pipeline components The new --raw-text argument and nlp.rehearse() method also gives us a good place to implement the the idea in the pseudo-rehearsal blog post in the parser. This works as follows: Add a new nlp.resume_training() method. This allocates copies of pre-trained models in the pipeline, setting things up for the rehearsal updates. It also returns an optimizer object. This also greatly reduces confusion around the nlp.begin_training() method, which randomises the weights, making it not suitable for adding new labels or otherwise fine-tuning a pre-trained model. Implement rehearsal updates on the Parser class, making it available for the dependency parser and NER. During rehearsal, the initial model is used to supervise the model being trained. The current model is asked to match the predictions of the initial model on some data. This minimises catastrophic forgetting, by keeping the model's predictions close to the original. See the blog post for details. Implement rehearsal updates for tagger Implement rehearsal updates for text categoriz	2018-12-10 16:25:33 +01:00
Matthew Honnibal	375f0dc529	💫 Make TextCategorizer default to a simpler, GPU-friendly model (#3038 ) Currently the TextCategorizer defaults to a fairly complicated model, designed partly around the active learning requirements of Prodigy. The model's a bit slow, and not very GPU-friendly. This patch implements a straightforward CNN model that still performs pretty well. The replacement model also makes it easy to use the LMAO pretraining, since most of the parameters are in the CNN. The replacement model has a flag to specify whether labels are mutually exclusive, which defaults to True. This has been a common problem with the text classifier. We'll also now be able to support adding labels to pretrained models again. Resolves #2934, #2756, #1798, #1748.	2018-12-10 14:37:39 +01:00
Matthew Honnibal	e5685d98a2	Fix averaging in textcat example (closes #2745 ) (#3032 ) [ci skip]	2018-12-08 13:27:05 +01:00
Ines Montani	5b2741f751	Remove unused cytoolz / itertools imports	2018-12-03 02:12:07 +01:00
Gavriel Loria	ae5601beae	Initialize trues to 0.0 in training example (#3004 ) * added contributor agreement * if there are no true positives, precision should be 0.0	2018-12-03 01:33:22 +01:00
Ines Montani	45798cc53e	Auto-format examples	2018-12-02 04:26:26 +01:00
Ines Montani	d33953037e	💫 Port master changes over to develop (#2979 ) * Create aryaprabhudesai.md (#2681) * Update _install.jade (#2688) Typo fix: "models" -> "model" * Add FAC to spacy.explain (resolves #2706) * Remove docstrings for deprecated arguments (see #2703) * When calling getoption() in conftest.py, pass a default option (#2709) * When calling getoption() in conftest.py, pass a default option This is necessary to allow testing an installed spacy by running: pytest --pyargs spacy * Add contributor agreement * update bengali token rules for hyphen and digits (#2731) * Less norm computations in token similarity (#2730) * Less norm computations in token similarity * Contributor agreement * Remove ')' for clarity (#2737) Sorry, don't mean to be nitpicky, I just noticed this when going through the CLI and thought it was a quick fix. That said, if this was intention than please let me know. * added contributor agreement for mbkupfer (#2738) * Basic support for Telugu language (#2751) * Lex _attrs for polish language (#2750) * Signed spaCy contributor agreement * Added polish version of english lex_attrs * Introduces a bulk merge function, in order to solve issue #653 (#2696) * Fix comment * Introduce bulk merge to increase performance on many span merges * Sign contributor agreement * Implement pull request suggestions * Describe converters more explicitly (see #2643) * Add multi-threading note to Language.pipe (resolves #2582) [ci skip] * Fix formatting * Fix dependency scheme docs (closes #2705) [ci skip] * Don't set stop word in example (closes #2657) [ci skip] * Add words to portuguese language _num_words (#2759) * Add words to portuguese language _num_words * Add words to portuguese language _num_words * Update Indonesian model (#2752) * adding e-KTP in tokenizer exceptions list * add exception token * removing lines with containing space as it won't matter since we use .split() method in the end, added new tokens in exception * add tokenizer exceptions list * combining base_norms with norm_exceptions * adding norm_exception * fix double key in lemmatizer * remove unused import on punctuation.py * reformat stop_words to reduce number of lines, improve readibility * updating tokenizer exception * implement is_currency for lang/id * adding orth_first_upper in tokenizer_exceptions * update the norm_exception list * remove bunch of abbreviations * adding contributors file * Fixed spaCy+Keras example (#2763) * bug fixes in keras example * created contributor agreement * Adding French hyphenated first name (#2786) * Fix typo (closes #2784) * Fix typo (#2795) [ci skip] Fixed typo on line 6 "regcognizer --> recognizer" * Adding basic support for Sinhala language. (#2788) * adding Sinhala language package, stop words, examples and lex_attrs. * Adding contributor agreement * Updating contributor agreement * Also include lowercase norm exceptions * Fix error (#2802) * Fix error ValueError: cannot resize an array that references or is referenced by another array in this way. Use the resize function * added spaCy Contributor Agreement * Add charlax's contributor agreement (#2805) * agreement of contributor, may I introduce a tiny pl languge contribution (#2799) * Contributors agreement * Contributors agreement * Contributors agreement * Add jupyter=True to displacy.render in documentation (#2806) * Revert "Also include lowercase norm exceptions" This reverts commit `70f4e8adf3`. * Remove deprecated encoding argument to msgpack * Set up dependency tree pattern matching skeleton (#2732) * Fix bug when too many entity types. Fixes #2800 * Fix Python 2 test failure * Require older msgpack-numpy * Restore encoding arg on msgpack-numpy * Try to fix version pin for msgpack-numpy * Update Portuguese Language (#2790) * Add words to portuguese language _num_words * Add words to portuguese language _num_words * Portuguese - Add/remove stopwords, fix tokenizer, add currency symbols * Extended punctuation and norm_exceptions in the Portuguese language * Correct error in spacy universe docs concerning spacy-lookup (#2814) * Update Keras Example for (Parikh et al, 2016) implementation (#2803) * bug fixes in keras example * created contributor agreement * baseline for Parikh model * initial version of parikh 2016 implemented * tested asymmetric models * fixed grevious error in normalization * use standard SNLI test file * begin to rework parikh example * initial version of running example * start to document the new version * start to document the new version * Update Decompositional Attention.ipynb * fixed calls to similarity * updated the README * import sys package duh * simplified indexing on mapping word to IDs * stupid python indent error * added code from https://github.com/tensorflow/tensorflow/issues/3388 for tf bug workaround * Fix typo (closes #2815) [ci skip] * Update regex version dependency * Set version to 2.0.13.dev3 * Skip seemingly problematic test * Remove problematic test * Try previous version of regex * Revert "Remove problematic test" This reverts commit `bdebbef455`. * Unskip test * Try older version of regex * 💫 Update training examples and use minibatching (#2830) <!--- Provide a general summary of your changes in the title. --> ## Description Update the training examples in `/examples/training` to show usage of spaCy's `minibatch` and `compounding` helpers ([see here](https://spacy.io/usage/training#tips-batch-size) for details). The lack of batching in the examples has caused some confusion in the past, especially for beginners who would copy-paste the examples, update them with large training sets and experienced slow and unsatisfying results. ### Types of change enhancements ## Checklist <!--- Before you submit the PR, go over this checklist and make sure you can tick off all the boxes. [] -> [x] --> - [x] I have submitted the spaCy Contributor Agreement. - [x] I ran the tests, and all new and existing tests passed. - [x] My changes don't require a change to the documentation, or if they do, I've added all required information. * Visual C++ link updated (#2842) (closes #2841) [ci skip] * New landing page * Add contribution agreement * Correcting lang/ru/examples.py (#2845) * Correct some grammatical inaccuracies in lang\ru\examples.py; filled Contributor Agreement * Correct some grammatical inaccuracies in lang\ru\examples.py * Move contributor agreement to separate file * Set version to 2.0.13.dev4 * Add Persian(Farsi) language support (#2797) * Also include lowercase norm exceptions * Remove in favour of https://github.com/explosion/spaCy/graphs/contributors * Rule-based French Lemmatizer (#2818) <!--- Provide a general summary of your changes in the title. --> ## Description <!--- Use this section to describe your changes. If your changes required testing, include information about the testing environment and the tests you ran. If your test fixes a bug reported in an issue, don't forget to include the issue number. If your PR is still a work in progress, that's totally fine – just include a note to let us know. --> Add a rule-based French Lemmatizer following the english one and the excellent PR for [greek language optimizations](https://github.com/explosion/spaCy/pull/2558) to adapt the Lemmatizer class. ### Types of change <!-- What type of change does your PR cover? Is it a bug fix, an enhancement or new feature, or a change to the documentation? --> - Lemma dictionary used can be found [here](http://infolingu.univ-mlv.fr/DonneesLinguistiques/Dictionnaires/telechargement.html), I used the XML version. - Add several files containing exhaustive list of words for each part of speech - Add some lemma rules - Add POS that are not checked in the standard Lemmatizer, i.e PRON, DET, ADV and AUX - Modify the Lemmatizer class to check in lookup table as a last resort if POS not mentionned - Modify the lemmatize function to check in lookup table as a last resort - Init files are updated so the model can support all the functionalities mentioned above - Add words to tokenizer_exceptions_list.py in respect to regex used in tokenizer_exceptions.py ## Checklist <!--- Before you submit the PR, go over this checklist and make sure you can tick off all the boxes. [] -> [x] --> - [X] I have submitted the spaCy Contributor Agreement. - [X] I ran the tests, and all new and existing tests passed. - [X] My changes don't require a change to the documentation, or if they do, I've added all required information. * Set version to 2.0.13 * Fix formatting and consistency * Update docs for new version [ci skip] * Increment version [ci skip] * Add info on wheels [ci skip] * Adding "This is a sentence" example to Sinhala (#2846) * Add wheels badge * Update badge [ci skip] * Update README.rst [ci skip] * Update murmurhash pin * Increment version to 2.0.14.dev0 * Update GPU docs for v2.0.14 * Add wheel to setup_requires * Import prefer_gpu and require_gpu functions from Thinc * Add tests for prefer_gpu() and require_gpu() * Update requirements and setup.py * Workaround bug in thinc require_gpu * Set version to v2.0.14 * Update push-tag script * Unhack prefer_gpu * Require thinc 6.10.6 * Update prefer_gpu and require_gpu docs [ci skip] * Fix specifiers for GPU * Set version to 2.0.14.dev1 * Set version to 2.0.14 * Update Thinc version pin * Increment version * Fix msgpack-numpy version pin * Increment version * Update version to 2.0.16 * Update version [ci skip] * Redundant ')' in the Stop words' example (#2856) <!--- Provide a general summary of your changes in the title. --> ## Description <!--- Use this section to describe your changes. If your changes required testing, include information about the testing environment and the tests you ran. If your test fixes a bug reported in an issue, don't forget to include the issue number. If your PR is still a work in progress, that's totally fine – just include a note to let us know. --> ### Types of change <!-- What type of change does your PR cover? Is it a bug fix, an enhancement or new feature, or a change to the documentation? --> ## Checklist <!--- Before you submit the PR, go over this checklist and make sure you can tick off all the boxes. [] -> [x] --> - [ ] I have submitted the spaCy Contributor Agreement. - [ ] I ran the tests, and all new and existing tests passed. - [ ] My changes don't require a change to the documentation, or if they do, I've added all required information. * Documentation improvement regarding joblib and SO (#2867) Some documentation improvements ## Description 1. Fixed the dead URL to joblib 2. Fixed Stack Overflow brand name (with space) ### Types of change Documentation ## Checklist <!--- Before you submit the PR, go over this checklist and make sure you can tick off all the boxes. [] -> [x] --> - [x] I have submitted the spaCy Contributor Agreement. - [x] I ran the tests, and all new and existing tests passed. - [x] My changes don't require a change to the documentation, or if they do, I've added all required information. * raise error when setting overlapping entities as doc.ents (#2880) * Fix out-of-bounds access in NER training The helper method state.B(1) gets the index of the first token of the buffer, or -1 if no such token exists. Normally this is safe because we pass this to functions like state.safe_get(), which returns an empty token. Here we used it directly as an array index, which is not okay! This error may have been the cause of out-of-bounds access errors during training. Similar errors may still be around, so much be hunted down. Hunting this one down took a long time...I printed out values across training runs and diffed, looking for points of divergence between runs, when no randomness should be allowed. * Change PyThaiNLP Url (#2876) * Fix missing comma * Add example showing a fix-up rule for space entities * Set version to 2.0.17.dev0 * Update regex version * Revert "Update regex version" This reverts commit `62358dd867`. * Try setting older regex version, to align with conda * Set version to 2.0.17 * Add spacy-js to universe [ci-skip] * Add spacy-raspberry to universe (closes #2889) * Add script to validate universe json [ci skip] * Removed space in docs + added contributor indo (#2909) * - removed unneeded space in documentation * - added contributor info * Allow input text of length up to max_length, inclusive (#2922) * Include universe spec for spacy-wordnet component (#2919) * feat: include universe spec for spacy-wordnet component * chore: include spaCy contributor agreement * Minor formatting changes [ci skip] * Fix image [ci skip] Twitter URL doesn't work on live site * Check if the word is in one of the regular lists specific to each POS (#2886) * 💫 Create random IDs for SVGs to prevent ID clashes (#2927) Resolves #2924. ## Description Fixes problem where multiple visualizations in Jupyter notebooks would have clashing arc IDs, resulting in weirdly positioned arc labels. Generating a random ID prefix so even identical parses won't receive the same IDs for consistency (even if effect of ID clash isn't noticable here.) ### Types of change bug fix ## Checklist <!--- Before you submit the PR, go over this checklist and make sure you can tick off all the boxes. [] -> [x] --> - [x] I have submitted the spaCy Contributor Agreement. - [x] I ran the tests, and all new and existing tests passed. - [x] My changes don't require a change to the documentation, or if they do, I've added all required information. * Fix typo [ci skip] * fixes symbolic link on py3 and windows (#2949) * fixes symbolic link on py3 and windows during setup of spacy using command python -m spacy link en_core_web_sm en closes #2948 * Update spacy/compat.py Co-Authored-By: cicorias <cicorias@users.noreply.github.com> * Fix formatting * Update universe [ci skip] * Catalan Language Support (#2940) * Catalan language Support * Ddding Catalan to documentation * Sort languages alphabetically [ci skip] * Update tests for pytest 4.x (#2965) <!--- Provide a general summary of your changes in the title. --> ## Description - [x] Replace marks in params for pytest 4.0 compat ([see here](https://docs.pytest.org/en/latest/deprecations.html#marks-in-pytest-mark-parametrize)) - [x] Un-xfail passing tests (some fixes in a recent update resolved a bunch of issues, but tests were apparently never updated here) ### Types of change <!-- What type of change does your PR cover? Is it a bug fix, an enhancement or new feature, or a change to the documentation? --> ## Checklist <!--- Before you submit the PR, go over this checklist and make sure you can tick off all the boxes. [] -> [x] --> - [x] I have submitted the spaCy Contributor Agreement. - [x] I ran the tests, and all new and existing tests passed. - [x] My changes don't require a change to the documentation, or if they do, I've added all required information. * Fix regex pin to harmonize with conda (#2964) * Update README.rst * Fix bug where Vocab.prune_vector did not use 'batch_size' (#2977) Fixes #2976 * Fix typo * Fix typo * Remove duplicate file * Require thinc 7.0.0.dev2 Fixes bug in gpu_ops that would use cupy instead of numpy on CPU * Add missing import * Fix error IDs * Fix tests	2018-11-29 16:30:29 +01:00
Matthew Honnibal	7ed9124a45	Fix Python2 error on example	2018-11-14 19:35:17 +01:00
Matthew Honnibal	09aa616182	Make pretraining script work without GPU	2018-11-04 17:09:52 +01:00
Matthew Honnibal	bc8cda818c	Improve pretrain textcat example	2018-11-04 00:17:09 +00:00
Matthew Honnibal	3e7a96f99d	Improve pretrain textcat example	2018-11-03 17:44:12 +00:00
Matthew Honnibal	c87c50af62	Rename new example	2018-11-03 13:09:46 +00:00
Matthew Honnibal	8e8ccc0f92	Work on pretraining script	2018-11-03 12:53:25 +00:00
Matthew Honnibal	0127f10ba3	Improve train tensorizer script	2018-11-03 10:54:20 +00:00
Matthew Honnibal	baf7feae68	Add tensorizer training example	2018-11-02 23:30:06 +00:00
Ines Montani	4cd9ec0f00	💫 Update training examples and use minibatching (#2830 ) <!--- Provide a general summary of your changes in the title. --> ## Description Update the training examples in `/examples/training` to show usage of spaCy's `minibatch` and `compounding` helpers ([see here](https://spacy.io/usage/training#tips-batch-size) for details). The lack of batching in the examples has caused some confusion in the past, especially for beginners who would copy-paste the examples, update them with large training sets and experienced slow and unsatisfying results. ### Types of change enhancements ## Checklist <!--- Before you submit the PR, go over this checklist and make sure you can tick off all the boxes. [] -> [x] --> - [x] I have submitted the spaCy Contributor Agreement. - [x] I ran the tests, and all new and existing tests passed. - [x] My changes don't require a change to the documentation, or if they do, I've added all required information.	2018-10-10 01:40:29 +02:00
ines	4339f64128	Merge branch 'master' into develop	2018-07-19 16:15:03 +02:00
ines	d489ffb78b	Fix formatting [ci skip]	2018-07-19 13:22:25 +02:00
Matthew Honnibal	2c4a6d66fa	Merge master into develop. Big merge, many conflicts -- need to review	2018-04-29 14:49:26 +02:00
Ines Montani	49cee4af92	💫 Interactive code examples, spaCy Universe and various docs improvements (#2274 ) * Integrate Python kernel via Binder * Add live model test for languages with examples * Update docs and code examples * Adjust margin (if not bootstrapped) * Add binder version to global config * Update terminal and executable code mixins * Pass attributes through infobox and section * Hide v-cloak * Fix example * Take out model comparison for now * Add meta text for compat * Remove chart.js dependency * Tidy up and simplify JS and port big components over to Vue * Remove chartjs example * Add Twitter icon * Add purple stylesheet option * Add utility for hand cursor (special cases only) * Add transition classes * Add small option for section * Add thumb object for small round thumbnail images * Allow unset code block language via "none" value (workaround to still allow unset language to default to DEFAULT_SYNTAX) * Pass through attributes * Add syntax highlighting definitions for Julia, R and Docker * Add website icon * Remove user survey from navigation * Don't hide GitHub icon on small screens * Make top navigation scrollable on small screens * Remove old resources page and references to it * Add Universe * Add helper functions for better page URL and title * Update site description * Increment versions * Update preview images * Update mentions of resources * Fix image * Fix social images * Fix problem with cover sizing and floats * Add divider and move badges into heading * Add docstrings * Reference converting section * Add section on converting word vectors * Move converting section to custom section and fix formatting * Remove old fastText example * Move extensions content to own section Keep weird ID to not break permalinks for now (we don't want to rewrite URLs if not absolutely necessary) * Use better component example and add factories section * Add note on larger model * Use better example for non-vector * Remove similarity in context section Only works via small models with tensors so has always been kind of confusing * Add note on init-model command * Fix lightning tour examples and make excutable if possible * Add spacy train CLI section to train * Fix formatting and add video * Fix formatting * Fix textcat example description (resolves #2246) * Add dummy file to try resolve conflict * Delete dummy file * Tidy up [ci skip] * Ensure sufficient height of loading container * Add loading animation to universe * Update Thebelab build and use better startup message * Fix asset versioning * Fix typo [ci skip] * Add note on project idea label	2018-04-29 02:06:46 +02:00
Matthew Honnibal	68ad366935	Improve train_new_entity_type example	2018-03-29 20:26:41 +02:00
Matthew Honnibal	1f7229f40f	Revert "Merge branch 'develop' of https://github.com/explosion/spaCy into develop" This reverts commit `c9ba3d3c2d`, reversing changes made to `92c26a35d4`.	2018-03-27 19:23:02 +02:00
Matthew Honnibal	00557c5fdd	Add example of NER multitask objective	2018-01-21 19:46:37 +01:00
mpuels	1e8147aec7	fix: Add missing period in train data	2017-12-13 10:51:05 +01:00
mpuels	ee4d6fdd40	Fix typo in comment	2017-12-09 13:14:57 +01:00
ines	726fb2d0b5	Use fewer iterations by default to avoid overfitting on blank model (resolves #1632 )	2017-11-23 15:27:12 +01:00
ines	ec08996000	Add note on tags matching tokenization (see #1613 )	2017-11-20 15:12:47 +01:00
ines	f36fab39b0	Don't rename component in intent parser example (resolves #1551 ) Otherwise, the default saved model won't know that it's supposed to create spaCy's 'parser'.	2017-11-10 23:35:38 +01:00
Ines Montani	1a23a0f87e	Remove broken link (resolves #1541 )	2017-11-10 12:28:39 +01:00
ines	89bd40b821	Fix print statement in textcat training example (resolves #1515 )	2017-11-08 17:17:40 +01:00
ines	a09c096d3c	Get docs ready for v2.0.0	2017-11-07 12:00:43 +01:00
ines	173b1551af	Update examples	2017-11-07 01:22:30 +01:00
ines	1b1c9105b4	Update example compatibility statements	2017-11-07 01:11:45 +01:00
ines	8fb48b9b91	Update and document new util functions	2017-11-07 00:22:43 +01:00
Matthew Honnibal	d7016d4050	Update intent parser example	2017-11-06 23:31:11 +01:00
ines	fe498b3d5e	Update training examples to use "simple style"	2017-11-06 23:14:04 +01:00
ines	2dca9e71a1	Add notes on catastrophic forgetting (see #1496 )	2017-11-06 13:17:02 +01:00
Matthew Honnibal	e033162a1d	Update tagger training example	2017-11-01 21:49:08 +01:00
ines	8f1d3fc3ee	Update textcat example	2017-11-01 17:09:22 +01:00
Matthew Honnibal	dad8f09fba	Fix print statements in text classifier example	2017-11-01 16:34:31 +01:00
ines	bfe17b7df1	Fix begin_training if get_gold_tuples is None	2017-11-01 13:14:31 +01:00
ines	4b196fdf7f	Fix formatting	2017-11-01 00:43:22 +01:00
ines	33af6ac69a	Use even smaller examle size 100 was still too much, so try 20 instead	2017-10-30 19:46:45 +01:00
ines	f02b0af821	Fix path and use smaller example size 500 was too larger and caused laggy rendering	2017-10-30 19:44:35 +01:00
ines	18dde7869a	Update training data docs and add vocab JSONL	2017-10-30 19:40:05 +01:00
ines	b5643d8575	Update intent parser docs and add to usage docs	2017-10-27 04:49:05 +02:00
ines	9dfca0f2f8	Add example for custom intent parser	2017-10-27 03:55:11 +02:00
ines	4d272e25ee	Fix examples	2017-10-27 03:55:04 +02:00
ines	a7b9074b4c	Update textcat training example and docs	2017-10-27 00:48:45 +02:00
ines	b61866a2e4	Update textcat example	2017-10-27 00:32:19 +02:00
ines	f81cc0bd1c	Fix usage of disable_pipes	2017-10-27 00:31:30 +02:00
ines	f57043e6fe	Update docstring	2017-10-26 16:29:08 +02:00
ines	b90e958975	Update tagger and parser examples and add to docs	2017-10-26 16:27:42 +02:00
ines	f1529463a8	Update tagger training example	2017-10-26 16:19:02 +02:00
ines	e44bbb5361	Remove old example	2017-10-26 16:12:41 +02:00
ines	421c3837e8	Fix formatting	2017-10-26 16:11:25 +02:00
ines	4d896171ae	Use plac annotations for arguments	2017-10-26 16:11:20 +02:00
ines	c3b681e5fb	Use plac annotations for arguments and add n_iter	2017-10-26 16:11:05 +02:00
ines	bc2c92f22d	Use plac annotations for arguments	2017-10-26 16:10:56 +02:00
ines	b5c74dbb34	Update parser training example	2017-10-26 15:15:37 +02:00
ines	586b9047fd	Use create_pipe instead of importing the entity recognizer	2017-10-26 15:15:26 +02:00
ines	d425ede7e9	Fix example	2017-10-26 15:15:08 +02:00
ines	9d58673aaf	Update train_ner example for spaCy v2.0	2017-10-26 14:24:12 +02:00
ines	e904075f35	Remove stray print statements	2017-10-26 14:24:00 +02:00
ines	c30258c3a2	Remove old example	2017-10-26 14:23:52 +02:00
ines	615c315d70	Update train_new_entity_type example to use disable_pipes	2017-10-25 14:56:53 +02:00
ines	2b8e7c45e0	Use better training data JSON example	2017-10-24 16:00:56 +02:00
ines	9bf5751064	Pretty-print JSON	2017-10-24 12:22:17 +02:00
ines	6675755005	Add training data JSON example	2017-10-24 12:05:10 +02:00
Jeroen Bobbeldijk	84c6c20d1c	Fix #1444 : fix pipeline logic and wrong paramater in update call	2017-10-22 15:18:36 +02:00
Jeffrey Gerard	5ba970b495	minor cleanup	2017-10-12 12:34:46 -07:00
Jeffrey Gerard	39d3cbfdba	Bugfix example script train_ner_standalone.py, fails after training	2017-10-12 11:39:12 -07:00
Matthew Honnibal	563f46f026	Fix multi-label support for text classification The TextCategorizer class is supposed to support multi-label text classification, and allow training data to contain missing values. For this to work, the gradient of the loss should be 0 when labels are missing. Instead, there was no way to actually denote "missing" in the GoldParse class, and so the TextCategorizer class treated the label set within gold.cats as complete. To fix this, we change GoldParse.cats to be a dict instead of a list. The GoldParse.cats dict should map to floats, with 1. denoting 'present' and 0. denoting 'absent'. Gradients are zeroed for categories absent from the gold.cats dict. A nice bonus is that you can also set values between 0 and 1 for partial membership. You can also set numeric values, if you're using a text classification model that uses an appropriate loss function. Unfortunately this is a breaking change; although the functionality was only recently introduced and hasn't been properly documented yet. I've updated the example script accordingly.	2017-10-05 18:43:02 -05:00
Matthew Honnibal	f1b86dff8c	Update textcat example	2017-10-04 15:12:28 +02:00
Matthew Honnibal	79a94bc166	Update textcat exampe	2017-10-04 14:55:30 +02:00
Matthew Honnibal	cbb1fbef80	Update train_ner_standalone example	2017-10-03 18:49:38 +02:00
Matthew Honnibal	027a5d8b75	Update train_ner_standalone example	2017-09-15 10:36:46 +02:00
Matthew Honnibal	683d81bb49	Update example for adding entity type	2017-09-14 16:15:59 +02:00
Matthew Honnibal	c16ef0a85c	Clarify train textcat example	2017-07-29 21:59:27 +02:00
Matthew Honnibal	54a539a113	Finish text classifier example	2017-07-23 00:34:12 +02:00
Matthew Honnibal	2bc7d87c70	Add example for training text classifier	2017-07-22 20:15:32 +02:00
ines	992559bf9a	Fix formatting and remove unused imports	2017-06-01 12:47:18 +02:00
Matthew Honnibal	5c30466c95	Update NER training example	2017-05-31 13:42:12 +02:00
Matthew Honnibal	2da16adcc2	Add dropout optin for parser and NER Dropout can now be specified in the `Parser.update()` method via the `drop` keyword argument, e.g. nlp.entity.update(doc, gold, drop=0.4) This will randomly drop 40% of features, and multiply the value of the others by 1. / 0.4. This may be useful for generalising from small data sets. This commit also patches the examples/training/train_new_entity_type.py example, to use dropout and fix the output (previously it did not output the learned entity).	2017-04-27 13:18:39 +02:00
Matthew Honnibal	0605b95f2e	Merge branch 'master' of https://github.com/explosion/spaCy	2017-04-18 13:48:00 +02:00
Matthew Honnibal	2f84626417	Fix train_new_entity_type example	2017-04-18 13:47:36 +02:00
Ines Montani	734b0a4e4a	Update train_new_entity_type.py	2017-04-16 23:42:16 +02:00

1 2 3 4

169 Commits