spaCy

mirror of https://github.com/explosion/spaCy.git synced 2024-12-26 18:06:29 +03:00

Author	SHA1	Message	Date
Matthew Honnibal	92c26a35d4	Update get_cuda_stream	2018-03-27 16:42:00 +00:00
Matthew Honnibal	bede11b67c	Improve label management in parser and NER (#2108 ) This patch does a few smallish things that tighten up the training workflow a little, and allow memory use during training to be reduced by letting the GoldCorpus stream data properly. Previously, the parser and entity recognizer read and saved labels as lists, with extra labels noted separately. Lists were used becaue ordering is very important, to ensure that the label-to-class mapping is stable. We now manage labels as nested dictionaries, first keyed by the action, and then keyed by the label. Values are frequencies. The trick is, how do we save new labels? We need to make sure we iterate over these in the same order they're added. Otherwise, we'll get different class IDs, and the model's predictions won't make sense. To allow stable sorting, we map the new labels to negative values. If we have two new labels, they'll be noted as having "frequency" -1 and -2. The next new label will then have "frequency" -3. When we sort by (frequency, label), we then get a stable sort. Storing frequencies then allows us to make the next nice improvement. Previously we had to iterate over the whole training set, to pre-process it for the deprojectivisation. This led to storing the whole training set in memory. This was most of the required memory during training. To prevent this, we now store the frequencies as we stream in the data, and deprojectivize as we go. Once we've built the frequencies, we can then apply a frequency cut-off when we decide how many classes to make. Finally, to allow proper data streaming, we also have to have some way of shuffling the iterator. This is awkward if the training files have multiple documents in them. To solve this, the GoldCorpus class now writes the training data to disk in msgpack files, one per document. We can then shuffle the data by shuffling the paths. This is a squash merge, as I made a lot of very small commits. Individual commit messages below. * Simplify label management for TransitionSystem and its subclasses * Fix serialization for new label handling format in parser * Simplify and improve GoldCorpus class. Reduce memory use, write to temp dir * Set actions in transition system * Require thinc 6.11.1.dev4 * Fix error in parser init * Add unicode declaration * Fix unicode declaration * Update textcat test * Try to get model training on less memory * Print json loc for now * Try rapidjson to reduce memory use * Remove rapidjson requirement * Try rapidjson for reduced mem usage * Handle None heads when projectivising * Stream json docs * Fix train script * Handle projectivity in GoldParse * Fix projectivity handling * Add minibatch_by_words util from ud_train * Minibatch by number of words in spacy.cli.train * Move minibatch_by_words util to spacy.util * Fix label handling * More hacking at label management in parser * Fix encoding in msgpack serialization in GoldParse * Adjust batch sizes in parser training * Fix minibatch_by_words * Add merge_subtokens function to pipeline.pyx * Register merge_subtokens factory * Restore use of msgpack tmp directory * Use minibatch-by-words in train * Handle retokenization in scorer * Change back-off approach for missing labels. Use 'dep' label * Update NER for new label management * Set NER tags for over-segmented words * Fix label alignment in gold * Fix label back-off for infrequent labels * Fix int type in labels dict key * Fix int type in labels dict key * Update feature definition for 8 feature set * Update ud-train script for new label stuff * Fix json streamer * Print the line number if conll eval fails * Update children and sentence boundaries after deprojectivisation * Export set_children_from_heads from doc.pxd * Render parses during UD training * Remove print statement * Require thinc 6.11.1.dev6. Try adding wheel as install_requires * Set different dev version, to flush pip cache * Update thinc version * Update GoldCorpus docs * Remove print statements * Fix formatting and links [ci skip]	2018-03-19 02:58:08 +01:00
Matthew Honnibal	31b156d60b	Fix itershuffle	2018-03-10 22:32:59 +01:00
Johannes Dollinger	012e874d09	Add contributor agreement for emulbreh	2018-02-13 13:40:33 +01:00
Johannes Dollinger	bf94c13382	Don't fix random seeds on import	2018-02-13 12:42:23 +01:00
ines	35653bef3a	Add missing import (fixes #1546 )	2017-11-10 19:05:18 +01:00
Matthew Honnibal	726f689da4	Fix missing import	2017-11-07 13:20:12 +01:00
ines	8fb48b9b91	Update and document new util functions	2017-11-07 00:22:43 +01:00
Matthew Honnibal	1cab703bba	Move minibatch function to util	2017-11-06 23:45:36 +01:00
ines	39e0586192	Add deprecated helper Uses warning to show DeprecationWarning and custom stack trace	2017-11-01 16:32:36 +01:00
Matthew Honnibal	a7bf38bf31	Remove misleading comment on util.get_cuda_stream()	2017-11-01 13:57:25 +01:00
ines	ea4a41c8fb	Tidy up util and helpers	2017-10-27 14:39:09 +02:00
Matthew Honnibal	9baa8fe7ec	Convert closure to functools.partial, to promote pickling	2017-10-17 18:20:52 +02:00
Matthew Honnibal	df488274b1	Fix deserialization of vectors	2017-10-16 20:55:00 +02:00
ines	d5418553eb	Fix whitespace	2017-10-16 18:30:04 +02:00
ines	6ceadcdb5c	Make sure from_disk passes string to numpy (see #1421 ) If path is a WindowsPath, numpy does not recognise it as a path and as a result, doesn't open the file. https://github.com/numpy/numpy/blob/master/numpy/lib/npyio.py#L369	2017-10-16 18:29:56 +02:00
ines	b39409173e	Add disable option and True/False/None values for pipeline	2017-10-07 00:29:08 +02:00
ines	212c8f0711	Implement new Language methods and pipeline API	2017-10-07 00:25:54 +02:00
Matthew Honnibal	f24c2e3a8a	Fix evaluate for non-GPU	2017-10-03 22:47:31 +02:00
ines	8dbe49ecb8	Always compare lowercase package names Otherwise, is_package will return False if model name contains uppercase characters. See this issue: https://support.prodi.gy/t/saving-a-trained-ner-model-as-a-loadable-modu le/46/6	2017-09-29 20:55:17 +02:00
ines	153c2589d4	Revert "Always compare lowercase package names" This reverts commit `7d77dc490f`.	2017-09-29 20:53:36 +02:00
ines	7d77dc490f	Always compare lowercase package names Otherwise, is_package will return False if model name contains uppercase characters. See this issue: https://support.prodi.gy/t/saving-a-trained-ner-model-as-a-loadable-modu le/46/6	2017-09-29 20:52:28 +02:00
Matthew Honnibal	ffda38356a	Add util function to enable GPU	2017-09-20 19:16:35 -05:00
ines	68f66aebf8	Use pkg_resources instead of pip for is_package (resolves #1293 )	2017-09-16 20:27:59 +02:00
Matthew Honnibal	30e35d9666	Fix syntax error	2017-08-30 17:35:39 -05:00
ines	173089a45a	Add more validation for model meta	2017-08-29 11:21:46 +02:00
Matthew Honnibal	ed95009b5c	Fix data loading on Python 2	2017-08-18 21:57:06 +02:00
Dan O'Huiginn	ebf5a3ce59	Allow loading with python < 3.6 Don't rely on recent python features to load models Fixes Issue #1271	2017-08-17 15:15:47 +00:00
ines	ea167e14db	Fix model package loading from link	2017-06-05 13:10:49 +02:00
ines	dd6dc4c120	Update spacy.load() helper functions	2017-06-05 13:02:31 +02:00
ines	7db1a0e83e	Make sure printed values are always strings	2017-06-04 21:27:20 +02:00
ines	070e026ed9	Ensure path on read_json	2017-06-04 20:44:37 +02:00
ines	e1e73936b1	Raise correct error	2017-06-04 20:44:27 +02:00
ines	4c2bbc3ccc	Add add_lookups util function	2017-06-03 19:44:47 +02:00
ines	924c58bde3	Fix serialization of optional elements	2017-06-02 18:18:17 +02:00
Matthew Honnibal	1d18cedae8	Fiddle with msgpack bytes vs unicode	2017-06-01 10:48:43 -05:00
Matthew Honnibal	3ff7d7fcef	Merge for updated requirements	2017-06-01 04:57:47 -05:00
Matthew Honnibal	ae8010b526	Move weight serialization to Thinc	2017-06-01 02:56:12 -05:00
Matthew Honnibal	c8a58cfcf8	Fix Python2/3 load bug	2017-05-31 15:21:44 -05:00
Matthew Honnibal	8dfb9546f0	Merge branch 'develop' of https://github.com/explosion/spaCy into develop	2017-05-31 07:21:14 -05:00
Matthew Honnibal	92f9e5cc9a	Silence env_opt, and fix serialization for GPU	2017-05-31 07:14:11 -05:00
Matthew Honnibal	33e5ec737f	Fix to/from disk methods	2017-05-31 13:43:10 +02:00
Matthew Honnibal	2a061e2777	Fix serialisation, for reals this time	2017-05-29 17:52:08 -05:00
Matthew Honnibal	35d981241f	Fix model deserialization	2017-05-29 14:46:31 -05:00
Matthew Honnibal	5b29f227ae	Fix serialization	2017-05-29 14:35:53 -05:00
Matthew Honnibal	1e6df0a2a1	Merge branch 'develop' of https://github.com/explosion/spaCy into develop	2017-05-29 14:30:12 -05:00
ines	08382f21e3	Pass model meta to nlp object in load_model	2017-05-29 20:44:11 +02:00
Matthew Honnibal	f1acdaab55	Fix serialization of weight offsets	2017-05-29 13:23:11 -05:00
Matthew Honnibal	c044e9c21c	Merge branch 'develop' of https://github.com/explosion/spaCy into develop	2017-05-29 08:41:02 -05:00
Matthew Honnibal	aa4c33914b	Work on serialization	2017-05-29 08:40:45 -05:00
ines	567485a818	Fix and document model loading with pipeline and overrides	2017-05-29 14:10:10 +02:00
Matthew Honnibal	deac7eb01c	Fix for serialization	2017-05-29 13:54:18 +02:00
Matthew Honnibal	04c32aa091	Fix for serialization	2017-05-29 13:53:32 +02:00
Matthew Honnibal	a1960c2d09	Fix for serialization	2017-05-29 13:47:42 +02:00
Matthew Honnibal	7b06bb896e	Fix for serialization	2017-05-29 13:42:55 +02:00
Matthew Honnibal	f4aafca222	Merge changes to test_misc	2017-05-29 12:26:02 +02:00
Matthew Honnibal	ff26aa6c37	Work on to/from bytes/disk serialization methods	2017-05-29 11:45:45 +02:00
ines	df920ba0e7	Add tests for displaCy and util functions and fix util typo	2017-05-29 10:51:19 +02:00
Matthew Honnibal	c91b121aeb	Move serialization functions to util	2017-05-29 10:13:42 +02:00
Matthew Honnibal	6dad4117ad	Work on serialization for models	2017-05-29 01:37:57 +02:00
ines	c1983621fb	Update util functions for model loading	2017-05-28 00:22:40 +02:00
ines	c8543c8237	Fix formatting and docstrings and remove deprecated function	2017-05-28 00:22:40 +02:00
ines	51882c4984	Fix formatting	2017-05-26 12:37:45 +02:00
Matthew Honnibal	80cf42e33b	Fix compounding and decaying utils	2017-05-25 17:15:39 -05:00
Matthew Honnibal	b9cea9cd93	Add compounding and decaying functions	2017-05-25 16:16:10 -05:00
ines	b5fb43fdd8	Allow sys.exit status as exits keyword arg in util.prints()	2017-05-22 12:29:15 +02:00
Matthew Honnibal	5db89053aa	Merge docstrings	2017-05-21 13:46:23 -05:00
Matthew Honnibal	0731971bfc	Add itershuffle utility function. Maybe belongs in thinc	2017-05-21 09:05:05 -05:00
ines	3871157d84	Update spacy.util documentation	2017-05-21 01:12:09 +02:00
Matthew Honnibal	238be0f16a	Merge branch 'develop' of https://github.com/explosion/spaCy into develop	2017-05-18 08:32:22 -05:00
Matthew Honnibal	c214c0decb	Improve env_opt reporting	2017-05-18 08:32:03 -05:00
ines	489d2fb4ba	Add is_in_jupyter() helper for displaCy (see #1058 )	2017-05-18 14:13:14 +02:00
ines	abf0188b0a	Move cupy and CudaStream to compat	2017-05-18 14:12:45 +02:00
Matthew Honnibal	fc8d3a112c	Add util.env_opt support: Can set hyper params through environment variables.	2017-05-18 04:36:53 -05:00
Matthew Honnibal	1d7c18e58a	Merge branch 'develop' of https://github.com/explosion/spaCy into develop	2017-05-15 21:53:47 +02:00
Matthew Honnibal	a9edb3aa1d	Improve integration of NN parser, to support unified training API	2017-05-15 21:53:27 +02:00
ines	c31792aaec	Add displaCy visualisers (see #1058 )	2017-05-14 17:50:23 +02:00
ines	b462076d80	Merge load_lang_class and get_lang_class	2017-05-14 01:31:10 +02:00
ines	36bebe7164	Update docstrings	2017-05-14 01:30:29 +02:00
Matthew Honnibal	4b9d69f428	Merge branch 'v2' into develop * Move v2 parser into nn_parser.pyx * New TokenVectorEncoder class in pipeline.pyx * New spacy/_ml.py module Currently the two parsers live side-by-side, until we figure out how to organize them.	2017-05-14 01:10:23 +02:00
Matthew Honnibal	f8c02b4341	Remove cupy imports from parser, so it can work on CPU	2017-05-14 00:37:53 +02:00
ines	1694c24e52	Add docstrings, error messages and fix consistency	2017-05-13 21:22:49 +02:00
ines	ee7dcf65c9	Fix expand_exc to make sure it returns combined dict	2017-05-13 21:22:25 +02:00
ines	824d09bb74	Move resolve_load_name to deprecated	2017-05-13 21:21:47 +02:00
ines	c4857bc7db	Remove unused argument	2017-05-12 15:37:54 +02:00
ines	86d9c29f30	Reorder util functions	2017-05-08 23:51:15 +02:00
ines	9a0d2fdef1	Add load_lang_class() util function	2017-05-08 23:50:45 +02:00
ines	2edc0aee12	Update warning message	2017-05-08 19:53:36 +02:00
ines	b9ba58ba5c	Add function to resolve load name Warn if old 'path' keyword argument is used.	2017-05-08 16:33:37 +02:00
ines	607ba458e7	Fix whitespace	2017-05-08 15:42:31 +02:00
ines	60db497525	Add update_exc and expand_exc to util Doesn't require separate language data util anymore	2017-05-08 15:42:12 +02:00
ines	95edd9e896	Let parse_package_meta take full path	2017-05-08 15:30:48 +02:00
ines	326746eb15	Add util function to resolve arg to model path 1. check if in data dir or shortcut link 2. check if installed as a pip package 3. check if string is path to model 4. check if Path or Path-like object	2017-05-08 15:29:47 +02:00
ines	94697e9afc	Fix typo	2017-05-08 02:00:37 +02:00
ines	c4492d260a	Fix kwargs	2017-05-08 01:05:24 +02:00
ines	59c3b9d4dd	Tidy up CLI and fix print functions	2017-05-07 23:25:29 +02:00
ines	e34069db9f	Move is_package and get_model_package_path to util	2017-05-07 23:24:51 +02:00
Ben Eyal	d8098a8be2	Use `regex` instead of `re`	2017-04-20 02:22:52 +03:00
ines	97647c46cd	Add docstring and todo note	2017-04-16 22:14:45 +02:00
ines	5c5f8c0a72	Check if full string is found in lang classes first This allows users to set arbitrary strings. (Otherwise, custom lang class "my_custom_class" would always load Burmese "my" tokenizer if one was available.)	2017-04-16 22:14:38 +02:00

1 2 3 4 5

245 Commits