spaCy

mirror of https://github.com/explosion/spaCy.git synced 2024-09-22 03:49:17 +03:00

Author	SHA1	Message	Date
Matthew Honnibal	62cfcd76fe	* Add supersense sets to lexemes, from WordNet. Look-up via lemmatization.	2015-07-01 18:48:59 +02:00
Matthew Honnibal	ef97b90833	* Fix token scoring	2015-06-28 06:22:18 +02:00
Matthew Honnibal	34c0ef2ee8	* Don't compile the orig_arc_eager and tree_arc_eager modules used for the EMNLP paper	2015-06-23 05:38:17 +02:00
Matthew Honnibal	59e9f9153c	* Remove projectivity constraint in train.py, but raise Exception if non-projective sentence is encountered, since we've told GoldParse to projectivize	2015-06-23 05:04:46 +02:00
Matthew Honnibal	839e5038b7	* Raise exception on non-projective input	2015-06-23 00:01:55 +02:00
Matthew Honnibal	4dad4058c3	* Uncomment NER training	2015-06-16 23:36:54 +02:00
Matthew Honnibal	5699585278	* Use tree_arc_eager system as baseline in experiments	2015-06-15 08:23:43 +02:00
Matthew Honnibal	4841f8ad5e	* Set transition system early	2015-06-15 02:54:12 +02:00
Matthew Honnibal	bcfdf126a4	* Add toggle for OrigArcEager system	2015-06-14 20:28:14 +02:00
Matthew Honnibal	c500d72dc2	* Temporarily disable NER, and wire up the verbose flag during training	2015-06-14 17:45:31 +02:00
Matthew Honnibal	ac422492cf	* Fix write_parses mode of bin/parser/train.py	2015-06-07 19:08:48 +02:00
Matthew Honnibal	4073533e28	* Upd munge_ewtb for the new json format	2015-06-06 02:10:33 +02:00
Matthew Honnibal	6a1341b29e	* Add tb pre-process script	2015-06-06 01:59:44 +02:00
Matthew Honnibal	1736fc5a67	* Add more options to bin/parser/train	2015-06-05 23:49:26 +02:00
Matthew Honnibal	362f87dc3a	* Update input corruption method to work with lists as well as trings	2015-06-05 19:33:32 +02:00
Matthew Honnibal	0aed9c9a33	* Fix train.py	2015-06-05 15:50:24 +02:00
Matthew Honnibal	8466600add	* Clean up train.py, removing unused tag jackknifing code	2015-06-05 15:01:28 +02:00
Matthew Honnibal	e772b48dcd	* Skip sentences of length 1 in training	2015-06-05 02:29:03 +02:00
Matthew Honnibal	e822df0867	* Fix bugs in new greedy/beam parser	2015-06-02 02:01:33 +02:00
Matthew Honnibal	70a7ad89ca	* Removed unused imports from train.py	2015-06-02 00:59:09 +02:00
Matthew Honnibal	a3de20118e	* Wire up beam-width command line argument	2015-06-02 00:54:12 +02:00
Matthew Honnibal	08044ea70c	* Remove try/except around parser.train	2015-05-31 15:21:56 +02:00
Matthew Honnibal	c8a553fe91	* Fix cluster initialization	2015-05-31 15:21:28 +02:00
Matthew Honnibal	d7cc2338e7	* Fix bug in train.py	2015-05-31 06:49:06 +02:00
Matthew Honnibal	c037f80638	* Add case expansion to Brown clusters	2015-05-31 05:50:50 +02:00
Matthew Honnibal	5ab0f233a1	* Ensure words in Brown clusters make it into the vocab, even if they're not in our probs list	2015-05-31 05:46:16 +02:00
Matthew Honnibal	d42dda0372	* Shuffle docs before doing jackknife partition --- otherwise we'll not get the right genre mixes...	2015-05-31 01:25:02 +02:00
Matthew Honnibal	4d8d490547	* Exclude empty sentences in prepare_treebank	2015-05-31 01:12:46 +02:00
Matthew Honnibal	d512d20d81	* Allow parser to jackknife POS tags before training.	2015-05-31 01:11:11 +02:00
Matthew Honnibal	6bbdcc5db5	* Fix gold_preproc flag in train.py	2015-05-30 05:23:02 +02:00
Matthew Honnibal	76300bbb1b	* Use updated JSON format, with sentences below paragraphs. Allows use of gold preprocessing flag.	2015-05-30 01:25:46 +02:00
Matthew Honnibal	2d11739f28	* Change data format of JSON corpus, putting sentences into lists with the paragraph	2015-05-30 01:25:00 +02:00
Matthew Honnibal	784e577f45	* Check NER length matches conll length in prepare_treebank	2015-05-29 03:54:06 +02:00
Matthew Honnibal	b76bbbd12c	* Read json files recursively from a directory, instead of requiring a single .json file	2015-05-29 03:52:55 +02:00
Matthew Honnibal	ef67ef7a4c	* Recomment in training in train.py	2015-05-28 22:40:26 +02:00
Matthew Honnibal	5eb64eeb11	* Print json treebank by genre, instead of by large file	2015-05-28 22:40:01 +02:00
Matthew Honnibal	f42dc1f7d8	* Fix evaluate method in train.py, to use sentences which don't have raw text	2015-05-28 16:30:23 +02:00
Matthew Honnibal	a7cee46fe9	* Update train.py, to support paragraphs where there's no raw_text	2015-05-27 19:14:02 +02:00
Matthew Honnibal	ef1333cf89	* Have prepare_treebank read train/dev/test IDs.	2015-05-27 17:35:05 +02:00
Matthew Honnibal	e140e03516	* Read in OntoNotes. Doesn't support train/test/dev split yet	2015-05-27 17:04:29 +02:00
Matthew Honnibal	895060e774	* Ensure tagger and NER are trained, even if non-projective problem	2015-05-27 03:16:21 +02:00
Matthew Honnibal	32ae2cdabe	* In prepare_treebank, move ner into the token descriptions	2015-05-26 19:52:39 +02:00
Matthew Honnibal	61885aee76	* Work on prepare_treebank script, adding NER to it	2015-05-26 19:28:29 +02:00
Matthew Honnibal	15bbbf4901	* Remove cruft from train.py	2015-05-25 07:54:10 +02:00
Matthew Honnibal	fc75210941	* Move spacy.syntax.conll to spacy.gold	2015-05-24 21:35:02 +02:00
Matthew Honnibal	541c62c126	* Remove import of removed read_docparse_file function	2015-05-24 20:05:13 +02:00
Matthew Honnibal	bfeb29ebd1	* Tmp commit	2015-05-24 02:50:14 +02:00
Matthew Honnibal	983d954ef4	* Tmp commit, while switch to new format that assumes alignment happens during training	2015-05-23 17:39:04 +02:00
Matthew Honnibal	f35503018e	* Tmp commit of train, while I move to better alignment in gold standard	2015-05-23 17:21:25 +02:00
Matthew Honnibal	3d6b3fc6fb	* Restore shuffling, and remove print statements from train.py	2015-05-12 20:27:56 +02:00

1 2

96 Commits