spaCy

mirror of https://github.com/explosion/spaCy.git synced 2025-09-10 14:12:37 +03:00

Author	SHA1	Message	Date
Matthew Honnibal	5ab0f233a1	* Ensure words in Brown clusters make it into the vocab, even if they're not in our probs list	2015-05-31 05:46:16 +02:00
Matthew Honnibal	e77940565d	* Add length cap to distance feature	2015-05-31 05:25:30 +02:00
Matthew Honnibal	fd596351ba	* Fix valency features	2015-05-31 05:24:33 +02:00
Matthew Honnibal	d42dda0372	* Shuffle docs before doing jackknife partition --- otherwise we'll not get the right genre mixes...	2015-05-31 01:25:02 +02:00
Matthew Honnibal	4d8d490547	* Exclude empty sentences in prepare_treebank	2015-05-31 01:12:46 +02:00
Matthew Honnibal	87d6551d19	* Allow gold parse to cut non-projective arcs	2015-05-31 01:11:56 +02:00
Matthew Honnibal	d512d20d81	* Allow parser to jackknife POS tags before training.	2015-05-31 01:11:11 +02:00
Matthew Honnibal	c4f0914b4e	* Fix POS tag evaluation in scorer.py: do evaluate punctuation tags	2015-05-30 18:24:32 +02:00
Matthew Honnibal	9e39a206da	* Fix efficiency of JSON reading, by using ujson instead of stream	2015-05-30 17:54:52 +02:00
Matthew Honnibal	6bbdcc5db5	* Fix gold_preproc flag in train.py	2015-05-30 05:23:02 +02:00
Matthew Honnibal	76300bbb1b	* Use updated JSON format, with sentences below paragraphs. Allows use of gold preprocessing flag.	2015-05-30 01:25:46 +02:00
Matthew Honnibal	2d11739f28	* Change data format of JSON corpus, putting sentences into lists with the paragraph	2015-05-30 01:25:00 +02:00
Matthew Honnibal	784e577f45	* Check NER length matches conll length in prepare_treebank	2015-05-29 03:54:06 +02:00
Matthew Honnibal	b76bbbd12c	* Read json files recursively from a directory, instead of requiring a single .json file	2015-05-29 03:52:55 +02:00
Matthew Honnibal	8f31d3b864	* Relax constraint on Break transition for non-monotonic parsing.	2015-05-28 23:39:52 +02:00
Matthew Honnibal	ef67ef7a4c	* Recomment in training in train.py	2015-05-28 22:40:26 +02:00
Matthew Honnibal	5eb64eeb11	* Print json treebank by genre, instead of by large file	2015-05-28 22:40:01 +02:00
Matthew Honnibal	6b2e5c4b8a	* Avoid NER scoring for sentences with some missing NER values.	2015-05-28 22:39:08 +02:00
Matthew Honnibal	f42dc1f7d8	* Fix evaluate method in train.py, to use sentences which don't have raw text	2015-05-28 16:30:23 +02:00
Matthew Honnibal	d25d31442d	* Hackishly support broken NER annotations. Should fix this.	2015-05-27 19:14:31 +02:00
Matthew Honnibal	a7cee46fe9	* Update train.py, to support paragraphs where there's no raw_text	2015-05-27 19:14:02 +02:00
Matthew Honnibal	7a2725bca4	* Read input json in a streaming way	2015-05-27 19:13:11 +02:00
Matthew Honnibal	b7fd77779a	* Add some tests for reading NER data	2015-05-27 17:37:03 +02:00
Matthew Honnibal	6a1c91675e	* Add file to read ENAMEX ner data	2015-05-27 17:36:23 +02:00
Matthew Honnibal	ef1333cf89	* Have prepare_treebank read train/dev/test IDs.	2015-05-27 17:35:05 +02:00
Matthew Honnibal	e140e03516	* Read in OntoNotes. Doesn't support train/test/dev split yet	2015-05-27 17:04:29 +02:00
Matthew Honnibal	732fa7709a	* Edits to align_raw script, for use in prepare_treebank	2015-05-27 04:23:31 +02:00
Matthew Honnibal	4010b9b6d9	* Pass parameter for regularization in parser.pyx	2015-05-27 03:18:50 +02:00
Matthew Honnibal	4c6058baa7	* Fix evaluation of NER in scorer.py	2015-05-27 03:18:16 +02:00
Matthew Honnibal	6016ee83a6	* Fix reading of NER in gold.pyx	2015-05-27 03:17:50 +02:00
Matthew Honnibal	04bda8648d	* Pass parameter for regularization to model	2015-05-27 03:16:58 +02:00
Matthew Honnibal	895060e774	* Ensure tagger and NER are trained, even if non-projective problem	2015-05-27 03:16:21 +02:00
Matthew Honnibal	f69fe6a635	* Fix heads problem in read_conll	2015-05-27 01:14:54 +02:00
Matthew Honnibal	0eec1d12af	* Add comment about zipf reweighting	2015-05-27 01:14:07 +02:00
Matthew Honnibal	4d37b66c55	* Make Zipf regularization a bit more efficient	2015-05-27 01:12:50 +02:00
Matthew Honnibal	7fc24821bc	* Experiment with Zipfian corruptions when calculating prediction	2015-05-26 22:17:15 +02:00
Matthew Honnibal	32ae2cdabe	* In prepare_treebank, move ner into the token descriptions	2015-05-26 19:52:39 +02:00
Matthew Honnibal	61885aee76	* Work on prepare_treebank script, adding NER to it	2015-05-26 19:28:29 +02:00
Matthew Honnibal	15bbbf4901	* Remove cruft from train.py	2015-05-25 07:54:10 +02:00
Matthew Honnibal	eba7b34f66	* Add flag to disable loading of word vectors	2015-05-25 01:02:42 +02:00
Matthew Honnibal	89c3364041	* Update tests, preventing the parser from being loaded if possible	2015-05-25 01:02:03 +02:00
Matthew Honnibal	a9c70c9447	* Add tests for ontonotes sgml extraction	2015-05-24 21:52:12 +02:00
Matthew Honnibal	f460a8d2b6	* Comment out failing test in test_conjuncts	2015-05-24 21:51:41 +02:00
Matthew Honnibal	cc7439a16b	* Don't use alignment.pyx file, move functionality to spacy.gold	2015-05-24 21:51:15 +02:00
Matthew Honnibal	3593babd35	* Add functions for Levenshtein distance alignment	2015-05-24 21:50:48 +02:00
Matthew Honnibal	744f06abf5	* Add script to read OntoNotes source documents	2015-05-24 21:49:58 +02:00
Matthew Honnibal	13a8595a4b	* Add tests for Levenshtein alignment of training data	2015-05-24 21:46:11 +02:00
Matthew Honnibal	fc75210941	* Move spacy.syntax.conll to spacy.gold	2015-05-24 21:35:02 +02:00
Matthew Honnibal	765b61cac4	* Update spacy.scorer, to use P/R/F to support tokenization errors	2015-05-24 20:07:18 +02:00
Matthew Honnibal	efe7a7d7d6	* Clean unused functions from spacy.syntax.conll	2015-05-24 20:06:46 +02:00

... 7 8 9 10 11 ...

1512 Commits