| 
							
							
								 Matthew Honnibal | c4f0914b4e | * Fix POS tag evaluation in scorer.py: do evaluate punctuation tags | 2015-05-30 18:24:32 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | 9e39a206da | * Fix efficiency of JSON reading, by using ujson instead of stream | 2015-05-30 17:54:52 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | 6bbdcc5db5 | * Fix gold_preproc flag in train.py | 2015-05-30 05:23:02 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | 76300bbb1b | * Use updated JSON format, with sentences below paragraphs. Allows use of gold preprocessing flag. | 2015-05-30 01:25:46 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | 2d11739f28 | * Change data format of JSON corpus, putting sentences into lists with the paragraph | 2015-05-30 01:25:00 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | 784e577f45 | * Check NER length matches conll length in prepare_treebank | 2015-05-29 03:54:06 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | b76bbbd12c | * Read json files recursively from a directory, instead of requiring a single .json file | 2015-05-29 03:52:55 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | 8f31d3b864 | * Relax constraint on Break transition for non-monotonic parsing. | 2015-05-28 23:39:52 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | ef67ef7a4c | * Recomment in training in train.py | 2015-05-28 22:40:26 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | 5eb64eeb11 | * Print json treebank by genre, instead of by large file | 2015-05-28 22:40:01 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | 6b2e5c4b8a | * Avoid NER scoring for sentences with some missing NER values. | 2015-05-28 22:39:08 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | f42dc1f7d8 | * Fix evaluate method in train.py, to use sentences which don't have raw text | 2015-05-28 16:30:23 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | d25d31442d | * Hackishly support broken NER annotations. Should fix this. | 2015-05-27 19:14:31 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | a7cee46fe9 | * Update train.py, to support paragraphs where there's no raw_text | 2015-05-27 19:14:02 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | 7a2725bca4 | * Read input json in a streaming way | 2015-05-27 19:13:11 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | b7fd77779a | * Add some tests for reading NER data | 2015-05-27 17:37:03 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | 6a1c91675e | * Add file to read ENAMEX ner data | 2015-05-27 17:36:23 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | ef1333cf89 | * Have prepare_treebank read train/dev/test IDs. | 2015-05-27 17:35:05 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | e140e03516 | * Read in OntoNotes. Doesn't support train/test/dev split yet | 2015-05-27 17:04:29 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | 732fa7709a | * Edits to align_raw script, for use in prepare_treebank | 2015-05-27 04:23:31 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | 4010b9b6d9 | * Pass parameter for regularization in parser.pyx | 2015-05-27 03:18:50 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | 4c6058baa7 | * Fix evaluation of NER in scorer.py | 2015-05-27 03:18:16 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | 6016ee83a6 | * Fix reading of NER in gold.pyx | 2015-05-27 03:17:50 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | 04bda8648d | * Pass parameter for regularization to model | 2015-05-27 03:16:58 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | 895060e774 | * Ensure tagger and NER are trained, even if non-projective problem | 2015-05-27 03:16:21 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | f69fe6a635 | * Fix heads problem in read_conll | 2015-05-27 01:14:54 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | 0eec1d12af | * Add comment about zipf reweighting | 2015-05-27 01:14:07 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | 4d37b66c55 | * Make Zipf regularization a bit more efficient | 2015-05-27 01:12:50 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | 7fc24821bc | * Experiment with Zipfian corruptions when calculating prediction | 2015-05-26 22:17:15 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | 32ae2cdabe | * In prepare_treebank, move ner into the token descriptions | 2015-05-26 19:52:39 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | 61885aee76 | * Work on prepare_treebank script, adding NER to it | 2015-05-26 19:28:29 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | 15bbbf4901 | * Remove cruft from train.py | 2015-05-25 07:54:10 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | eba7b34f66 | * Add flag to disable loading of word vectors | 2015-05-25 01:02:42 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | 89c3364041 | * Update tests, preventing the parser from being loaded if possible | 2015-05-25 01:02:03 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | a9c70c9447 | * Add tests for ontonotes sgml extraction | 2015-05-24 21:52:12 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | f460a8d2b6 | * Comment out failing test in test_conjuncts | 2015-05-24 21:51:41 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | cc7439a16b | * Don't use alignment.pyx file, move functionality to spacy.gold | 2015-05-24 21:51:15 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | 3593babd35 | * Add functions for Levenshtein distance alignment | 2015-05-24 21:50:48 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | 744f06abf5 | * Add script to read OntoNotes source documents | 2015-05-24 21:49:58 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | 13a8595a4b | * Add tests for Levenshtein alignment of training data | 2015-05-24 21:46:11 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | fc75210941 | * Move spacy.syntax.conll to spacy.gold | 2015-05-24 21:35:02 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | 765b61cac4 | * Update spacy.scorer, to use P/R/F to support tokenization errors | 2015-05-24 20:07:18 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | efe7a7d7d6 | * Clean unused functions from spacy.syntax.conll | 2015-05-24 20:06:46 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | 78487f3e66 | * Update parser oracle for missing heads | 2015-05-24 20:05:58 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | 541c62c126 | * Remove import of removed read_docparse_file function | 2015-05-24 20:05:13 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | 1044a13413 | * Begin refactoring scorer to use recall over gold dependencies | 2015-05-24 17:40:15 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | acd1245ad4 | * Remove cruft from conll.pyx --- unused stuff about evlauation, which now lives in spacy.scorer | 2015-05-24 17:35:49 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | bfeb29ebd1 | * Tmp commit | 2015-05-24 02:50:14 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | 20f1d868a3 | * Tmp commit. Working on whole document parsing | 2015-05-24 02:49:56 +02:00 |  | 
			
				
					| 
							
							
								 Matthew Honnibal | 983d954ef4 | * Tmp commit, while switch to new format that assumes alignment happens during training | 2015-05-23 17:39:04 +02:00 |  |