Matthew Honnibal
|
70a7ad89ca
|
* Removed unused imports from train.py
|
2015-06-02 00:59:09 +02:00 |
|
Matthew Honnibal
|
75658b2ed3
|
* Remove use of new beam.loss property, to maintain compatibility with older versions of thinc for now.
|
2015-06-02 00:57:09 +02:00 |
|
Matthew Honnibal
|
a3de20118e
|
* Wire up beam-width command line argument
|
2015-06-02 00:54:12 +02:00 |
|
Matthew Honnibal
|
7c29362d60
|
* Rename parser class in parser.pxd, now that beam parsing is supported
|
2015-06-02 00:53:49 +02:00 |
|
Matthew Honnibal
|
58d5ac0944
|
* Add beam search capabilities to Parser. Rename GreedyParser to Parser.
|
2015-06-02 00:28:02 +02:00 |
|
Matthew Honnibal
|
62424e6c76
|
* Remove unused regularize argument from _ml.Model
|
2015-06-02 00:27:07 +02:00 |
|
Matthew Honnibal
|
adeb57cb1e
|
* Fix long line
|
2015-06-01 23:07:00 +02:00 |
|
Matthew Honnibal
|
e09a08bd00
|
* Add copy_state function
|
2015-06-01 23:06:30 +02:00 |
|
Matthew Honnibal
|
c7876aa8b6
|
* Add get_valid method
|
2015-06-01 23:06:00 +02:00 |
|
Matthew Honnibal
|
d82f9d958d
|
* Remove regularization cruft from _ml, move score from .pxd file to .pyx
|
2015-05-31 18:48:05 +02:00 |
|
Matthew Honnibal
|
08044ea70c
|
* Remove try/except around parser.train
|
2015-05-31 15:21:56 +02:00 |
|
Matthew Honnibal
|
c8a553fe91
|
* Fix cluster initialization
|
2015-05-31 15:21:28 +02:00 |
|
Matthew Honnibal
|
5e99ff94c8
|
* Edits to arc eager oracle. Couldn't figure out how the non-monotonic lines made sense. They seem covered by children_in_stack
|
2015-05-31 15:14:37 +02:00 |
|
Matthew Honnibal
|
6c5632b71c
|
* Roll back proposed change to Break transition while investigate effect
|
2015-05-31 06:49:52 +02:00 |
|
Matthew Honnibal
|
d7cc2338e7
|
* Fix bug in train.py
|
2015-05-31 06:49:06 +02:00 |
|
Matthew Honnibal
|
6bba793df3
|
* Disable the Zipf-reweighting thing while investigate effect
|
2015-05-31 06:48:43 +02:00 |
|
Matthew Honnibal
|
c037f80638
|
* Add case expansion to Brown clusters
|
2015-05-31 05:50:50 +02:00 |
|
Matthew Honnibal
|
5ab0f233a1
|
* Ensure words in Brown clusters make it into the vocab, even if they're not in our probs list
|
2015-05-31 05:46:16 +02:00 |
|
Matthew Honnibal
|
e77940565d
|
* Add length cap to distance feature
|
2015-05-31 05:25:30 +02:00 |
|
Matthew Honnibal
|
fd596351ba
|
* Fix valency features
|
2015-05-31 05:24:33 +02:00 |
|
Matthew Honnibal
|
d42dda0372
|
* Shuffle docs before doing jackknife partition --- otherwise we'll not get the right genre mixes...
|
2015-05-31 01:25:02 +02:00 |
|
Matthew Honnibal
|
4d8d490547
|
* Exclude empty sentences in prepare_treebank
|
2015-05-31 01:12:46 +02:00 |
|
Matthew Honnibal
|
87d6551d19
|
* Allow gold parse to cut non-projective arcs
|
2015-05-31 01:11:56 +02:00 |
|
Matthew Honnibal
|
d512d20d81
|
* Allow parser to jackknife POS tags before training.
|
2015-05-31 01:11:11 +02:00 |
|
Matthew Honnibal
|
c4f0914b4e
|
* Fix POS tag evaluation in scorer.py: do evaluate punctuation tags
|
2015-05-30 18:24:32 +02:00 |
|
Matthew Honnibal
|
9e39a206da
|
* Fix efficiency of JSON reading, by using ujson instead of stream
|
2015-05-30 17:54:52 +02:00 |
|
Matthew Honnibal
|
6bbdcc5db5
|
* Fix gold_preproc flag in train.py
|
2015-05-30 05:23:02 +02:00 |
|
Matthew Honnibal
|
76300bbb1b
|
* Use updated JSON format, with sentences below paragraphs. Allows use of gold preprocessing flag.
|
2015-05-30 01:25:46 +02:00 |
|
Matthew Honnibal
|
2d11739f28
|
* Change data format of JSON corpus, putting sentences into lists with the paragraph
|
2015-05-30 01:25:00 +02:00 |
|
Matthew Honnibal
|
784e577f45
|
* Check NER length matches conll length in prepare_treebank
|
2015-05-29 03:54:06 +02:00 |
|
Matthew Honnibal
|
b76bbbd12c
|
* Read json files recursively from a directory, instead of requiring a single .json file
|
2015-05-29 03:52:55 +02:00 |
|
Matthew Honnibal
|
8f31d3b864
|
* Relax constraint on Break transition for non-monotonic parsing.
|
2015-05-28 23:39:52 +02:00 |
|
Matthew Honnibal
|
ef67ef7a4c
|
* Recomment in training in train.py
|
2015-05-28 22:40:26 +02:00 |
|
Matthew Honnibal
|
5eb64eeb11
|
* Print json treebank by genre, instead of by large file
|
2015-05-28 22:40:01 +02:00 |
|
Matthew Honnibal
|
6b2e5c4b8a
|
* Avoid NER scoring for sentences with some missing NER values.
|
2015-05-28 22:39:08 +02:00 |
|
Matthew Honnibal
|
f42dc1f7d8
|
* Fix evaluate method in train.py, to use sentences which don't have raw text
|
2015-05-28 16:30:23 +02:00 |
|
Matthew Honnibal
|
d25d31442d
|
* Hackishly support broken NER annotations. Should fix this.
|
2015-05-27 19:14:31 +02:00 |
|
Matthew Honnibal
|
a7cee46fe9
|
* Update train.py, to support paragraphs where there's no raw_text
|
2015-05-27 19:14:02 +02:00 |
|
Matthew Honnibal
|
7a2725bca4
|
* Read input json in a streaming way
|
2015-05-27 19:13:11 +02:00 |
|
Matthew Honnibal
|
b7fd77779a
|
* Add some tests for reading NER data
|
2015-05-27 17:37:03 +02:00 |
|
Matthew Honnibal
|
6a1c91675e
|
* Add file to read ENAMEX ner data
|
2015-05-27 17:36:23 +02:00 |
|
Matthew Honnibal
|
ef1333cf89
|
* Have prepare_treebank read train/dev/test IDs.
|
2015-05-27 17:35:05 +02:00 |
|
Matthew Honnibal
|
e140e03516
|
* Read in OntoNotes. Doesn't support train/test/dev split yet
|
2015-05-27 17:04:29 +02:00 |
|
Matthew Honnibal
|
732fa7709a
|
* Edits to align_raw script, for use in prepare_treebank
|
2015-05-27 04:23:31 +02:00 |
|
Matthew Honnibal
|
4010b9b6d9
|
* Pass parameter for regularization in parser.pyx
|
2015-05-27 03:18:50 +02:00 |
|
Matthew Honnibal
|
4c6058baa7
|
* Fix evaluation of NER in scorer.py
|
2015-05-27 03:18:16 +02:00 |
|
Matthew Honnibal
|
6016ee83a6
|
* Fix reading of NER in gold.pyx
|
2015-05-27 03:17:50 +02:00 |
|
Matthew Honnibal
|
04bda8648d
|
* Pass parameter for regularization to model
|
2015-05-27 03:16:58 +02:00 |
|
Matthew Honnibal
|
895060e774
|
* Ensure tagger and NER are trained, even if non-projective problem
|
2015-05-27 03:16:21 +02:00 |
|
Matthew Honnibal
|
f69fe6a635
|
* Fix heads problem in read_conll
|
2015-05-27 01:14:54 +02:00 |
|