Matthew Honnibal
|
468ca6c760
|
Merge branch 'develop' of https://github.com/explosion/spaCy into develop
|
2017-06-03 14:33:51 -05:00 |
|
Matthew Honnibal
|
c647a0d33e
|
Fix training counter for gold preprocessing
|
2017-06-03 14:33:39 -05:00 |
|
ines
|
e47eef5e03
|
Update German tokenizer exceptions and tests
|
2017-06-03 21:07:44 +02:00 |
|
ines
|
d77c2cc8bb
|
Add tests for English norm exceptions
|
2017-06-03 20:59:50 +02:00 |
|
ines
|
0d6fa8b241
|
Add German norm exceptions
|
2017-06-03 20:54:18 +02:00 |
|
ines
|
5bd311c77e
|
Fix update of norm exceptions
|
2017-06-03 20:54:09 +02:00 |
|
Matthew Honnibal
|
94e063ae2a
|
Merge branch 'develop' of https://github.com/explosion/spaCy into develop
|
2017-06-03 13:31:40 -05:00 |
|
Matthew Honnibal
|
fea1144e6d
|
Set max batch size in evaluate
|
2017-06-03 13:31:33 -05:00 |
|
Matthew Honnibal
|
805495af27
|
Fix off-by-one in number of tags
|
2017-06-03 13:29:23 -05:00 |
|
Matthew Honnibal
|
e62f46d39f
|
Clarify gold.pyx slightly
|
2017-06-03 13:28:52 -05:00 |
|
Matthew Honnibal
|
43353b5413
|
Improve train CLI script
|
2017-06-03 13:28:20 -05:00 |
|
ines
|
746653880c
|
Add English norm exceptions to lex_attrs
|
2017-06-03 20:27:28 +02:00 |
|
ines
|
095eeeb12f
|
Update English tokenizer exceptions and add norms
|
2017-06-03 20:27:16 +02:00 |
|
ines
|
e5d426406a
|
Add base norm exceptions
|
2017-06-03 20:27:05 +02:00 |
|
ines
|
4c2bbc3ccc
|
Add add_lookups util function
|
2017-06-03 19:44:47 +02:00 |
|
ines
|
05fe6758a7
|
Set lexeme attributes for tokenizer special cases
|
2017-06-03 19:44:39 +02:00 |
|
ines
|
3152ee5ca2
|
Update serialization tests for tokenizer
|
2017-06-03 17:05:28 +02:00 |
|
ines
|
7c919aeb09
|
Make sure serializers and deserializers are ordered
|
2017-06-03 17:05:09 +02:00 |
|
ines
|
1ebd0d3f27
|
Add assert_packed_msg_equal util function
|
2017-06-03 17:04:30 +02:00 |
|
ines
|
de974f7bef
|
Add serializer tests for tokenizer
|
2017-06-03 13:26:34 +02:00 |
|
ines
|
0153b66a86
|
Return self in Tokenizer.from_bytes
|
2017-06-03 13:26:13 +02:00 |
|
ines
|
82154a1861
|
Add letter spacing to arrow label
|
2017-06-03 13:25:41 +02:00 |
|
ines
|
32c6f05de9
|
Adjust spacing and sizing in compact mode
|
2017-06-03 13:25:32 +02:00 |
|
ines
|
cc8c8617a4
|
Shut down displaCy server on KeyboardInterrupt
|
2017-06-03 13:24:56 +02:00 |
|
ines
|
70fbba7d08
|
Clone Doc to never merge punctuation on original Doc
|
2017-06-03 13:24:43 +02:00 |
|
ines
|
459a1e8470
|
Fix whitespace
|
2017-06-03 11:31:18 +02:00 |
|
ines
|
5109bba910
|
Port over fix from #1070
|
2017-06-03 11:31:11 +02:00 |
|
ines
|
d21459f87d
|
Update serializer tests
|
2017-06-02 21:42:26 +02:00 |
|
ines
|
6669583f4e
|
Use OrderedDict
|
2017-06-02 21:07:56 +02:00 |
|
ines
|
2f1025a94c
|
Port over Spanish changes from #1096
|
2017-06-02 19:09:58 +02:00 |
|
ines
|
d86e7cde93
|
Add entity recognizer to parser serialization tests
|
2017-06-02 18:40:06 +02:00 |
|
ines
|
0051c05964
|
Add tests for serializing parser
|
2017-06-02 18:37:19 +02:00 |
|
ines
|
fdd0923be4
|
Translate model=True in exclude to lower_model and upper_model
|
2017-06-02 18:37:07 +02:00 |
|
ines
|
cef547a9f0
|
Add serialization tests for tensorizer
|
2017-06-02 18:18:30 +02:00 |
|
ines
|
924c58bde3
|
Fix serialization of optional elements
|
2017-06-02 18:18:17 +02:00 |
|
ines
|
f74a45c1fe
|
Remove unnecessary argument
|
2017-06-02 18:17:46 +02:00 |
|
ines
|
43b4d63f85
|
Add serialization tests for tagger
|
2017-06-02 17:29:34 +02:00 |
|
ines
|
1b593bbd6d
|
Fix encoding on tagger serialization
|
2017-06-02 17:29:21 +02:00 |
|
Matthew Honnibal
|
5f4d328e2c
|
Fix serialization of tag_map in NeuralTagger
|
2017-06-02 10:18:37 -05:00 |
|
Matthew Honnibal
|
ed6f575e06
|
Merge branch 'develop' of https://github.com/explosion/spaCy into develop
|
2017-06-02 04:26:39 -05:00 |
|
ines
|
acd65c00f6
|
Add serialization tests for StringStore and Vocab
|
2017-06-02 10:57:42 +02:00 |
|
ines
|
41a6adf1f6
|
Initialise Vocab length correctly
|
2017-06-02 10:57:25 +02:00 |
|
ines
|
53b82f972a
|
Add strings to Vocab in init, instead of StringStore
|
2017-06-02 10:57:06 +02:00 |
|
ines
|
023f38bdd4
|
Fix return value of Vocab.from_bytes
|
2017-06-02 10:56:40 +02:00 |
|
ines
|
9692c98f57
|
Add test utils for temp file and temp dir
|
2017-06-02 10:56:09 +02:00 |
|
Matthew Honnibal
|
c650bc481c
|
Merge branch 'develop' of https://github.com/explosion/spaCy into develop
|
2017-06-01 13:03:57 -05:00 |
|
Matthew Honnibal
|
307d615c5f
|
Fix serialization for tagger when tag_map has changed
|
2017-06-01 12:18:36 -05:00 |
|
Matthew Honnibal
|
1d18cedae8
|
Fiddle with msgpack bytes vs unicode
|
2017-06-01 10:48:43 -05:00 |
|
ines
|
7a2380f617
|
Rename "nn_tagger" to "tagger"
|
2017-06-01 17:37:53 +02:00 |
|
ines
|
e5ae6ccf4e
|
Fix typo
|
2017-06-01 16:46:15 +02:00 |
|