Søren Lind Kristiansen
|
0ffd27b0f6
|
Add several Danish alternative spellings
|
2017-11-27 13:35:41 +01:00 |
|
Ines Montani
|
a7bb8f1b42
|
Merge pull request #1637 from sorenlind/da_tokenization
Improve Danish tokenization
|
2017-11-26 15:41:38 +00:00 |
|
ines
|
c699aec089
|
Add offsets_from_biluo_tags helper and tests (see #1626)
|
2017-11-26 16:38:01 +01:00 |
|
Søren Lind Kristiansen
|
6aa241bcec
|
Add day of month tokenizer exceptions for Danish.
|
2017-11-24 15:03:24 +01:00 |
|
Søren Lind Kristiansen
|
0c276ed020
|
Add weekday abbreviations and remove abiguous month abbreviations for Danish.
|
2017-11-24 14:43:29 +01:00 |
|
Søren Lind Kristiansen
|
056547e989
|
Add multiple tokenizer exceptions for Danish.
|
2017-11-24 11:51:26 +01:00 |
|
Søren Lind Kristiansen
|
8dc265ac0c
|
Add test for tokenization of 'i.' for Danish.
|
2017-11-24 11:29:37 +01:00 |
|
Matthew Honnibal
|
30ba81f881
|
Merge pull request #1576 from ligser/master
Actually reset caches in pipe [wip]
|
2017-11-23 12:54:48 +01:00 |
|
ines
|
c90fe92e15
|
Fix displaCy test
|
2017-11-22 05:04:39 +01:00 |
|
ines
|
a6f33ac27d
|
Fix displaCy test
|
2017-11-22 04:19:28 +01:00 |
|
Burton DeWilde
|
635792997c
|
Add regression test for #1612
|
2017-11-20 12:05:35 -06:00 |
|
ines
|
d70a64d78b
|
Fix syntax error and formatting in test (see #1617)
|
2017-11-20 14:01:25 +01:00 |
|
ines
|
17849dee4b
|
Fix French test (see #1617)
|
2017-11-20 13:59:59 +01:00 |
|
Motoki Wu
|
b818afaa0e
|
Added failing test for Issue #1207.
The noun chunk iterator should work for `Doc` but not for `Span`.
|
2017-11-17 17:04:27 -08:00 |
|
ines
|
a3d4dd1a5d
|
Test adding of lots of pipeline components (see #1585)
Just to make sure that there's no error now or in the future with adding a large number of pipeline components.
|
2017-11-15 17:28:06 +01:00 |
|
Roman Domrachev
|
505c6a2f2f
|
Completely cleanup tokenizer cache
Tokenizer cache can have be different keys than string
That modification can slow down tokenizer and need to be measured
|
2017-11-15 17:55:48 +03:00 |
|
Roman Domrachev
|
3e21680814
|
Use safer method to get string without hit
|
2017-11-14 22:58:46 +03:00 |
|
Roman Domrachev
|
4e378dc4a4
|
Remove all obsolete code and test only initial problem
|
2017-11-14 20:45:04 +03:00 |
|
Roman
|
47ce2347b0
|
Create test that fails when actual cleanup caused
|
2017-11-14 20:28:13 +03:00 |
|
Roman Domrachev
|
3d247d2bb8
|
Get back previous testcase
|
2017-11-14 18:01:37 +03:00 |
|
Roman Domrachev
|
a2745b0e84
|
StringStore now actually cleaned
Do not lose docs in ref tracking
|
2017-11-14 17:45:50 +03:00 |
|
Roman Domrachev
|
ee60a52ee7
|
Fix test imports and last batch cleanup
|
2017-11-11 11:32:16 +03:00 |
|
Roman Domrachev
|
3c600adf23
|
Try to fix StringStore clean up (see #1506)
|
2017-11-11 03:11:27 +03:00 |
|
ines
|
ee97fd3cb4
|
Add regression test for #1547
|
2017-11-11 00:14:03 +01:00 |
|
ines
|
2df27db671
|
Add unicode declaration
|
2017-11-11 00:13:56 +01:00 |
|
ines
|
1c218397f6
|
Ensure path in Doc.to_disk/from_disk (resolves ##1521)
Also add Doc serialization tests with both Path and string path options
|
2017-11-09 02:29:03 +01:00 |
|
Matthew Honnibal
|
a5ea0fdf5a
|
Fix #1518: vocab.vectors.resize() didn't work
|
2017-11-08 22:18:37 +01:00 |
|
Matthew Honnibal
|
4194bc5744
|
Xfail flakey serialization test
|
2017-11-08 13:55:13 +01:00 |
|
ines
|
42a0fbf291
|
Fix textcat simple train example
|
2017-11-07 01:25:54 +01:00 |
|
ines
|
5f43953536
|
Move test
|
2017-11-06 23:14:10 +01:00 |
|
Matthew Honnibal
|
1831dbd065
|
Add test of simple textcat workflow
|
2017-11-06 22:04:29 +01:00 |
|
Matthew Honnibal
|
2f7e9f390d
|
Make test less flakey
|
2017-11-06 17:34:50 +01:00 |
|
Matthew Honnibal
|
407b08017e
|
Make test less flakey
|
2017-11-06 17:31:40 +01:00 |
|
Matthew Honnibal
|
102f797933
|
Fix lemma ordering in test
|
2017-11-06 17:02:17 +01:00 |
|
Matthew Honnibal
|
63c6ae4191
|
Fix lemmatizer test
|
2017-11-06 11:57:06 +01:00 |
|
Matthew Honnibal
|
00435d8f0c
|
Add extra beam parsing test
|
2017-11-05 14:39:57 +01:00 |
|
ines
|
5e7d98f72a
|
Remove test for #1491
|
2017-11-03 22:10:57 +01:00 |
|
ines
|
718f1c50fb
|
Add regression test for #1491
|
2017-11-03 21:11:20 +01:00 |
|
Matthew Honnibal
|
144a93c2a5
|
Back-off to tensor for similarity if no vectors
|
2017-11-03 20:56:33 +01:00 |
|
Matthew Honnibal
|
d6e831bf89
|
Fix lemmatizer tests
|
2017-11-03 19:46:34 +01:00 |
|
ines
|
eef930c73e
|
Assert instead of print
|
2017-11-03 18:50:57 +01:00 |
|
ines
|
f0986df94b
|
Add test for #1488 (passes on v2.0.0a18?)
|
2017-11-03 14:44:36 +01:00 |
|
Matthew Honnibal
|
711278b667
|
Make test less flakey
|
2017-11-03 14:36:08 +01:00 |
|
Matthew Honnibal
|
0a534ae96a
|
Fix test for backprop d_pad
|
2017-11-03 14:04:16 +01:00 |
|
Matthew Honnibal
|
a22f96c3f1
|
Add test for backpropagating padding
|
2017-11-03 00:48:54 +01:00 |
|
ines
|
3af281a334
|
Update test model name
|
2017-11-01 23:02:00 +01:00 |
|
ines
|
8c2260e18c
|
Move span tests to /doc
|
2017-11-01 16:56:35 +01:00 |
|
ines
|
260cb37224
|
Catch deprecation warning
|
2017-11-01 16:49:18 +01:00 |
|
ines
|
5914faafbb
|
Fix .merge tests to not use deprecated API
|
2017-11-01 16:49:11 +01:00 |
|
Matthew Honnibal
|
9e0ebee81c
|
Add Token.is_sent_start property, so can deprecate Token.sent_start
|
2017-11-01 13:27:14 +01:00 |
|