Matthew Honnibal
|
36b47e3fa6
|
Fix (and test) vector pickling
|
2017-12-07 09:53:30 +01:00 |
|
Matthew Honnibal
|
05f41ff587
|
Set version to 2.0.4
|
2017-12-06 13:24:02 +01:00 |
|
Matthew Honnibal
|
04c38f7e87
|
Merge branch 'master' of https://github.com/explosion/spaCy
|
2017-12-06 12:15:52 +01:00 |
|
Matthew Honnibal
|
361944e512
|
If no rules are set, lemmatize by lookup
|
2017-12-06 12:12:11 +01:00 |
|
Matthew Honnibal
|
2ab0f2d186
|
Merge pull request #1664 from jimregan/italian-lemmatizer
BOM in Italian lemmatiser
|
2017-12-06 11:09:04 +01:00 |
|
Matthew Honnibal
|
3f247119d3
|
Merge pull request #1668 from sorenlind/da_morph
Add more Danish morph rules and clean up existing ones
|
2017-12-06 11:08:09 +01:00 |
|
Matthew Honnibal
|
b712de774e
|
Fix vectors pickling
|
2017-12-05 12:45:24 +01:00 |
|
Matthew Honnibal
|
04650e38c7
|
Set version to 2.0.4.dev0
|
2017-12-05 10:52:31 +01:00 |
|
Matthew Honnibal
|
07acb43a85
|
Merge branch 'master' of https://github.com/explosion/spaCy
|
2017-12-04 14:42:52 +01:00 |
|
Thomas Werkmeister
|
94eac75b7c
|
fix setup.py spacy req string for packaging
Requirement should be `spacy>=2.0.2` instead of `spacy2.0.2`
|
2017-12-03 04:16:28 -06:00 |
|
ines
|
f2ea6d4713
|
Add Dutch example sentences (see #1107)
|
2017-12-01 23:36:05 +01:00 |
|
Søren Lind Kristiansen
|
d86b537a38
|
Enable morph rules for Danish
|
2017-11-30 15:58:02 +01:00 |
|
Søren Lind Kristiansen
|
13a988adc3
|
Remove 'Number[psor]'
|
2017-11-30 15:55:04 +01:00 |
|
Søren Lind Kristiansen
|
dd6fde18a9
|
Add more Danish morph rules and clean up existing ones
|
2017-11-30 11:17:19 +01:00 |
|
Vadim Mazaev
|
4ba7ddf651
|
Bugfixies
|
2017-11-30 12:29:38 +03:00 |
|
Matthew Honnibal
|
6bc0f4d29f
|
Merge pull request #1611 from fsonntag/master
Solving #1494
|
2017-11-29 23:11:23 +01:00 |
|
Matthew Honnibal
|
f9ed9ea529
|
Merge pull request #1624 from GreenRiverRUS/russian
Add support for Russian
|
2017-11-29 23:10:01 +01:00 |
|
Jim O'Regan
|
ba6a23fd11
|
BOM in Italian lemmatiser
|
2017-11-29 17:40:07 +00:00 |
|
ines
|
a31506e060
|
Fix off-by-one error in nlp.add_pipe(after=name) (fixes #1654)
|
2017-11-28 20:37:55 +01:00 |
|
ines
|
b62739fbfe
|
Add regression test for #1654
|
2017-11-28 20:27:54 +01:00 |
|
ines
|
2e50dbb9d7
|
Simplify test
|
2017-11-28 20:27:27 +01:00 |
|
Felix Sonntag
|
724ae7dc55
|
Fixed issue of infix capturing prefixes
|
2017-11-28 17:17:12 +01:00 |
|
Ines Montani
|
9052643e2c
|
Merge pull request #1653 from sorenlind/da_example_typo
Fix typo
|
2017-11-27 14:47:42 +00:00 |
|
Søren Lind Kristiansen
|
5fe58b885b
|
Fix typo
|
2017-11-27 15:36:18 +01:00 |
|
Ines Montani
|
d52b1ab245
|
Add unicode_literals (hopefully fixes test failure on Python 2)
|
2017-11-27 15:16:54 +01:00 |
|
Søren Lind Kristiansen
|
0ffd27b0f6
|
Add several Danish alternative spellings
|
2017-11-27 13:35:41 +01:00 |
|
Ines Montani
|
6362024cf8
|
Merge pull request #1645 from GreenRiverRUS/fix_default_meta
Fixed spaCy version string in default meta
|
2017-11-27 11:58:02 +00:00 |
|
Vadim Mazaev
|
59f03ab1d7
|
Fixed spacy version string in default meta
|
2017-11-26 23:02:07 +03:00 |
|
Vadim Mazaev
|
53e7c38637
|
Fixed tests depends on pymorphy2
|
2017-11-26 21:04:44 +03:00 |
|
Vadim Mazaev
|
cacd859dcd
|
Added tag map, fixed tests fails, added more exceptions
|
2017-11-26 20:54:48 +03:00 |
|
Ines Montani
|
a7bb8f1b42
|
Merge pull request #1637 from sorenlind/da_tokenization
Improve Danish tokenization
|
2017-11-26 15:41:38 +00:00 |
|
ines
|
c699aec089
|
Add offsets_from_biluo_tags helper and tests (see #1626)
|
2017-11-26 16:38:01 +01:00 |
|
Søren Lind Kristiansen
|
ef03e9ea53
|
Remove unused import.
|
2017-11-25 13:04:02 +01:00 |
|
Søren Lind Kristiansen
|
6aa241bcec
|
Add day of month tokenizer exceptions for Danish.
|
2017-11-24 15:03:24 +01:00 |
|
Søren Lind Kristiansen
|
0c276ed020
|
Add weekday abbreviations and remove abiguous month abbreviations for Danish.
|
2017-11-24 14:43:29 +01:00 |
|
Søren Lind Kristiansen
|
056547e989
|
Add multiple tokenizer exceptions for Danish.
|
2017-11-24 11:51:26 +01:00 |
|
Søren Lind Kristiansen
|
8dc265ac0c
|
Add test for tokenization of 'i.' for Danish.
|
2017-11-24 11:29:37 +01:00 |
|
Søren Lind Kristiansen
|
ac8116510d
|
Fix tokenization of 'i.' for Danish.
|
2017-11-24 11:16:53 +01:00 |
|
Matthew Honnibal
|
79f11d4f85
|
Pickle vectors with vocab
|
2017-11-23 17:19:50 +01:00 |
|
Matthew Honnibal
|
f29c3925ee
|
Fix more efficient nonproj
|
2017-11-23 12:48:00 +00:00 |
|
Matthew Honnibal
|
e10e9ad2c5
|
Improve efficiency of Doc.to_array
|
2017-11-23 12:33:27 +00:00 |
|
Matthew Honnibal
|
2acc907d55
|
Improve profiling
|
2017-11-23 12:33:03 +00:00 |
|
Matthew Honnibal
|
fa62427300
|
Remove lookup-based lemmatization
|
2017-11-23 12:32:22 +00:00 |
|
Matthew Honnibal
|
fb26b2cb12
|
Use lookup lemmatizer if lemma unset
|
2017-11-23 12:31:58 +00:00 |
|
Matthew Honnibal
|
db5c714ad2
|
Improve efficiency of deprojectivization
|
2017-11-23 12:31:34 +00:00 |
|
Matthew Honnibal
|
8fec7268eb
|
Move string cleanup under a setting flag
|
2017-11-23 12:19:18 +00:00 |
|
Matthew Honnibal
|
5949777b12
|
Fix misleading multi-threading docstring
|
2017-11-23 12:18:59 +00:00 |
|
Matthew Honnibal
|
542e6fd4ea
|
Don't remove entries from specials
|
2017-11-23 12:17:42 +00:00 |
|
Matthew Honnibal
|
30ba81f881
|
Merge pull request #1576 from ligser/master
Actually reset caches in pipe [wip]
|
2017-11-23 12:54:48 +01:00 |
|
ines
|
c90fe92e15
|
Fix displaCy test
|
2017-11-22 05:04:39 +01:00 |
|