Adriane Boyd
673e2bc4c0
Add usage docs for streamed train corpora ( #7693 )
2021-04-09 16:15:38 +02:00
Sofie Van Landeghem
3e5bd5055e
expand quickstart widget with cuda 11.1 and 11.2 ( #7615 )
2021-04-08 12:25:42 +02:00
Sofie Van Landeghem
204c2f116b
Extend score_spans for overlapping & non-labeled spans ( #7209 )
...
* extend span scorer with consider_label and allow_overlap
* unit test for spans y2x overlap
* add score_spans unit test
* docs for new fields in scorer.score_spans
* rename to include_label
* spell out if-else for clarity
* rename to 'labeled'
Co-authored-by: Adriane Boyd <adrianeboyd@gmail.com>
2021-04-08 12:19:17 +02:00
broaddeep
ee159b8543
Support match alignments ( #7321 )
...
* Support match alignments
* change naming from match_alignments to with_alignments, add conditional flow if with_alignments is given, validate with_alignments, add related test case
* remove added errors, utilize bint type, cleanup whitespace
* fix no new line in end of file
* Minor formatting
* Skip alignments processing if as_spans is set
* Add with_alignments to Matcher API docs
* Update website/docs/api/matcher.md
Co-authored-by: Sofie Van Landeghem <svlandeg@users.noreply.github.com>
Co-authored-by: Adriane Boyd <adrianeboyd@gmail.com>
Co-authored-by: Sofie Van Landeghem <svlandeg@users.noreply.github.com>
2021-04-08 18:10:14 +10:00
Ines Montani
de4f4c9b8a
Add more link anchors [ci skip]
2021-04-06 14:15:21 +10:00
Ines Montani
5bbdd7dc4c
Update pipeline design docs [ci skip]
2021-04-06 14:13:22 +10:00
Ines Montani
1d1cfadbca
Fix formatting [ci skip]
2021-04-06 14:13:13 +10:00
Jaidev Deshpande
93ee74a0a6
Add Numerizer to SpaCy universe ( #7650 )
...
Numerizer is a spaCy extension that converts numbers written in natural language
into numeric strings.
2021-04-05 19:02:27 +02:00
Sam Edwardes
f6ad4684bd
Updates to universe.json for spaCyTextBlob ( #7647 )
...
* Updates to universe.json for spaCyTextBlob
Updated the documentation for spaCy 3.0.
* SamEdwardes.md
* Update SamEdwardes.md
2021-04-04 20:17:57 +02:00
Ayush Chaurasia
3c2ce41dd8
W&B integration: Optional support for dataset and model checkpoint logging and versioning ( #7429 )
...
* Add optional artifacts logging
* Update docs
* Update spacy/training/loggers.py
Co-authored-by: Sofie Van Landeghem <svlandeg@users.noreply.github.com>
* Update spacy/training/loggers.py
Co-authored-by: Sofie Van Landeghem <svlandeg@users.noreply.github.com>
* Update spacy/training/loggers.py
Co-authored-by: Sofie Van Landeghem <svlandeg@users.noreply.github.com>
* Bump WandbLogger Version
* Add documentation of v1 to legacy docs
* bump spacy-legacy to 3.0.2 (to be released)
Co-authored-by: Sofie Van Landeghem <svlandeg@users.noreply.github.com>
Co-authored-by: svlandeg <sofie.vanlandeghem@gmail.com>
2021-04-01 19:36:23 +02:00
vincent d warmerdam
8b3eec6e62
Add Tokenwiser to Projects ( #7541 )
...
* Add tokenwiser
* Update universe.json
2021-04-01 14:39:36 +02:00
Sofie Van Landeghem
59c2069eb1
Legacy docs ( #7601 )
...
* document legacy Tok2Vec architectures
* add TextCatEnsemble.v1 legacy documentation
* Separate legacy section in side bar
2021-03-30 12:43:14 +02:00
Santiago Castro
af07fc3bc1
Add support for CUDA 11.2 ( #7583 )
...
* Add support for CUDA 11.2
* Update the docs
* Format
Co-authored-by: Adriane Boyd <adrianeboyd@gmail.com>
2021-03-30 09:47:33 +02:00
Álvaro Abella Bascarán
5b4dde38a3
fix fn name: tokenizer.infixes_finditer -> tokenizer.infix_finditer ( #7606 )
2021-03-30 09:45:49 +02:00
Ines Montani
be55f43163
Merge pull request #7473 from adrianeboyd/docs/v3-pipeline-deps-order
2021-03-22 12:43:07 +01:00
Ines Montani
3ee2fcfba0
Merge pull request #7483 from adrianeboyd/docs/various-v3-4 [ci skip]
2021-03-22 12:37:06 +01:00
Ines Montani
88e5a0dc16
Merge pull request #7504 from polm/fix/lexeme-docs [ci skip]
...
Fix mismatched backtick in Lexeme docs
2021-03-22 12:36:44 +01:00
Adriane Boyd
0d2b723e8d
Update entity setting section
2021-03-20 11:38:55 +01:00
Paul O'Leary McCann
e39c0dcf33
Fix mismatched backtick in Lexeme docs
2021-03-20 18:40:00 +09:00
Adriane Boyd
c771ec22f0
Update matcher errors and docs
...
* Mention `tagger+attribute_ruler` in `POS`/`MORPH` error messages for
`Matcher` and `PhraseMatcher`
* Document `Matcher.__call__(allow_missing=)`
2021-03-19 10:11:18 +01:00
Adriane Boyd
6a9a467766
Update website/docs/usage/processing-pipelines.md
...
Co-authored-by: Ines Montani <ines@ines.io>
2021-03-19 08:12:49 +01:00
Adriane Boyd
6354b642c5
Fix typo
2021-03-18 19:01:10 +01:00
Adriane Boyd
40e5d3a980
Update saving/loading example
2021-03-18 16:56:10 +01:00
Adriane Boyd
0fb1881f36
Reformat processing pipelines
2021-03-18 13:31:42 +01:00
Adriane Boyd
acc58719da
Update custom similarity hooks example
2021-03-18 13:31:42 +01:00
Adriane Boyd
c9e1a9ac17
Add multiprocessing section
2021-03-18 13:31:42 +01:00
Adriane Boyd
9a254d3995
Include all en_core_web_sm components in examples
2021-03-18 13:31:42 +01:00
Adriane Boyd
83c1b919a7
Fix positional/option in CLI types
2021-03-18 13:31:42 +01:00
Adriane Boyd
9fd41d6742
Remove Language.pipe cleanup arg
2021-03-18 13:31:42 +01:00
Adriane Boyd
5da323fd86
Minor edits
2021-03-17 12:59:05 +01:00
Adriane Boyd
a5ffe8dfed
Add details about pretrained pipeline design
2021-03-17 11:31:26 +01:00
Paolo Arduin
00e59be966
Add SpikeX to spaCy universe
2021-03-16 18:22:03 +01:00
bsweileh
61472e7cb3
Update _training.md - Fix broken link on backpropagation ( #7431 )
...
* Update _training.md
Fix broken link on backpropagation
* Add agreement
add spacy contributor agreement
2021-03-15 09:21:35 +01:00
Ines Montani
c67d5a6eb0
Merge pull request #7394 from adrianeboyd/docs/ner-example-data-readme
2021-03-13 04:26:18 +01:00
Ines Montani
068b97a617
Merge pull request #7408 from adrianeboyd/bugfix/load-keyword-only
2021-03-13 04:25:50 +01:00
Adriane Boyd
3168103605
Fix type of spacy train --output in docs
2021-03-12 10:04:57 +01:00
Adriane Boyd
03e9e7b567
Add --code option to init fill-config
2021-03-12 10:03:57 +01:00
Adriane Boyd
124304b146
Add vocab kwarg back to spacy.load
...
* Additional minor formatting and docs cleanup
2021-03-11 10:58:59 +01:00
Adriane Boyd
84470d9b9e
Incorporate BILUO note from #7407
2021-03-11 10:11:21 +01:00
Adriane Boyd
4294bcf4ab
Align keyword-only in docs for init/util
2021-03-11 09:52:40 +01:00
Adriane Boyd
28726c25a1
Update docs for convert CLI and NER examples
2021-03-10 11:42:02 +01:00
Adriane Boyd
d746ea6278
Add warning about GPU selection in Jupyter notebooks ( #7075 )
...
* Initial warning
* Update check
* Redo edit
* Move jupyter warning to helper method
* Add link with details to warnings
2021-03-09 15:35:21 +01:00
Sofie Van Landeghem
932887b950
textcat scoring fix and multi_label docs ( #6974 )
...
* add multi-label textcat to menu
* add infobox on textcat API
* add info to v3 migration guide
* small edits
* further fixes in doc strings
* add infobox to textcat architectures
* add textcat_multilabel to overview of built-in components
* spelling
* fix unrelated warn msg
* Add textcat_multilabel to quickstart [ci skip]
* remove separate documentation page for multilabel_textcategorizer
* small edits
* positive label clarification
* avoid duplicating information in self.cfg and fix textcat.score
* fix multilabel textcat too
* revert threshold to storage in cfg
* revert threshold stuff for multi-textcat
Co-authored-by: Ines Montani <ines@ines.io>
2021-03-09 23:04:22 +11:00
Sofie Van Landeghem
cd70c3cb79
Fixing pretrain ( #7342 )
...
* initialize NLP with train corpus
* add more pretraining tests
* more tests
* function to fetch tok2vec layer for pretraining
* clarify parameter name
* test different objectives
* formatting
* fix check for static vectors when using vectors objective
* clarify docs
* logger statement
* fix init_tok2vec and proc.initialize order
* test training after pretraining
* add init_config tests for pretraining
* pop pretraining block to avoid config validation errors
* custom errors
2021-03-09 14:01:13 +11:00
Ines Montani
dfb23a419e
Merge branch 'spacy.io' [ci skip]
2021-03-06 17:38:54 +11:00
graue70
7d085d5b1c
Fix typo in docs
2021-03-05 18:30:09 +01:00
vincent d warmerdam
1b0d413e45
Removed Languages that were listed twice on Docs ( #7272 )
...
* removed languages that were listed twice
* sorted
* d0h
* the d0h strikes back when you dont hit save
2021-03-05 14:31:15 +01:00
svlandeg
682a6232e3
fix typo
2021-03-02 17:59:13 +01:00
svlandeg
d900c55061
consistently use registry as callable
2021-03-02 17:56:28 +01:00
graue70
0fddc0447c
Fix copy & paste error in API docs
2021-03-02 14:00:14 +01:00