sentencepiece (Q89656)

From MaRDI portal

!

This is the item page for this Wikibase entity, intended for internal use and editing purposes. Please use the normal view instead:

Text Tokenization using Byte Pair Encoding and Unigram Modelling
Language Label Description Also known as
default for all languages
No label defined
    English
    sentencepiece
    Text Tokenization using Byte Pair Encoding and Unigram Modelling

      Statements

      0 references
      0.2.3
      13 November 2022
      0 references
      PACKAGES.rds
      9 July 2026
      0 references
      0.1.1
      4 June 2020
      0 references
      0.1.2
      8 June 2020
      0 references
      0.2.1
      21 December 2021
      0 references
      0.2.2
      9 November 2022
      0 references
      0.2.4
      27 November 2025
      0 references
      0.2
      15 December 2021
      0 references
      0.2.5
      9 February 2026
      0 references
      0 references
      0 references
      0 references
      0 references
      9 February 2026
      0 references
      Unsupervised text tokenizer allowing to perform byte pair encoding and unigram modelling. Wraps the 'sentencepiece' library <https://github.com/google/sentencepiece> which provides a language independent tokenizer to split text in words and smaller subword units. The techniques are explained in the paper "SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing" by Taku Kudo and John Richardson (2018) <doi:10.18653/v1/D18-2012>. Provides as well straightforward access to pretrained byte pair encoding models and subword embeddings trained on Wikipedia using 'word2vec', as described in "BPEmb: Tokenization-free Pre-trained Subword Embeddings in 275 Languages" by Benjamin Heinzerling and Michael Strube (2018) <http://www.lrec-conf.org/proceedings/lrec2018/pdf/1049.pdf>.
      0 references

      Identifiers