Numeric Lyndon-based feature embedding of sequencing reads for machine learning approaches
From MaRDI portal
Abstract: Feature embedding methods have been proposed in literature to represent sequences as numeric vectors to be used in some bioinformatics investigations, such as family classification and protein structure prediction. Recent theoretical results showed that the well-known Lyndon factorization preserves common factors in overlapping strings. Surprisingly, the fingerprint of a sequencing read, which is the sequence of lengths of consecutive factors in variants of the Lyndon factorization of the read, is effective in preserving sequence similarities, suggesting it as basis for the definition of novels representations of sequencing reads. We propose a novel feature embedding method for Next-Generation Sequencing (NGS) data using the notion of fingerprint. We provide a theoretical and experimental framework to estimate the behaviour of fingerprints and of the -mers extracted from it, called -fingers, as possible feature embeddings for sequencing reads. As a case study to assess the effectiveness of such embeddings, we use fingerprints to represent RNA-Seq reads and to assign them to the most likely gene from which they were originated as fragments of transcripts of the gene. We provide an implementation of the proposed method in the tool lyn2vec, which produces Lyndon-based feature embeddings of sequencing reads.
Recommendations
- Can we replace reads by numeric signatures? Lyndon fingerprints as representations of sequencing reads for machine learning
- Low-dimensional representation of genomic sequences
- Sequence graph transform (SGT): a feature embedding function for sequence data mining
- Analysis method and algorithm design of biological sequence problem based on generalized k-mer vector
- Estimating sequence similarity from read sets for clustering next-generation sequencing data
Cites work
- A new characterization of maximal repetitions by Lyndon trees
- Classification of time series by shapelet transformation
- Factorizing words over an ordered alphabet
- Free differential calculus. IV: The quotient groups of the lower central series
- scientific article; zbMATH DE number 1737190 (Why is no real title available?)
- Inverse Lyndon words and inverse Lyndon factorizations of words
- Lyndon words versus inverse Lyndon words: queries on suffixes and bordered words
- On Burnside's Problem
- On the longest common prefix of suffixes in an inverse Lyndon factorization and other properties
- SPADE: An efficient algorithm for mining frequent sequences
- The origins of combinatorics on words
- Variable length local decoding and alignment-free sequence comparison
This page was built for publication: Numeric Lyndon-based feature embedding of sequencing reads for machine learning approaches
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6195191)