TDbasedUFEadv
Workflows
Standard Workflow
Integrate two omics datasets sharing the same features using memory-efficient SVD with partial summation.
library(TDbasedUFEadv)
library(RTCGA.rnaseq)
# Prepare drug and disease expression datasets
Cancer_cell_lines <- list(ACC.rnaseq, BLCA.rnaseq, BRCA.rnaseq, CESC.rnaseq)
Drug_and_Disease <- prepareexpDrugandDisease(Cancer_cell_lines)
expDrug <- Drug_and_Disease$expDrug
expDisease <- Drug_and_Disease$expDisease
# Compute SVD on the two matrices
SVD <- computeSVD(exprs(expDrug), exprs(expDisease))
# Generate the matrix product and prepare the tensor
Z <- t(exprs(expDrug)) %*% exprs(expDisease)
sample <- outer(
colnames(expDrug),
colnames(expDisease),
function(x, y) { paste(x, y) }
)
Z <- PrepareSummarizedExperimentTensor(
sample = sample,
feature = rownames(expDrug),
value = Z
)
Inputs are two expression matrices sharing features; output is a SummarizedExperiment-like tensor and SVD results.
Multi Omics Sharing Features
Integrate multiple omics datasets sharing the same features but having different samples using projection-based SVD and HOSVD.
library(TDbasedUFEadv)
library(RTCGA.rnaseq)
library(RTCGA.clinical)
# Prepare a list of matrices sharing features
Multi <- list(
BLCA.rnaseq[seq_len(100), 1 + seq_len(1000)],
BRCA.rnaseq[seq_len(100), 1 + seq_len(1000)],
CESC.rnaseq[seq_len(100), 1 + seq_len(1000)],
COAD.rnaseq[seq_len(100), 1 + seq_len(1000)]
)
# Prepare a tensor from the list of matrices
Z <- prepareTensorfromList(Multi, 10L)
# Permute the tensor modes to align features
Z <- aperm(Z, c(2, 1, 3))
# Wrap the tensor using PrepareSummarizedExperimentTensor
Clinical <- list(BLCA.clinical, BRCA.clinical, CESC.clinical, COAD.clinical)
Multi_sample <- list(
BLCA.rnaseq[seq_len(100), 1, drop = FALSE],
BRCA.rnaseq[seq_len(100), 1, drop = FALSE],
CESC.rnaseq[seq_len(100), 1, drop = FALSE],
COAD.rnaseq[seq_len(100), 1, drop = FALSE]
)
ID_column_of_Multi_sample <- c(770, 1482, 773, 791)
ID_column_of_Clinical <- c(20, 20, 12, 14)
Z <- PrepareSummarizedExperimentTensor(
feature = colnames(ACC.rnaseq)[1 + seq_len(1000)],
sample = array("", 1),
value = Z,
sampleData = prepareCondTCGA(
Multi_sample,
Clinical,
ID_column_of_Multi_sample,
ID_column_of_Clinical
)
)
# Compute HOSVD
HOSVD <- computeHosvd(Z)
# Select features using selectFeatureProj in batch mode
cond <- attr(Z, "sampleData")
index <- selectFeatureProj(HOSVD, Multi, cond, de = 1e-3, input_all = 3)
# Extract and display selected features
head(tableFeatures(Z, index))
Inputs are a list of matrices sharing features and clinical metadata; output is a table of selected features with p-values.
Multi Omics Sharing Samples
Integrate multiple omics datasets sharing the same samples using projection-based SVD and HOSVD.
library(TDbasedUFEadv)
# Prepare a tensor from the list of matrices and wrap in a SummarizedExperiment-like object
# Z <- PrepareSummarizedExperimentTensor(sample = sample, feature = feature, value = value)
# Compute HOSVD
# HOSVD <- computeHosvd(Z)
# Select features across individual profiles
# index <- selectFeature(HOSVD, input_all, de = 0.05)
# Extract, print, and save selected features
# head(tableFeatures(Z, index))
Inputs are multi-omics matrices sharing samples; output is a table of selected features.
When to Use
- Unsupervised feature extraction (e.g., gene selection) from multi-omics datasets where either features (genes) or samples are shared.
- When dealing with a small number of samples associated with a large number of features (common in genomics).
- Evaluating selected genes via enrichment analysis using tools like
enrichR::enrichr,STRINGdb::STRINGdb, orDOSE::enrichDGN.
When NOT to Use
- For supervised differential expression analysis where group labels are strictly used to guide the mathematical decomposition; use
DESeq2orlimmainstead. - When you prefer a fully automated, non-interactive pipeline without manual/interactive selection of singular value vectors; use standard
SVDorHOSVDpackages directly.
Data Requirements
- Input data should be matrix-like (e.g., gene expression matrices from
RTCGA.rnaseqlikeBLCA.rnaseq). - For feature-sharing multi-omics, a list of matrices (e.g.,
Multi) where rows represent features and columns represent samples. - Metadata/clinical data (e.g.,
BLCA.clinical) to construct condition vectors for selecting singular value vectors.
Key Parameters
- de (1e-3): Standard deviation threshold for feature selection in
selectFeatureProjorselectFeature. - input_all: Vector of selected singular value vectors in
selectSingularValueVectorLargeorselectFeatureProj.
Best Practices
- Master the basic
TDbasedUFEpackage workflows before moving to the advanced features ofTDbasedUFEadv. - Use memory-efficient SVD with partial summation (
computeSVD) when the number of features is too large to construct a full tensor. - Verify the success of singular value vector selection interactively using histograms and standard deviation optimization via
selectSingularValueVectorLarge. - Perform downstream enrichment analysis on selected genes using
enrichR::enrichrorSTRINGdb::STRINGdbto biologically validate the unsupervised feature selection.
Common Pitfalls
- Out of memory errors: Attempting to apply
computeHosvdon a full tensor with too many features. Fix: UsecomputeSVDwith partial summation to reduce memory requirements. - Incorrect tensor permutation: Forgetting to align features across datasets. Fix: Use
apermto permute the tensor modes appropriately before wrapping withPrepareSummarizedExperimentTensor.
Alternatives
DESeq2: For supervised differential expression analysis.limma: For linear modeling of gene expression data.TDbasedUFE: For simpler, standard unsupervised feature extraction workflows.
Citations
- Taguchi, Y-H. 2020. Unsupervised Feature Extraction Applied to Bioinformatics. Springer International Publishing. https://doi.org/10.1007/978-3-030-22456-1
- Taguchi, Y-H. 2023. TDbasedUFE: Tensor Decomposition Based Unsupervised Feature Extraction.
References
- Homepage: bioconductor.org/packages/TDbasedUFEadv
- Vignette: https://bioconductor.org/packages/release/bioc/vignettes/TDbasedUFEadv/inst/doc/Enrichment.html