×

Nextflow Modules

Clear

Showing module(s) with keyword "clustering"

Module Keywords Description
nf-core/ampcombi2/cluster antimicrobial peptides amps parsing reporting align clustering mmseqs2 A submodule that clusters the merged AMP hits generated from ampcombi2/parsetables and ampcombi2/complete using MMseqs2 cluster.
nf-core/autocycler/cluster autocycler genome-assembly clustering long-read Cluster replicons in compressed assemblies with Autocycler.
nf-core/bigscape/bigscape biosynthetic gene clusters BGC clustering similarity network antiSMASH gene cluster families secondary metabolites natural products BiG-SCAPE (Biosynthetic Gene Similarity Clustering and Prospecting Engine) clusters biosynthetic gene clusters (BGCs) into gene cluster families (GCFs) based on sequence similarity networks. It takes antiSMASH-annotated GenBank files and a Pfam database as input, and produces similarity networks and clustering results.
nf-core/comebin/runcomebin metagenomics binning clustering Effective binning of metagenomic contigs using COntrastive Multi-viEw representation learning
nf-core/custom/clustermetrics clustering metrics silhouette calinski-harabasz davies-bouldin evaluation Computes clustering quality metrics (silhouette, Calinski-Harabasz, Davies-Bouldin) and performs k-sweep analysis
nf-core/custom/clustervisualization clustering visualization pca umap tsne dimension-reduction Generates UMAP and t-SNE visualizations colored by cluster
nf-core/custom/orfmerge orf ribo-seq catalogue merge clustering Cluster normalised per-sample, per-caller ORF predictions into a single cohort-level catalogue. Pair with `custom/orfnormalise` upstream and (typically) `bedtools/getfasta` + `seqkit/translate` downstream to obtain the AA FASTA. Rows sharing an identical exon structure are grouped first and clustered as a single proxy chosen by class specificity, then the members are restored. Class selects the strategy, so without this an ORF that two callers class differently would be split across strategies and could never merge. Proxies are then partitioned by clustering strategy. The harmonised `orf_class` written by `custom/orfnormalise` selects the strategy but is never part of a grouping key, because callers disagree on class for the same ORF (Ribo-TISH reports `5'UTR` for both uORFs and CDS-overlapping uORFs) and keying on it would emit one catalogue row per disagreeing caller: - canonical_cds: grouped by (transcript_id, strand), then reciprocal-overlap clustered within the transcript so a short truncated variant is not folded into the full-length CDS. - uORF, uoORF, dORF, collapse by (transcript_id, strand, start, doORF, intORF, other: end). A single transcript can host multiple distinct uORFs / dORFs / internal ORFs, so keying on the outer span keeps them in separate clusters while still merging cross-caller calls that agree on coordinates. - novel_u: greedy reciprocal-overlap clustering on summed exon-block intersection at `--reciprocal-overlap` (default 0.8). Catches fuzzy cross-caller matches and exact-coordinate collapses in one pass. Order-dependent at the boundary: a chain A-B-C where A-B and B-C overlap at ~0.85 but A-C only at ~0.75 may cluster as {A,B,C} or {A,B}+{C} depending on iteration order. Rare in practice at 0.8. Cross-caller consensus is recorded in two column families on the catalogue TSV: - `called_by_<caller>`: 0/1 indicator per supported caller (ribotish, ribocode, ribotricer, rpbp, price). - `score_<caller>`: best score from that caller within the cluster. Score direction is per-caller (p-values are minimised; Bayes factors / phase scores are maximised). Cross-sample recurrence is recorded in two further columns: - `n_samples`: number of distinct samples contributing to the cluster (a cohort recurrence metric). - `samples`: sorted, comma-separated list of those sample ids. Emits a small MultiQC custom-content TSV (per-class counts) for inclusion in downstream MultiQC reports. Alongside the full catalogue, emits a consensus view (`*.consensus.*`) filtered to ORFs supported by at least `--min-callers` distinct callers and recurring in at least `--min-samples` samples (both default 1, i.e. no filtering, so the consensus view equals the full catalogue). Raising either threshold yields a higher-confidence catalogue without altering the full one.
nf-core/custom/pcaclustering clustering kmeans dbscan pca embeddings Performs KMeans or DBSCAN clustering on a sample-by-feature numeric matrix (e.g. principal components, embeddings)
nf-core/diamond/cluster clustering alignment genomics proteomics calculate clusters of highly similar sequences
nf-core/diamond/deepclust clustering protein diamond deepclust proteomics Fast graph-based protein sequence clustering using DIAMOND deepclust
nf-core/diamond/linclust clustering protein diamond linclust proteomics Fast protein sequence clustering using a greedy incremental approach
nf-core/humid umi fastq deduplication hamming-distance clustering HUMID is a tool to quickly and easily remove duplicate reads from FASTQ files, with or without UMIs.
nf-core/leafcutter/clusterregtools splicing RNA-seq clustering junctions Cluster RNA-seq junction reads extracted by regtools and refine them based on read counts and ratios for alternative splicing analysis
nf-core/lsa/cosine similarity cosine clustering rnaseq heatmap Calculates the cosine similarity matrix between samples based on a gene expression matrix.
nf-core/mgnifam/generatefamilies protein families