Pipeline and methods#
BASALT is organized as a sequence of candidate generation, selection, refinement, and reassembly stages. The description below distinguishes executable operations from biological interpretation.
assemblies + reads
→ candidate generation
→ within-assembly selection
→ cross-assembly dereplication
→ contig screening and retrieval
→ reassembly and OLC
→ final MAG set + quality report
Stage 0: preprocessing and coverage#
BASALT filters assembly sequences shorter than 1,000 bp, maps the supplied reads, and derives depth and connectivity files. Paired-end reads are mapped with Bowtie 2. Long reads use long-read-aware mapping paths, including minimap2 where applicable.
Interpretive boundary. Coverage is informative only when the supplied reads and assemblies belong to a coherent study design. BASALT does not test whether files were paired correctly at the biological level.
Stage 1: candidate generation#
The selected --sensitive preset controls candidate diversity:
Preset |
Candidate generators |
|---|---|
|
MetaBAT 2 and SemiBin 2 |
|
MetaBAT 2, CONCOCT, and SemiBin 2 |
|
MaxBin 2.0, MetaBAT 2, CONCOCT, and SemiBin 2 |
MetaBAT 2 is run across four internal threshold settings. CONCOCT cluster counts are derived from preliminary bin counts. MaxBin 2.0 uses four probability thresholds in more-sensitive mode. Optional adapters can add MetaBinner, VAMB, or LorBin candidates.
Interpretive boundary. Candidate multiplicity increases the space evaluated by later selection stages. It does not make every candidate independent or biologically valid.
Stages 2–4: bin characterization and selection#
BASALT computes abundance and paired-end connectivity summaries, compares candidates within an assembly, and then compares selected bins across assemblies. The selection logic uses combinations of contig overlap, coverage, sequence composition, and CheckM or CheckM2 quality estimates.
For biologically related samples, contigs from the same population genome can exhibit concordant coverage changes across samples, whereas foreign or mismatched contigs may show discordant profiles. This differential signal is most informative when the sample group retains substantial community continuity while population abundances vary. It is complementary evidence, not an independent proof of contig origin; see Study design patterns.
These stages reduce redundant candidates and retain representatives for downstream refinement.
Interpretive boundary. Computational dereplication is not a taxonomic or strain-definition procedure. Use an explicit downstream species or strain criterion if the scientific question requires one.
Stage 5: contig-level contamination screening#
The outlier-removal stage derives tetranucleotide-frequency and coverage-related features for contigs. An ensemble of five multilayer perceptrons classifies contigs into retained and predicted-contaminant classes. The ensemble combines model predictions by voting.
The model architecture contains linear, batch-normalization, and rectified-linear layers with residual blocks. Model weights are distributed separately and are located through BASALT_WEIGHT.
Interpretive boundary. The output is a model prediction conditioned on the training distribution and available features. It is not direct experimental proof that a contig is foreign to a population genome. Report the model version or checksum when refinement affects the main result.
Stages 6–8: contig retrieval and OLC#
BASALT evaluates unbinned or excluded contigs using paired-end connectivity, long-read connectivity, coverage consistency, and sequence-composition criteria. Compatible contigs may be added to a bin. Overlap–layout–consensus (OLC) operations can merge or extend sequences in applicable routes.
--refinepara deep extends retrieval in compatible code paths. The effect depends on the available read types and checkpoint state.
Interpretive boundary. Retrieved contigs should be treated as read- and feature-supported candidates. Evaluate quality changes, taxonomic consistency, and unexpected sequence composition before interpreting an expanded bin.
Stages 9–10: reassembly and final OLC#
When paired-end short reads are available, BASALT performs bin-level reassembly using SPAdes-related paths. Hybrid routes use long reads where supported. Reassembled and original candidates are compared, followed by another OLC step in compatible workflows.
Reassembly is skipped when the required short-read inputs are absent. Individual bins may also lack sufficient reads for successful reassembly.
Interpretive boundary. Higher contiguity is not sufficient evidence of higher genome accuracy. Examine completeness, contamination, assembly size, N50, read support, and taxonomic coherence together.
Finalization#
The final selected directory is moved to the name supplied by -o. FASTA files are normalized to .fa names where possible, and a final CheckM2 report is generated when one is not already present. Cleanup then archives or removes several intermediate groups.
Retain logs and checksums before applying additional filtering or renaming.
Checkpoint model#
Basalt_checkpoint.txt records textual completion markers for major stages. --mode continue reads these markers and attempts to resume from the last recognized state.
Checkpointing provides operational recovery, not environment reproducibility. A valid resume requires compatible inputs, code, models, databases, software versions, and intermediate filenames.
Implementation map#
Stage |
Main implementation |
|---|---|
CLI and dispatch |
|
CheckM2 orchestration |
|
Autobinning |
|
Optional binners |
|
Abundance and PE connectivity |
|
Within-assembly comparison |
|
Cross-assembly comparison |
|
Outlier screening |
|
PE-supported retrieval |
|
Additional retrieval and polishing |
|
OLC |
|
Reassembly |
|
Final OLC |
|
The dated filenames are implementation identifiers, not software release numbers. Cite the BASALT release or commit used for analysis.
Methodological reporting#
A Methods section should state:
BASALT release or Git commit;
input assemblies and read datasets, including preprocessing;
selected quality-control backend and database;
sensitivity and refinement settings;
completeness and contamination thresholds;
optional binners and their versions;
compute environment and external-tool versions;
post-BASALT filtering, dereplication, taxonomy, and quality criteria.
See Reproducibility and reporting for a reusable template.