Skip to content

Known Issues — 2026 Modernization Audit

Running log of every defect found while redeploying ViroProfiler dev_ru on a fresh machine (NVIDIA DGX Spark, aarch64/GB10, Apptainer 1.5.2, Nextflow 26.04.3, 2026-08-12).

Bugs are recorded first and fixed in batches; the Status column tracks that. Severity: P0 blocks any run · P1 blocks a common use case · P2 correctness or usability defect · P3 hygiene.

ID Severity Area Status
I-01 P0 Containers — amd64-only images Partly fixed — arm64 images build from docker/, except iPHoP. DeepVirFinder was retired in favour of geNomad, which builds on aarch64
I-02 P0 Containers — viroprofiler-viewer has no Dockerfile Fixed
I-03 P0 Config — params.db never bind-mounted into containers Fixed
I-04 P1 Test data — stub samplesheet uses launch-dir-relative paths Fixed
I-05 P1 Config — params.tracedir frozen to the default outdir Fixed
I-06 P1 Databases — iPHoP DB directory name hardcoded and stale Fixed
I-07 P1 Databases — Bracken/Kraken2 DB paths disagree Fixed
I-08 P2 Databases — NCBI taxonomy pinned to a 2022 archive snapshot Fixed — its only consumer, the MMseqs2 LCA, was retired and DB_VREFSEQ deleted
I-09 P2 Databases — VOGDB host renamed; plain-HTTP URL Fixed — HTTPS, canonical host, release pinned via --vogdb_version
I-10 P2 Databases — setup steps are not resumable and never verified Open
I-11 P2 Workflow — --mode fastqc / fastp / contiglib not honoured Fixed
I-12 P2 Config — docker.userEmulation removed in modern Nextflow Fixed
I-13 P3 Repo — stub output directories committed despite .gitignore Fixed
I-14 P3 Docs — CLAUDE.md references an MCP server that is not part of the repo Fixed
I-15 P1 Config — contamref_idx ignores --db and nothing ever creates it Partly fixed — follows --db, still not built by setup
I-16 P1 Config — modules.config loaded after profiles, so containers were unoverridable Fixed
I-17 P2 Containers — Dockerfiles call wget that is only present transitively Fixed
I-18 P2 Containers — DeepVirFinder bundled into the binning image Fixed
I-19 P2 Modules — vendored nf-core modules and their containers are from 2022 Open
I-20 P3 Assets — samplesheet_contigs.csv contains a literal ${HOME} Fixed
I-21 P1 Containers — 2022-era tools break on modern Python/setuptools Fixed
I-22 P0 Containers — DRAM 1.4 source file copied over a DRAM 1.3.5 install Fixed
I-23 P1 Databases — DRAM CONFIG hand-written from 2021 filenames and today's date Fixed
I-24 P0 Modules — DRAMV writes a symlink into the container image Fixed
I-25 P1 Containers — eggNOG-mapper 2.1.9 needs distutils, gone in Python 3.12 Fixed
I-26 P0 Databases — VIBRANT setup reports success with an unusable database Fixed
I-27 P1 Containers — iPHoP built from bioconda carries three known defects Fixed
I-28 P0 Databases — DRAM's dbCAN downloads return an HTML landing page Fixed
I-29 P1 Databases — DRAM setup cannot tell a finished database from an abandoned one Fixed
I-30 P0 Databases — VOGDB moved its profiles into a subdirectory; DRAM builds an empty HMM file Fixed
I-31 P1 Databases — files unpacked by tar can land unreadable by their own owner Fixed
I-32 P1 Modules — vConTACT2 taxonomy derived from a SPAdes naming convention Fixed — replaced by vConTACT3
I-33 P0 Containers — RESULTS_TSE cannot read its own gzipped abundance inputs Fixed
I-34 P2 Modules — -max_target_seqs makes contig dereplication depend on library size Fixed — BLAST chain replaced by Vclust
I-35 P3 Modules — exact duplicate contigs entered the O(n²) dereplication stage Fixed
I-36 P1 Modules — PHAMB_RF calls a CLI the installed phamb does not have Fixed
I-37 P0 Containers — run_RF.py copied onto PATH cannot find its own model Fixed
I-38 P0 Containers — the binning image never installed VAMB Fixed on amd64; VAMB has no aarch64 build
I-39 P1 Containers — the binning image's second vrhyme prefix had drifted to Python 3.14 Fixed
I-40 P2 Modules — VAMB was given a depth table covering contigs its FASTA does not contain Fixed
I-41 P3 Config — CONTIGLIB_CLUSTER requested one CPU while Vclust threads every stage Fixed
I-42 P0 Containers — conda post-link scripts are not run by pixi, leaving a data package uninstalled Fixed
I-43 P0 Containers — the VAMB version pinned has no --jgi, the flag the process passes Fixed
I-44 P1 Modules — VAMB's --jgi pairs depths to contigs by row, not by name Fixed
I-45 P0 Config — the pipeline does not parse under Nextflow's default config parser Fixed
I-46 P1 Modules — iPHoP, CheckAMG and DRAM-v results never reached RESULTS_TSE Fixed
I-47 P2 Config — 432 lines of iGenomes reference config that nothing reads Fixed
I-48 P1 Config — a scheduler's TMPDIR points outside the container, and DRAM-v dies on it Fixed
I-49 P2 Workflow — the completion summary was printed twice on every run Fixed
I-50 P0 Config — a boolean given on the command line arrived as a string, and "false" is true Fixed
I-51 P1 Config — process.resourceLimits above the profiles block ignores every profile Fixed
I-52 P0 Databases — DB_VCONTACT3 cannot download: curl is not in the image Fixed
I-53 P2 Databases — a symlinked database that is not bind-mounted is silently re-downloaded Open
I-54 P1 CI — docker.yml was invalid YAML, so the lockfile gate never ran once Fixed
I-55 P2 Modules — two of ABUNDANCE's six CoverM passes feed no downstream process Closed — kept as published side products, by decision
I-56 P0 Databases — neither PHAMB database could be built, and the guard hid both Fixed
I-57 P1 Config — Docker creates a missing --db as root, so setup cannot write to it Fixed
I-58 P0 Containers — the viewer image could not install vpfkit on amd64: here undeclared, present on aarch64 only transitively Fixed
I-59 P1 Containers — viroprofiler-host could not be built by CI: the iPHoP fork it installs was private Fixed — the fork is public, the image is published

I-01 — All 11 container images are amd64-only (P0)

docker manifest inspect on every image referenced by conf/modules.config returns a single-architecture amd64 manifest:

denglab/viroprofiler-base:v0.2           amd64
denglab/viroprofiler-abundance:v0.2      amd64
denglab/viroprofiler-bracken:v0.2        amd64
denglab/viroprofiler-vibrant:v0.2        amd64
denglab/viroprofiler-binning:v0.2        amd64
denglab/viroprofiler-geneannot:v0.2      amd64
denglab/viroprofiler-host:v0.2.6         amd64
denglab/viroprofiler-replicyc:v0.1       amd64
denglab/viroprofiler-taxa:v0.1           amd64
denglab/viroprofiler-virsorter2:v0.2.5   amd64
denglab/viroprofiler-viewer:v0.2         amd64

The four vendored nf-core modules (FASTQC, FASTP, SPADES, MULTIQC) pull quay.io/biocontainers/*, which are likewise amd64-only.

Impact. The pipeline cannot start on any arm64 host (Apple silicon, AWS Graviton, NVIDIA GB10/GB200, Ampere). Apptainer fails with Exec format error unless a binfmt_misc QEMU interpreter is registered, and emulation is not a supportable answer for multi-threaded native aligners.

Decision. Publish arm64 builds from the Dockerfiles already in docker/.

Known obstacles for the arm64 rebuild (from api.anaconda.org, 2026-08-12):

Package linux-aarch64 in bioconda?
mmseqs2, spades, fastp, bbmap, seqkit, prodigal-gv, bowtie2, coverm yes
virsorter no — linux-64 + noarch only
fastqc, abricate, eggnog-mapper, multiqc no — linux-64 + noarch only
checkv, vibrant, iphop, bacphlip, replidec, vrhyme noarch (installable, but their compiled dependencies must resolve)
vcontact3 no — and two of its dependencies have no aarch64 conda build either (I-32)
deepvirfinder not in bioconda at all (comes from the hcc channel)

The retired docker/viroprofiler-taxa/Dockerfile additionally downloaded mmseqs-linux-avx2.tar.gz, an x86-64 binary that has no arm64 equivalent under that name.

I-02 — viroprofiler-viewer image has no Dockerfile in this repo (P0 for rebuilds)

conf/modules.config maps withLabel: viroprofiler_vpfkit to denglab/viroprofiler-viewer:v0.2, which RESULTS_TSE (modules/local/base.nf) uses to build the final TreeSummarizedExperiment. There is no docker/viroprofiler-viewer/ directory, so the image cannot be rebuilt or audited from this repository. The label name (vpfkit) and the image name (viewer) also disagree, which makes the mapping hard to follow.

I-03 — params.db is used inside containers but never bind-mounted (P0)

Thirteen call sites reference the database root as a bare path inside a container script, for example:

  • modules/local/viral_detection.nf:22,26,81,160
  • modules/local/annotation.nf:21,22,62,85,107
  • modules/local/taxonomy.nf:61,65,66
  • modules/local/viral_host.nf:18
  • modules/local/bracken.nf:27

params.db is a plain string, not a staged Nextflow input, so singularity.autoMounts / apptainer.autoMounts do not mount it. Nothing in nextflow.config adds a bind either — the only bind directives in the repository live in the untracked, site-specific custom.config.

Impact. The default --db $HOME/viroprofiler works only by accident, because Apptainer mounts $HOME by default. Any user who follows the documentation and puts the databases on a scratch or project filesystem (--db /mnt/scratch/db/viroprofiler, --db /scratch/$USER/db, …) gets "No such file or directory" from inside every container. This is a strong candidate for the failure users have been reporting.

I-04 — Stub samplesheet paths resolve against the launch directory (P1)

tests/data/samplesheet_stub.csv and samplesheet_stub_se.csv list tests/data/stub_R1.fastq.gz. Nextflow resolves a relative file() path against launchDir, not projectDir, so the stub test only works when it is launched from the repository root:

ERROR ~ No such file or directory: <launchDir>/tests/data/stub_R1.fastq.gz
 -- Check script 'subworkflows/local/input_check.nf' at line: 14

Introduced by commit b7689d1 ("use relative paths in stub samplesheet for local compatibility"). The CI workflow happens to run from the repo root, so CI never caught it.

I-05 — params.tracedir captures the default outdir (P1)

nextflow.config:88 defines tracedir = "${params.outdir}/pipeline_info" inside the same params block that defines outdir. The GString is interpolated when the block is parsed, before any profile or --outdir on the command line is applied. Running with -profile test_stub (which sets outdir = "output_stub") still reports tracedir : output/pipeline_info, and the timeline/report/trace/DAG files are written to the wrong tree — or fail to render at all:

WARN: Failed to render execution report -- see the log file for details
WARN: Failed to render execution timeline -- see the log file for details

I-06 — iPHoP database directory name is hardcoded and inconsistent (P1)

  • modules/local/viral_host.nf:18 passes --db_dir ${params.db}/iphop/Aug_2023_pub_rw.
  • modules/local/setup_db.nf (DB_IPHOP) runs iphop download -d $params.db/iphop -n and then deletes ${params.db}/iphop/iPHoP_db_Sept21.tar.gz.

The setup step therefore cleans up an artefact from the September 2021 release while the analysis step expects the August 2023 directory. Whichever release iphop download currently fetches, at most one of the two is right, and the layout is not verified after download. iPHoP host prediction is on by default (use_iphop = true), so this breaks the default configuration.

I-07 — Bracken and Kraken2 database paths disagree (P1)

DB_KRAKEN2 builds into ${params.db}/kraken2/..., workflows/viroprofiler.nf:255 passes "${params.db}/kraken2" to BRACKEN, but modules/local/bracken.nf:27 links ${params.db}/bracken/taxonomy/*. The bracken subdirectory is never created by any setup process. Only reachable with --use_kraken2 true, hence P1 rather than P0.

I-08 — NCBI taxonomy pinned to the 2022-08-01 archive (P2)

DB_VREFSEQ downloaded https://ftp.ncbi.nih.gov/pub/taxonomy/taxdump_archive/taxdmp_2022-08-01.zip. The URL still resolves (HTTP 200 as of 2026-08-12), so this was not a dead link, but every ViroProfiler installation was silently frozen to a four-year-old taxonomy. Taxa described since 2022 — including the entire post-2022 ICTV phage reclassification — could not be assigned.

Resolution. The snapshot existed for one consumer: mmseqs createtaxdb, which built the LCA database TAXONOMY_MMSEQS searched. That module was the lowest-priority taxonomy source of four, behind VITAP, geNomad and vConTACT3, and it has been retired along with DB_VREFSEQ, bin/parse_mmseqsTaxa.py and params.taxa_db_source. VITAP and vConTACT3 carry their own reference sets, so nothing in the pipeline reads an NCBI taxdump any more.

I-09 — VOGDB host renamed; URL is plain HTTP (P2)

DB_VOGDB fetched http://fileshare.csb.univie.ac.at/vog/latest/vog.hmm.tar.gz. Measured 2026-08-15, that URL still resolves, through two redirects: plain HTTP to HTTPS, then csb.univie.ac.at to fileshare.lisc.univie.ac.at. It worked only because wget follows them. latest was also unpinned, so two installations built months apart got different databases with nothing recording which.

Both are fixed: the process now requests https://fileshare.lisc.univie.ac.at/vog/vog${params.vogdb_version}/vog.hmm.tar.gz directly, with vogdb_version defaulting to 236 — the newest release at the time of writing, 554 MB compressed.

The same download was also broken in a second, worse way; see I-56.

The other external URLs were re-verified on 2026-08-12 and all return HTTP 200: Zenodo record 7044674 (mmseqs_vrefseq.tar.gz), the phamb RF_model.sav on raw.githubusercontent.com, and the miComplete Bact105.hmm on Bitbucket.

I-10 — Database setup is not resumable and never verified (P2)

Every process in modules/local/setup_db.nf follows the same pattern:

if [ ! -d ${params.db}/<tool> ]; then <download> ; else echo "already exists" ; fi

Consequences:

  1. Interrupted downloads are treated as success. The guard tests only for directory existence. A download killed halfway leaves the directory in place, so every later run skips it and the pipeline fails much later with a confusing tool-level error.
  2. No checksum or content validation. Nothing confirms that the expected files landed.
  3. No declared outputs. The processes have neither input: nor output: blocks and write outside the task work directory, so -resume cannot reason about them and Nextflow's provenance tracking does not cover the databases.
  4. DB_VIROPROFILER is a stub that prints "Please download checkv database manually". It is dead code — never included by subworkflows/local/init.nf.

I-11 — --mode fastqc / fastp / contiglib are not implemented (P2)

nextflow.config and the documentation advertised mode = ["setup", "fastqc", "fastp", "contiglib", "all"], but workflows/viroprofiler.nf only branched on setup and all. Any other value ran everything up to and including CONTIGLIB_CLUSTER and then silently stopped, so --mode fastqc still ran assembly and CheckV. The schema also defaulted mode to test, a value the workflow never mentions.

Resolution. The stages are numbered in WorkflowViroprofiler.MODE_STAGE and each block in the workflow is gated on that number, so a mode is a prefix of the pipeline rather than a branch. --mode fastqc now runs 3 processes, fastp 4, contiglib 8 and all 26 on the stub dataset. Reporting — CUSTOM_DUMPSOFTWAREVERSIONS and MULTIQC — runs in every mode, so a run stopped early still says what it did. initialise() rejects an unknown mode, and rejects fastqc/fastp together with --reads_type clean, which skips both stages and would otherwise produce an empty run.

I-12 — docker.userEmulation was removed from Nextflow (P2)

nextflow.config:144 sets docker.userEmulation = true. The option was deprecated in Nextflow 23.x and removed in 24.x; manifest.nextflowVersion still claims >=22.04.0. Users on a current Nextflow release who pick -profile docker hit an unknown-option error. The manifest's version floor needs to state a range this pipeline has actually been tested against.

I-13 — Stub output directories are committed (P3)

.gitignore lists output_stub*, yet 121 files under output_stub_v2/ and output_stub_contiganno_v2/ are tracked in git — they were added before the ignore rule. They are run artefacts, not fixtures, and they inflate every clone.

I-14 — CLAUDE.md mandates an MCP server that is not part of the repo (P3)

CLAUDE.md instructs assistants to "ALWAYS use the code-review-graph MCP tools BEFORE using Grep/Glob/Read". That server is a local, personal setup; it is not configured in this repository and is unavailable in a clean checkout, so the instruction misfires for every other contributor.

I-15 — contamref_idx ignores --db and nothing ever creates it (P1)

nextflow.config defaults contamref_idx = "${HOME}/viroprofiler/contamination_refs/hg19/ref". Two problems:

  1. The path is anchored to $HOME, not to params.db, so moving the databases with --db leaves the decontamination reference behind.
  2. No process in modules/local/setup_db.nf builds a contamination_refs index, and --mode setup never mentions it. --use_decontam true therefore fails on a clean installation with a bare Nextflow file-not-found:
ERROR ~ No such file or directory: /home/<user>/viroprofiler/contamination_refs/hg19/ref

The default should follow params.db, resolved lazily, and either a setup process should build the index or the failure should carry an actionable message.

I-16 — conf/modules.config was loaded after profiles (P1)

includeConfig 'conf/modules.config' sat at the bottom of nextflow.config, after the profiles block. In Nextflow the later include wins, so the container assignments in modules.config overrode anything a profile set. No profile could redirect the pipeline to a different registry, a local mirror, or a locally built image — an offline or air-gapped cluster had no supported way to substitute images.

Moving the include above profiles makes profile overrides possible; conf/arm64_local.config relies on it.

I-17 — Dockerfiles call wget that was only present transitively (P2)

docker/viroprofiler-taxa/Dockerfile and docker/viroprofiler-binning/Dockerfile download reference data with wget, but neither env_taxa.yml nor env_binning.yml listed it — the binary arrived as an incidental dependency of a pinned package. As soon as the solve changes, the build fails late with wget: command not found (exit 127) after the multi-GB conda step. wget is now declared explicitly.

I-18 — DeepVirFinder was bundled into the binning image (P2)

docker/viroprofiler-binning/Dockerfile created a second conda environment from env_dvf.yml, which pins theano=1.0.3 and keras=2.2.4 — both frozen to 2018, and theano has no linux-aarch64 build at all. One unmaintained tool therefore made the whole binning image (metabat2, vRhyme, phamb) unbuildable on arm64.

DeepVirFinder now lives in docker/viroprofiler-dvf/, and the DVF process carries its own viroprofiler_dvf label. This changes nothing for amd64 users beyond one extra image.

I-19 — Vendored nf-core modules and their containers are from 2022 (P2)

modules/nf-core/modules/ pins FastQC 0.11.9, fastp 0.23.2, SPAdes 3.15.4 and MultiQC 1.12 against quay.io/biocontainers, which publishes amd64 only. They also use the pre-nf-core-3 conda (params.enable_conda ? ... : null) idiom, and params.enable_conda is itself deprecated. Current bioconda has FastQC 0.12.1, fastp 1.3.6, SPAdes 4.3.0 and MultiQC 1.35, all with linux-aarch64 builds.

I-20 — assets/samplesheet_contigs.csv contains a literal ${HOME} (P3)

sample,contigs
contigs,${HOME}/viroprofiler/testdata/viroprofiler-test/contigs.fasta

splitCsv does no shell expansion, so this resolves to a directory literally named ${HOME}. The file is also unused: contig-only runs are driven by --input_contigs, not by a samplesheet.

Resolution. The path is /path/to/contigs.fasta. docs/quickstart.md had the same ${HOME} in its example samplesheet and got the same substitution.

I-21 — 2022-era tools break on a modern Python/setuptools (P1)

Loosening the conda pins so the environments solve on linux-aarch64 also lets the solver pick current Python and setuptools, and three tools break there:

Tool Failure
VirSorter2 2.2.4 ImportError: cannot import name 'load_configfile' from 'snakemake' — removed in snakemake 8
DRAM 1.3.5 ModuleNotFoundError: No module named 'pkg_resources' on Python 3.14
vConTACT2 0.11.3 same pkg_resources failure

pkg_resources cannot be restored just by adding setuptools: setuptools 81 dropped it, and the solver picks 84 by default. The environments therefore pin python=3.10 and setuptools<81, and VirSorter2 now comes from bioconda instead of a pip install -e of git master, so its recipe constrains snakemake for us.

Upgrading DRAM to 1.4.6 does not lift its half of this: mag_annotator/database_handler.py still opens with from pkg_resources import resource_filename, so setuptools<81 stays.

This is the general hazard when reviving a pipeline whose environments were captured in 2022: an unpinned solve is not a "newer, better" environment, it is an untested one.

I-22 — A DRAM 1.4 source file was copied over a DRAM 1.3.5 install (P0)

docker/viroprofiler-geneannot/Dockerfile copied a vendored database_handler.py over mag_annotator/database_handler.py, but env_dram.yml left dram unpinned and the solver resolved 1.3.5. The vendored file is from 1.4.x and imports a helper that does not exist in 1.3.5, so every DRAM-setup.py invocation died on import:

File ".../mag_annotator/database_handler.py", line 17, in <module>
    from mag_annotator.utils import divide_chunks, setup_logger
ImportError: cannot import name 'setup_logger' from 'mag_annotator.utils'

The environment now pins DRAM 1.4.6, whose own database_handler.py is byte-identical to the vendored copy (diff reports only a missing trailing newline). Both the overlay file and its database_handler_bak.py sibling are deleted along with the COPY/cp steps.

DRAM 1.4.6 is installed with pip install --no-deps DRAM-bio==1.4.6 rather than from bioconda: every bioconda build of dram 1.4.6 constrains scikit-bio to <0.6, and conda-forge publishes no scikit-bio below 0.6.0 for linux-aarch64, so the conda solve is unsatisfiable on arm64. bioconda builds that package from the same PyPI sdist, so the installed code is identical; env_dram.yml now lists the run dependencies explicitly.

I-23 — DRAM's CONFIG was hand-written from 2021 filenames and today's date (P1)

bin/create_dram_config.py wrote DRAM's CONFIG by hand before DRAM-setup.py prepare_databases ran. It was wrong in three independent ways:

  • Stale filenames. dbCAN-HMMdb-V10.txt, CAZyDB.07292021.fam-activities.txt, vog_latest_hmms.txt and friends are 2021 release names that DRAM 1.4 no longer produces.
  • Today's date baked into paths. Entries such as refseq_viral.{today}.mmsdb only line up if prepare_databases finishes on the same calendar day the config was written.
  • Wrong schema. DRAM 1.3 read a flat {key: path} dict; 1.4 uses a nested document with search_databases / database_descriptions / dram_sheets / description_db / dram_version keys. A flat file is routed through DatabaseHandler.__construct_from_dram_pre_1_4_0(), a legacy import path.

None of it is needed. prepare_databases builds a DatabaseHandler, calls clear_config() and then set_database_paths() / write_config() after every single database it processes, recording each path.realpath() as it goes. The only thing it cannot do is create the file: DatabaseHandler.load_config() opens it unconditionally, so a missing DRAM_CONFIG_LOCATION is a FileNotFoundError. DB_DRAM therefore seeds the file with DRAM's own DRAM-setup.py export_config --output_file, which copies the empty template that ships inside mag_annotator, and bin/create_dram_config.py is deleted.

export_config has to run with DRAM_CONFIG_LOCATION unset — it resolves its source through the same get_config_loc(), so with the variable already exported it would try to read the file it is supposed to create.

# Due to limitation of container, DRAM database path is hardset to /opt/conda/db2
ln -s ${params.db} /opt/conda/db2

Writing into /opt/conda needs a writable container filesystem. Under Apptainer and Singularity that means --writable-tmpfs, which was only ever set in the untracked, site-specific custom.config, so DRAMV could not run for anybody else.

The symlink is not needed at all. The CONFIG that prepare_databases writes holds absolute paths under --db, and --db is bind-mounted into every task at the identical path (see I-03), so exporting DRAM_CONFIG_LOCATION is sufficient. Verified by replaying the setup and annotate config paths in a docker run --read-only container: the ln -s fails with Read-only file system while the rest of the chain succeeds and leaves the in-image mag_annotator/CONFIG untouched.

One residual sharp edge: set_database_paths() stores path.realpath(), whereas the bind mount is computed with Groovy's getAbsolutePath(), which does not resolve symlinks. A --db that is itself a symlink will therefore be recorded under its resolved name and that name will not be mounted.

I-25 — eggNOG-mapper 2.1.9 needs distutils, removed in Python 3.12 (P1)

env_emapper.yml pinned only eggnog-mapper=2.1.9, so the solver picked Python 3.14 and emapper.py could not start:

File ".../eggnogmapper/common.py", line 6, in <module>
    from distutils.spawn import find_executable
ModuleNotFoundError: No module named 'distutils'

distutils left the standard library in Python 3.12. The environment now pins python<3.12. Same class of defect as I-21, in a second environment of the same image.

I-26 — VIBRANT setup reports success with an unusable database (P0)

DB_VIBRANT calls download-db.sh, whose last two lines are:

echo "VIBRANT databases are downloaded successfully. Please see log file for any error messages."

exit 0

The exit status is unconditional. On a first run against a filesystem that applies a default ACL, the files download-db.sh copies out of the image land without owner read permission, so python VIBRANT_setup.py never starts:

Set VIBRANT_DATA_PATH to <db>/vibrant
Downloading VIBRANT databases to <db>/vibrant...
python: can't open file 'VIBRANT_setup.py': [Errno 13] Permission denied
VIBRANT databases are downloaded successfully. Please see log file for any error messages.

The process exits 0 with a 4 MB directory instead of the expected ~11 GB of pressed HMM profiles. Because the guard is if [ ! -d <db>/vibrant ], every later run then skips the step, and the pipeline fails much later inside VIBRANT itself with an unrelated-looking error.

The setup step now builds into the task work directory, fixes permissions, asserts that VOGDB94_phage, KEGG_profiles_prokaryotes and Pfam-A_v32 all have a non-empty hmmpress index, and only then publishes the directory.

The three sources VIBRANT_setup.py downloads from were re-checked on 2026-08-12 and all resolve: the VOG host redirects fileshare.csb.univie.ac.at to fileshare.lisc.univie.ac.at, and the Pfam and KEGG profile archives are reachable. Both of the latter are fetched over ftp://, however, which many clusters block outbound; the HTTPS mirrors https://ftp.ebi.ac.uk/... and https://www.genome.jp/ftp/... serve the same files.

I-27 — iPHoP built from bioconda carries three known defects (P1)

docker/viroprofiler-host/ installed the bioconda iphop package, which ships:

  1. perl-bioperl<=1.7, which excludes every build bioconda actually publishes (1.7.x). The solver silently skipped the package and RaFAH crashed at run time with Can't locate Bio/SeqIO.pm in @INC.
  2. an unpinned protobuf, so pip resolves 5.x and TensorFlow 2.7 fails at import with TypeError: Descriptors cannot be created directly.
  3. classifier weights tracked with Git LFS, so a plain checkout yields ~130-byte pointer files and the step-8 integrator dies in tf.saved_model.load().

The image is now built from the hardened fork at github.com/rujinlong/iphop, pinned by ARG IPHOP_REF, which fixes all three. It is amd64-only; see ARM64.md.

Separately, iphop --version does not exist in any iPHoP release, so versions.yml recorded an empty iPHoP version. The process now reads iphop.__version__ instead.

I-28 — DRAM's dbCAN downloads return an HTML landing page (P0)

mag_annotator/database_processing.py fetches all three dbCAN files from bcb.unl.edu. That host now redirects every path below /dbCAN2/download/ to the dbCAN home page and answers 200:

$ curl -sIL -o /dev/null -w '%{http_code} %{content_type} %{url_effective}\n' \
    http://bcb.unl.edu/dbCAN2/download/dbCAN-HMMdb-V11.txt
200 text/html; charset=UTF-8 https://pro.unl.edu/dbCAN2/

download_file() uses urlretrieve, which treats that as a successful download and writes the 8 KB landing page to dbCAN-HMMdb-V11.txt, CAZyDB.08062022.fam-activities.txt and CAZyDB.08062022.fam.subfam.ec.txt. hmmpress at least rejects the first one; the other two are description files that nothing validates, so they are parsed straight into description_db.sqlite and every dbCAN annotation comes out as fragments of HTML.

The files themselves are unchanged and still published — only the host moved:

$ curl -sIL -o /dev/null -w '%{http_code} %{content_type}\n' \
    https://pro.unl.edu/dbCAN2/download/dbCAN-HMMdb-V11.txt
200 text/plain

docker/viroprofiler-geneannot/Dockerfile therefore rewrites the host in the installed mag_annotator, and the same RUN asserts that the rewrite matched so a future DRAM release cannot silently skip it. Patching the library is the last resort, but it is the only option here: prepare_databases() collects user-supplied files with

locs = {remove_suffix(i, '_loc'): j for i, j in locals().items() if i.endswith('_loc') and j is not None}

and neither dbcan_fam_activities nor dbcan_subfam_ec carries the _loc suffix, so --dbcan_fam_activities is accepted and then ignored, and the sub-family EC file has no command-line option at all. The same defect makes --vog_annotations a no-op.

Every other source DRAM 1.4.6 uses was re-checked at the same time and is alive (2026-08-12). All of them are tried over ftp:// first with an http(s):// fallback, and on this network FTP is reachable for all three hosts, so both routes work:

Database URL Status
KOfam profiles, KO list ftp.genome.jp/pub/db/kofam/ 200, ~8 MB/s over FTP
Pfam-A.full, Pfam-A.hmm.dat ftp.ebi.ac.uk/pub/databases/Pfam/current_release/ 200, ~3 MB/s; Pfam-A.full.gz is 22.3 GiB
MEROPS pepunit.lib ftp.ebi.ac.uk/pub/databases/merops/current_release/ 200, 436 MiB
RefSeq viral proteins ftp.ncbi.nlm.nih.gov/refseq/release/viral/ 200; one viral.N.protein.faa.gz exists, which is what NUMBER_OF_VIRAL_FILES = 1 expects
VOGDB hmms, annotations fileshare.csb.univie.ac.at/vog/latest/ 301 to fileshare.lisc.univie.ac.at, followed automatically
DRAM distillation sheets raw.githubusercontent.com/WrightonLabCSU/DRAM/master/data/ 200
dbCAN HMMs, family activities, sub-family EC bcb.unl.edu/dbCAN2/download/ dead, see above

DB_VOGDB reaches the same VOGDB host over plain HTTP; that is tracked separately as I-09.

I-29 — DRAM setup cannot tell a finished database from an abandoned one (P1)

DB_DRAM guarded its work with [ ! -d ${params.db}/dram ]. prepare_databases downloads and processes sixteen databases over several hours and cannot resume, so any interruption leaves a directory that satisfies the guard forever: the next run prints "DRAM database already exists", the pipeline reports success, and DRAM-v.py annotate fails much later against a database that was never finished.

The guard is now a completeness check of the CONFIG that DRAM will actually read. Every path it names must exist and be non-empty; none may begin with an HTML document (the failure mode in I-28); the sidecar files that mmseqs and HMMER need but the CONFIG does not name must be present — .h3f/.h3i/.h3m/.h3p for the HMM databases, and for the mmseqs ones both the .idx* k-mer index that mmseqs search needs and the _h* header database that DRAM opens directly to turn a hit into a description; and every description table in description_db.sqlite must have rows. The same check runs after the build, so a database that fails it makes the process exit non-zero instead of being published. An incomplete directory is deleted and rebuilt rather than reused.

The build now runs in the task work directory and is published to ${params.db}/dram only once it is complete, which is what DB_VIBRANT and DB_VREFSEQ already do and what I-31 requires. Because set_database_paths() records every database under the path it was built at, the CONFIG is repointed at the published location afterwards and then re-checked.

That needs room: mmseqs convertmsa turns the 22.3 GiB Pfam-A.full.gz into a 143 GB intermediate, and with the 25 GB of downloads and 25 GB of finished databases alongside it the build peaks at about 190 GB. DB_DRAM refuses to start below 250 GB rather than fill the filesystem two hours in.

prepare_databases never deletes the intermediates it feeds to a step that has finished, and the CONFIG never refers to them, so DB_DRAM removes pfam.mmsmsa, the mmseqs tmp directory and the unpacked KOfam and VOGDB profile trees before publishing. They are larger than everything the build publishes put together.

--skip_uniref is kept: UniRef90 adds several hundred GB and DRAM's own documentation states it does not affect distillation. KEGG is licensed and cannot be downloaded, so kegg and gene_ko_link stay unset; DRAM substitutes KOfam for KEGG orthology.

I-30 — VOGDB moved its profiles into a subdirectory and DRAM silently builds nothing (P0)

vog.hmm.tar.gz used to hold its profiles at the root of the archive. It now nests them:

$ tar -tzf vog.hmm.tar.gz | head -2
hmm/VOG00001.hmm
hmm/VOG00003.hmm

process_vogdb() unpacks the archive and then collects the profiles with glob(path.join(hmm_dir, 'VOG*.hmm')), which matches the top level only. It finds none of the 49116 files, merge_files() writes a zero-byte vog_latest_hmms.txt, and hmmpress stops with "File exists, but appears to be empty?" — two hours into the build, after Pfam has been processed and with no way to resume.

docker/viroprofiler-geneannot/Dockerfile makes the glob recursive, so it no longer depends on the archive's internal layout. --vogdb_loc is not a way out: DRAM unpacks whatever file it is handed and then applies the same glob, so the pipeline would have to download and repack the archive purely to satisfy a hardcoded path.

I-31 — Files unpacked by tar can land unreadable by their own owner (P1)

On a filesystem whose default ACL leaves the owner class empty — access being granted through a named entry instead — GNU tar restores each member's stored mode and produces files that their owner cannot open:

$ getfacl -p /mnt/scratch/db
user::---
user:allen:rwx
default:user::---
default:user:allen:rwx

$ tar xzf probe.tar.gz -C /mnt/scratch/db/probe && ls -l /mnt/scratch/db/probe
**Resolution.** `output_stub_v2/` and `output_stub_contiganno_v2/` are untracked and
deleted; `.gitignore` already carried `output_stub*`, which is why they were invisible
after the first commit that added them.

----rw---- 1 allen uucp 1 probe          # archived as -rw-rw-r--

Only tar is affected: open(), touch + chmod, install -m and cp all produce the mode they asked for on the same directory. DRAM unpacks the KOfam profiles with tar and then opens all 26000 of them, so the build dies with PermissionError an hour in. DB_VREFSEQ hits the same thing when it unpacks mmseqs_vrefseq.tar.gz and works around it with an explicit chmod, attributing it there to the archive's stored modes.

Two things in DB_DRAM follow from this. The build happens in the task work directory rather than under --db, so the databases are assembled where the pipeline computes rather than wherever the user keeps storage; and before any of it starts, the process unpacks a one-file archive and checks that it can read the result, so an unsuitable work directory is reported in seconds with the reason and the fix (-w) instead of an hour later as a bare PermissionError. The published database is then made owner-readable explicitly, and check_dram_db.py opens every file the CONFIG names rather than trusting its mode bits.

I-32 — vConTACT2 replaced by vConTACT3 (P1)

vConTACT2 clusters contigs but does not assign taxonomy, so bin/parse_vContact2_vc.py had to derive one: it split genome_by_genome_overview.csv into query contigs and reference genomes by testing whether the contig name contains NODE_, computed a per-cluster LCA over the reference genomes, and transferred it to the queries. That test is a SPAdes naming convention, so the whole taxonomy assignment silently produced nothing for MEGAHIT, Flye or any pre-assembled contig set, and --assembler other only inverted the test rather than fixing it. vConTACT3 predicts taxonomy natively and marks reference genomes with a Reference boolean, so both the guess and the hand-rolled LCA are gone, along with bin/parse_vContact2_vc.py and bin/combine_taxa.py.

Four things about vConTACT3 are not obvious and each one costs a build or a wrong result.

It cannot be installed from conda on aarch64, and there is no PyPI package. pixi global install -c bioconda vcontact3 and every other conda route fail because fastcluster and jenkspy have no linux-aarch64 conda build. There is no vcontact3 distribution on PyPI at all, so the Bitbucket source is the only option — and the source tree is 3.2.4 against bioconda's 3.0.3. Both blocking packages build from their PyPI sdists, which is why docker/viroprofiler-vcontact3/Dockerfile installs build-essential and pins the source by commit (the project publishes no tags).

Exactly one database version works per release. 3.2.4 accepts version 232 and rejects 223, 228 and 230 outright. --db-version is therefore passed explicitly by TAXONOMY_VCONTACT3; without it vConTACT3 globs --db-path and takes whatever is numerically newest, which on a host that has ever held another release is the wrong one. The version lives in params.vcontact3_db_version so that both DB_VCONTACT3 and TAXONOMY_VCONTACT3 read the same value.

prepare_databases prints [ERROR] Unable to retrieve database 232 ... on runs that succeed. Its exit status is not evidence either way. DB_VCONTACT3 ignores both and checks the artifact instead: mmseqs dbtype <dir>/v232/RefSeq.232.0.3.mmseq_0.3_clu must print Clustering, which a truncated download, an HTML error page or a directory that only got as far as being created cannot do. The database is built in the task work directory and moved into --db only after that check passes.

Its pandas>=2.1.1 has no upper bound. A fresh resolve installs pandas 3.x, published years after this commit and changing copy-on-write and string-dtype semantics throughout — the same trap that broke vConTACT2 with numpy 1.24 and scipy's COO refactor (I-21). The Dockerfile pins pandas>=2.1.1,<3.

Two smaller findings. vConTACT3 3.2.4 never invokes diamond: find_tools looks up only mmseqs (required) and vclust (optional, and it gates only the ani export this pipeline does not request). And it bundles pyrodigal and pyrodigal-gv, so the external gene caller and bin/gene_to_genome.py that vConTACT2 needed are gone too.

docker/viroprofiler-taxa is retired with vConTACT2; TAXONOMY_MERGE is pure Python over the callers' tables and runs in viroprofiler-base, which already carries pandas and click.

I-33 — RESULTS_TSE cannot read its own gzipped abundance inputs (P0)

Unrelated to taxonomy, but it is what the pipeline now fails on once taxonomy completes, and it had never been reached before: no run in this repository has ever produced an .rds, because every earlier one stopped at vConTACT2.

ABUNDANCE writes abundance_contigs_{count,tpm,covered_fraction}.tsv.gz, and vpfkit::read_coverm() opens them with data.table::fread(). fread() cannot decompress a .gz without the R.utils package, and denglab/viroprofiler-viewer does not have it:

$ apptainer exec viroprofiler-viewer.sif Rscript -e 'requireNamespace("R.utils")'
FALSE

Error in fread(fpath) :
  To read gz files directly, fread() requires 'R.utils' package which cannot be
  found. Please install 'R.utils' using 'install.packages('R.utils')'.
Calls: <Anonymous> ... read_coverm -> fread -> stopf -> raise_condition -> signal

The taxonomy input is unaffected: taxa_mmseqs_formatted_all.tsv is not compressed, and read_coverm never opens it.

This cannot be fixed from this repository as it stands, which is the point of I-02: the viewer image has no Dockerfile here, so R.utils cannot be added to it. Either that image gains a Dockerfile and the package, or RESULTS_TSE decompresses the three files before calling create_tse.r.

Resolution. R.utils is in the viewer image and vpfkit lists it under Imports. Verified in the pixi-built image, and end to end: the sixteen-sample run reads the gzipped abundance tables and writes its .rds.

I-34 — -max_target_seqs makes contig dereplication depend on library size (P2)

CONTIGLIB_CLUSTER dereplicated the pooled contig library with the MIUViG recipe: an all-vs-all blastn, anicalc.py to turn the HSPs into ANI and coverage, and aniclust.py to cluster them greedily. The blastn call carried -max_target_seqs 25000, which reads like "keep the best 25000 hits" and is not that. It is a cutoff applied during the search, and NCBI documents that the hits it keeps are not guaranteed to be the best ones — the point of Shah et al., "Misunderstood parameter of NCBI BLAST impacts the correctness of bioinformatics workflows", Bioinformatics 35(9):1613–1614 (2019).

The consequence for this pipeline is that a contig pair can stop being reported because of sequences that have nothing to do with either contig. A query, a 98%-identity full-length partner and -max_target_seqs 5 (the same mechanism as 25000, at a size that is quick to run):

$ blastn -query q.fasta -db db_small -perc_identity 90 -max_target_seqs 5 ...
library=small subjects=1  reported=1 partner_reported=1

$ blastn -query q.fasta -db db_big   -perc_identity 90 -max_target_seqs 5 ...
library=big   subjects=31 reported=5 partner_reported=0

The 30 sequences added to db_big are unrelated to the query–partner relationship, yet the partner is no longer reported at all, so aniclust.py never sees the edge and the two contigs land in different clusters. Representatives are the read-mapping reference for abundance and the input to viral detection, so a pair silently lost this way propagates into apparent abundance and viral calls. Nothing in the outputs records that it happened, and the effect grows with the number of samples pooled.

Fixed. The chain is now Vclust 1.3.1 — a Kmer-db prefilter, LZ-ANI alignment, and greedy clustering — which has no equivalent per-query cap, so the result no longer depends on how many other contigs are in the library. bin/anicalc.py, bin/aniclust.py and bin/parse_NRCLib_clusters.py are deleted; bin/parse_vclust_clusters.py writes the same repid/ctgid table the pipeline published before. The thresholds keep their names and their meaning:

Old New Measure
--min_ani 95 --ani 0.95 Vclust ani — identical nucleotides over the aligned region only
--min_tcov 85, --min_qcov 0 --qcov 0.85 Vclust qcov — aligned fraction of the shorter contig, longer one unconstrained

gani and tani are the wrong measures here: gani divides by the whole query length, so it scores a contained contig as poorly as a diverged one, and tani is symmetric, so it cannot express containment at all. Setting --rcov alongside --qcov would demand that both sequences be covered, which is reciprocal-overlap clustering rather than containment. --algorithm cd-hit is used rather than Vclust's default leiden because it is the same greedy longest-first centroid scheme aniclust.py implemented.

One class of representative does change. When two contigs in a cluster are exactly the same length, aniclust.py kept whichever came first in the FASTA, so the representative followed the order the samples happened to be pooled in; Vclust breaks the tie the same way every time. Both contigs are equally valid representatives, and the new choice is the reproducible one:

input order alpha,beta : aniclust -> alpha_first   vclust -> alpha_first
input order beta,alpha : aniclust -> beta_second   vclust -> alpha_first

On the contig library of the two-sample run in this repository (23 contigs, 22 clusters), the substitution reproduces contigs_nrclib.fasta and contigs_nrclib.dict byte for byte and contigs_ANIclst.tsv row for row. That library is far too small to bound the change on its own, so it was repeated on 1600 sequences — 800 CheckV reference genomes plus derived fragments placed on both sides of the 95 %/85 % boundary: 824 of 825 clusters identical, 1599 of 1600 contigs assigned the same representative, and the single disagreement a pair whose identity the two aligners estimate as 0.9507 (BLAST) and 0.9496 (LZ-ANI), i.e. astride the threshold rather than a difference in what the threshold means.

End to end, the two-sample run (-profile apptainer,arm64_local --use_dram false) still completes and its two TreeSummarizedExperiment objects are byte identical to the ones the BLAST chain produced:

$ cmp <blast-run>/results/viroprofiler_output.rds <vclust-run>/results/viroprofiler_output.rds
$ cmp <blast-run>/results/viroprofiler_output_all_contigs.rds <vclust-run>/results/viroprofiler_output_all_contigs.rds

CONTIGLIB_CLUSTER itself went from 3.3 s to 1.5 s on that library, which is far too small to say anything about how the two scale.

I-35 — Exact duplicate contigs entered the O(n²) dereplication stage (P3)

CONTIGLIB pools the contigs of every sample into one library, so a contig that several samples assembled identically appears once per sample. The all-vs-all blastn compared each of those copies against everything else, and aniclust.py then collapsed them again at the end.

CONTIGLIB_CLUSTER now runs vclust deduplicate first, which drops contigs whose sequence is identical to another contig's, in either orientation, before the prefilter sees them. On the 1600-sequence set above that removed 115 sequences (7 %) with no change to the clustering. The ids it drops are not lost: vclust deduplicate writes a .duplicates.txt companion file, and bin/parse_vclust_clusters.py reads it back so that every contig in the library still appears in contigs_ANIclst.tsv under the representative of its cluster.

I-36 — PHAMB_RF calls a CLI the installed phamb does not have (P1)

The process invoked

run_RF.py -f <contigs> -d <dvf> -p <micomplete> -g <vog> -c <clusters> \
          -l <minlen> -m /opt/phamb/workflows/mag_annotation/dbs/RF_model.python39.sav \
          -s <minbin> -o .

but the phamb the image installs takes four positional arguments and two options:

run_RF.py <fastafile> <clusterspath> <annotationdir> <directoryout> [-m MIN_BIN_SIZE] [-s SEPARATOR]

Every flag was wrong, and two were actively dangerous: -m had become an int minimum bin size and was being handed a filesystem path, while -s had become a binsplit separator and was being handed a size in bases. The model path did not exist either — phamb moved its dbs/ from workflows/mag_annotation/ into the package itself.

The three annotation files are also not passed individually any more. run_RF.py looks them up by fixed name inside annotationdir: all.DVF.predictions.txt, all.hmmVOG.tbl and all.hmmMiComplete105.tbl — none of which matches what MICOMPLETEDB and VOGDB emit (hmmMiComplete.tbl, hmmVOG.tbl).

Resolution. PHAMB_RF builds the annotation directory under the expected names and calls the current CLI. Nothing was passing before, so --binning phamb had been broken for longer than the missing DeepVirFinder table alone would explain.

I-37 — run_RF.py copied onto PATH cannot find its own model (P0)

run_RF.py loads the random forest from a path relative to its own source file:

rf_model_file = Path(__file__).parent / "dbs/RF_model.python39.sav"

The binning Dockerfile did cp /opt/phamb/phamb/*.py /opt/conda/bin/, so the copy that PATH resolves has __file__ in /opt/conda/bin, where there is no dbs/:

$ command -v run_RF.py               -> /opt/conda/bin/run_RF.py
$ ls /opt/conda/bin/dbs/             -> No such file or directory
$ ls /opt/phamb/phamb/dbs/           -> RF_model.python39.sav

joblib.load would therefore have failed with FileNotFoundError — after VAMB and both HMM searches had already run.

Resolution. The image installs a wrapper at /usr/local/bin/run_RF.py that executes the file in its installed location, so __file__ still resolves beside the model. The Dockerfile asserts both the script and the model exist, and the smoke test deserialises the forest.

I-38 — The binning image never installed VAMB (P0)

vMAG_PHAMB runs VAMB before PHAMB_RF, but vamb was not in env_binning.yml and not in the image:

$ apptainer exec viroprofiler-binning.sif vamb --version
/bin/bash: line 1: vamb: command not found

So --binning phamb could not have worked on any architecture, independently of I-36 and I-37.

Resolution on amd64. vamb 4.1.3 is declared in docker/viroprofiler-binning/pixi.toml as a linux-64-only feature and installed when BuildKit's TARGETARCH is amd64.

Not resolvable on aarch64. VAMB has no linux-aarch64 artifact in any release line, and 5.x depends on pycoverm, which has none either. WorkflowViroprofiler.binningIsAvailable() refuses --binning phamb on aarch64 at start-up rather than letting the run reach VAMB. See ARM64.md.

I-39 — The binning image's second vrhyme prefix had drifted to Python 3.14 (P1)

The image built two prefixes containing vRhyme: the base environment, and a separate viroprofiler-vrhyme one from env_vrhyme.yml. Only the base environment was on PATH, so the second was never used — and its unpinned solve had drifted far enough to stop working:

$ /opt/conda/envs/viroprofiler-vrhyme/bin/vRhyme --version
  File "/opt/conda/envs/viroprofiler-vrhyme/bin/vRhyme", line 16, in <module>
    import pkg_resources
ModuleNotFoundError: No module named 'pkg_resources'

It had resolved to Python 3.14, numpy 2.4.6, pandas 3.0.5 and setuptools ≥81, which no longer ships pkg_resources. The vRhyme that VRHYME actually runs, from the base environment, was on Python 3.10 with scikit-learn 1.0.2 and works.

Resolution. There is one environment now, binning, and vrhyme is declared in it — which is where the working executable always came from. Its Python and scikit-learn are pinned, the latter because PHAMB's forest is a joblib-serialised scikit-learn estimator and phamb's own setup.py requires exactly 1.0.2.

I-40 — VAMB was given a depth table covering contigs its FASTA does not contain (P2)

vMAG_PHAMB passes the viral subset as --fasta but the BAMs are mapped against the whole dereplicated library, and jgi_summarize_bam_contig_depths summarises whatever the BAMs contain. VAMB pairs --jgi with --fasta by row, so every contig after the first extra one would have been given another contig's abundance.

The two cut/paste lines that preceded the filter were a no-op: cut -f1-3 and cut -f1-3 --complement pasted back together reproduce the input table exactly.

Resolution. The depth table is restricted to the names in the FASTA before the length filter is applied.

I-41 — CONTIGLIB_CLUSTER requested one CPU while Vclust threads every stage (P3)

vclust deduplicate, prefilter, align and cluster all take -t $task.cpus, and the process passed task.cpus to each of them, but nextflow.config requested cpus = 1 * task.attempt. The Kmer-db prefilter and the LZ-ANI alignment — the two stages that dominate the runtime — ran single-threaded.

Memory was also sized for the viral subset rather than for what this process actually sees: it dereplicates the whole pooled contig library, because its output is the read-mapping reference for abundance.

Resolution. 4 CPUs and 12 GB, both scaled by task.attempt and still bounded by check_max.

I-42 — conda post-link scripts are not run by pixi (P0)

bioconductor-genomeinfodbdata ships no R package. Its conda artifact contains exactly two files, both shell scripts:

$ python3 -c "import json; print(json.load(open('conda-meta/bioconductor-genomeinfodbdata-1.2.11-r43hdfd78af_1.json'))['files'])"
['bin/.bioconductor-genomeinfodbdata-post-link.sh', 'bin/.bioconductor-genomeinfodbdata-pre-unlink.sh']

$ cat bin/.bioconductor-genomeinfodbdata-post-link.sh
#!/bin/bash
installBiocDataPackage.sh "genomeinfodbdata-1.2.11"

The data package itself is downloaded and installed by that post-link script. micromamba runs post-link scripts; pixi does not, and says so only as a build-time notice. So the environment solved and installed cleanly while GenomeInfoDb — a dependency of every SummarizedExperiment package — could not be loaded, and library(vpfkit) failed with

Error in loadNamespace(i, ...) : there is no package called 'GenomeInfoDbData'
ERROR: lazy loading failed for package 'vpfkit'

several layers below the package that was actually missing.

Resolution. docker/viroprofiler-viewer/Dockerfile runs that one script explicitly and asserts the resulting directory exists, rather than setting run-post-link-scripts insecure, which would silently execute any post-link script any future dependency brings in. A second step fails the build if a package other than bioconductor-genomeinfodbdata ever ships one, so the next occurrence is found at build time rather than in a run.

This is a general hazard of the micromamba-to-pixi move, not a defect in either tool: it is invisible in the solve, invisible in the lockfile, and only shows up when something tries to load the package.

I-43 — the pinned VAMB has no --jgi, which is the flag the process passes (P0)

VAMB runs

vamb --outdir out_vamb --fasta $contigs -m ... --jgi depth_clean.txt -o __ --minfasta ...

but the version first pinned in docker/viroprofiler-binning/pixi.toml was 4.1.3, and VAMB 4 removed --jgi entirely. Its argument parser offers only --bamfiles and --rpkm; the word jgi does not appear anywhere in the 4.1.3 package. The process would have died on unrecognized arguments: --jgi.

Nor is 4.x a matter of renaming the flag. It computes depths from BAMs itself and hashes the BAM reference names against the FASTA, refusing a mismatch — and this pipeline deliberately hands it a viral subset FASTA with BAMs mapped against the whole dereplicated library. VAMB 5 goes further and restructures the cluster file phamb.vambtools.read_clusters parses.

Resolution. Pinned to vamb 3.0.2, the release PHAMB was developed against: it has --jgi, and its write_clusters emits the two-column clustername<TAB>contigname file phamb reads. It runs on Python 3.6, which is why it stays in its own pixi environment with its own solve group; the lockfile is what keeps a Python 3.6 environment installable at all.

I-44 — VAMB's --jgi pairs depths to contigs by row, not by name (P1)

vamb.vambtools.load_jgi discards the contig names:

header = next(filehandle)
...
columns = tuple([i for i in range(3, len(fields)) if not fields[i].endswith("-var")])
array = _np.loadtxt(filehandle, dtype=_np.float32, usecols=columns)
return validate_input_array(array)

It returns an N_contigs x N_samples matrix that VAMB then pairs with --fasta by row. So the depth table must hold exactly the sequences of the FASTA, in exactly their order. Anything else silently gives each contig another contig's abundance, and the clustering is computed on the wrong numbers with no error anywhere.

The process filtered the table with csvtk grep -f contigName -P <list>, which preserves the order of the depth file — the BAM reference order — not the order of the FASTA. Those two happen to agree today, because the viral subset is produced by seqkit grep from the same library the bowtie2 index was built from, and both preserve input order. That is a property of the tools involved, not a guarantee either makes.

Resolution. The table is now built in FASTA order explicitly, by looking each sequence up by name, and two assertions follow it: a contig absent from the depth table is fatal (it would mean the BAMs were built against a different library), and the emitted row count must equal the number of FASTA sequences that pass VAMB's own -m length filter.

Two related hardenings went in with it, neither of which had produced a wrong answer yet: bin/genomad_to_dvf.py now passes geNomad's score through as the source text rather than re-formatting it with %.4f, because PHAMB applies round(x, 2) to whatever it reads and 0.00499 re-emitted as 0.0050 rounds to 0.01 instead of 0.00; and the vRhyme model download is checked against a pinned SHA-256 rather than trusted because its URL carries a commit.


I-45

Config — the pipeline does not parse under Nextflow's default config parser. P0. Fixed.

Nextflow 25.x made a new, restricted config language the default. nextflow.config is not written in it, and a run dies before the first process with a message that points at a line of config rather than at the cause:

Error nextflow.config:286:27: Unexpected input: '\n'
 286 |             case 'docker':
ERROR ~ Config parsing failed

Four constructs are rejected, each independently fatal:

  • the switch in the containerOptions closure;
  • def check_max(obj, type), a function definition, used at 133 call sites;
  • the top-level if (!params.igenomes_ignore) — "If statements cannot be mixed with config statements", so a conditional includeConfig has no v2 spelling at all;
  • def trace_timestamp = ..., a variable declaration.

The script parser is stricter too: for loops are gone (workflows/viroprofiler.nf:15), and a top-level statement such as WorkflowMain.initialise(...) in main.nf must move inside a workflow block.

Why it has been invisible. NXF_SYNTAX_PARSER=v1 happens to be set in the interactive shell on the development machine. Every run to date inherited it. A sbatch --export=NIL job does not, which is how it surfaced, and neither does a new user's shell.

The two parsers cannot both be satisfied. Moving the Groovy into lib/ fixes the config under v2 — and breaks it under v1, where an unknown identifier in a config closure resolves to a ConfigObject instead of the class: Unknown method invocation 'checkMax' on ConfigObject type. Likewise env('HOME') is v2-only. So this cannot be done incrementally: config and scripts have to migrate together, in one change, and the pipeline then requires a Nextflow new enough to have the v2 parser. manifest.nextflowVersion would move from >=22.04.0 accordingly.

Resolution. Config and scripts migrated together, and manifest.nextflowVersion is now >=26.04.0. NXF_SYNTAX_PARSER must not be set at all.

  • check_max() and its 137 call sites are gone, replaced by process.resourceLimits. Moving the helper into lib/ was tried first and does not work: a config closure resolves Utils to a ConfigObject, so the call fails with No signature of method: groovy.util.ConfigObject.checkMax(). The assignment sits below the profiles block — see I-51 for what happens when it does not.
  • The switch is an if/else chain. def and if inside a config closure remain legal.
  • The report file names use params.trace_timestamp, assigned by nextflow.config itself from new java.util.Date().format(...). The earlier note here — that the timestamp would have to come from the command line — was wrong: an arbitrary expression is a legal assignment value, and params is what carries it into the four reporting scopes. It is in schema_ignore_params, being an implementation detail rather than an option.
  • "${HOME}" is "${env('HOME')}", which is what makes the legacy parser unusable.
  • Statements at file scope moved into the workflow bodies, for became .each, and the no_file closure became a function — the strict parser resolves no_file(...) as a function name.

Two further defects surfaced during the migration and have their own entries: the completion summary was being printed twice (I-49), and command-line booleans stopped being coerced (I-50).

Verified with the default parser: nextflow lint reports no errors where it reported 54, and all five CI stub jobs, the four-mode ladder and both negative tests pass.


I-46

iPHoP, CheckAMG and DRAM-v results never reached RESULTS_TSE. P1. Fixed.

All three ran, all three published their tables under --outdir, and none of them was wired into the process that builds the TreeSummarizedExperiment. The R object that is the pipeline's headline output therefore carried no host prediction, no auxiliary-gene calls and no per-gene annotation. On the two-sample test CheckAMG produced 210 curated protein calls that nothing downstream could see.

DRAMV did not even emit its annotation table as a named channel — only genes.faa and scaffolds.fna — and ABUNDANCE did not emit log_contig_count.txt, the only place CoverM records each sample's library size.

Resolution. RESULTS_TSE takes eight optional inputs: VIBRANT, geNomad, CheckAMG, iPHoP, DRAM-v, the candidate contig list, the CoverM log and a sample metadata table. When a tool did not run its slot carries an empty placeholder from assets/optional/ and the corresponding create_tse.r argument is omitted, so vpfkit leaves those columns out of rowData rather than joining an empty table — which keeps "switched off" distinguishable from "found nothing".

Sample metadata needs --sample_metadata rather than extra samplesheet columns: INPUT_CHECK rejects a samplesheet that is not exactly three columns. Until this, colData held nothing but the sample name, so no group-wise analysis was possible on the pipeline's own output at all.

The two placeholder tables with fabricated header rows (assets/no_dvf_scores.tsv, assets/no_vibrant_quality.tsv) are gone. They existed only because create_vpftse() indexed columns without checking for them; vpfkit now treats both inputs as optional.


I-47

432 lines of iGenomes reference config that nothing reads. P2. Fixed.

conf/igenomes.config, params.genome, params.igenomes_base, params.igenomes_ignore, WorkflowMain.getGenomeAttribute() and WorkflowViroprofiler.genomeExistsError() were nf-core template scaffolding. No process reads params.genome or params.fasta; getGenomeAttribute is defined and never called. genomeExistsError validated --genome on every run, against a table only it consulted.

Removed. workflows/contig_anno.nf also imported RESULTS_TSE without ever calling it; that import is gone too.


I-48

A scheduler's TMPDIR points outside the container, and DRAM-v dies on it. P1. Fixed.

TMPDIR is inherited from whatever launched the pipeline. Under a batch scheduler it normally names node-local scratch — /localscratch/<user>/slurm/<jobid>/tmp here. Nextflow binds the task directory into the container and nothing else, so a tool that writes to $TMPDIR by absolute path finds no such directory:

tRNAscan-SE ... experienced an error: Unable to open
/localscratch/allen/slurm/2615/tmp/tscan87374.fpass for writing.  Aborting program.

DRAM-v reached that call nine minutes in, after kofam, viral, peptidase, pfam, dbCAN and VOGDB had all completed, and the whole process was lost. Only tRNAscan-SE was affected because it is the one tool in this pipeline that builds an absolute scratch path from TMPDIR; the rest write relative paths into the task directory and never noticed.

It does not reproduce interactively, where TMPDIR is /tmp and Apptainer provides one.

Resolution. conf/base.config sets beforeScript = 'export TMPDIR="$PWD"' for every process. The task directory is bound by definition and lives on the work filesystem, which is where large scratch files belong. Verified by checking that the export reaches all 26 task scripts in a stub run, rather than by assuming a directive took effect.


I-49

The completion summary was printed twice on every run. P2. Fixed.

workflows/viroprofiler.nf and workflows/contig_anno.nf each ended in a file-scope workflow.onComplete { ... NfcoreTemplate.summary(...) }. main.nf includes both entry workflows and invokes one of them, but a module's file-scope statements run at include time, not at invocation — so both handlers were registered and both fired. Every run printed the completion summary twice, and with --email set would have tried to send two e-mails.

Confirmed in isolation: two modules, one file-scope handler each, only one workflow invoked, both handlers ran.

Resolution. There is one handler, in main.nf's entry workflow, and it is registered once. See I-45 for why it is an onComplete: section rather than a closure.


I-50

A boolean given on the command line arrived as a string, and "false" is true. P0. Fixed.

Under the strict parser Nextflow no longer infers a parameter's type from its default, so --use_dram false reaches the pipeline as the String "false". Groovy treats any non-empty string as true, so if (params.use_dram) is true and DRAM-v runs when the user asked for it not to. The same applies to all 19 boolean, 7 integer and 8 number parameters — --max_cpus 4 arrives as "4".

What made this survivable is accidental: JSON-schema validation rejects the run with expected type: Boolean, found: String (false) before the workflow starts. Switch validate_params off and the inversion is silent.

Resolution. main.nf declares the types of all 34 non-string parameters in a params { } block, which is where the strict language expects a parameter's type; values stay in nextflow.config, which still overrides the declaration, and a profile and then the command line override that in turn. Only Boolean, Integer and Float convert a command-line string — Number, Double and BigDecimal reject it outright. Float is the right choice for the schema's number parameters even where the value is a whole number: it leaves 95 an Integer rather than rendering it into a tool's command line as 95.0.

The declaration mirrors "type" in nextflow_schema.json. A new non-string parameter has to be added to both.

Verified: --use_dram false now leaves DRAMV out of the run rather than failing validation.


I-51

process.resourceLimits above the profiles block ignores every profile. P1. Fixed.

resourceLimits replaced check_max() (I-45), but the two are not evaluated at the same time. check_max() ran inside a per-task closure and read params.max_cpus at submission; resourceLimits is a plain map, evaluated where it is written. Written in conf/base.config — the natural home, next to the resource defaults — it freezes to the values in the params block of nextflow.config, because that file is included before profiles is applied.

The effect is silent and specific: --max_cpus on the command line still works, because command-line parameters are resolved before the config is, but test (2 cpus, 6.GB), test_stub (2 cpus, 4.GB) and custom.config (4 cpus, 20.GB) are all ignored, and every process runs at the unrestricted defaults.

Resolution. The assignment lives in nextflow.config below profiles, next to the reporting scopes that are placed there for the same reason. Verified per profile with nextflow config -flat, and end to end: process_high requests 12 cpus, 72.GB and 16.h, and under -profile test_stub no task received more than 2 cpus, 4 GB and 1h.

What that placement still cannot reach. A -c file is applied after the project config, so one that sets only params.max_cpus moves the number the parameter summary prints without moving the cap. Measured: -c with max_cpus = 7 against -profile test_stub, whose ceiling is 2, leaves every task at 2 while the summary says 7. The routes that do work are a profile, -params-file, --max_cpus on the command line, and setting process.resourceLimits directly in the -c file.

WorkflowMain.resourceLimitsAgree() warns at startup when params.max_* and the effective process.resourceLimits disagree, naming both values and the fix. It warns rather than fails because overriding process.resourceLimits in a -c file is the supported route and makes the two disagree by design.


I-52

DB_VCONTACT3 cannot download: curl is not in the image. P0. Fixed.

vcontact3 prepare_databases does not fetch anything itself; it shells out to curl, and denglab/viroprofiler-vcontact3 has neither curl nor wget. The download therefore dies inside vConTACT3's own downloader:

File ".../vcontact3/databases.py", line 103, in download_version
    version_dl_res = utils.execute_stdout(['curl', '-s', '-o', str(dest_path), src_path])
FileNotFoundError: [Errno 2] No such file or directory: 'curl'

The path had never been run. Every database on the development host was built or copied by hand, and --mode setup had only ever been exercised against a --db that already held one.

What kept it from shipping a broken database is the process's own check, which is the shape I-10 and the verification discipline exist for. prepare_databases returned non-zero, the process said so and carried on to verify anyway, and the verification failed on content rather than on status:

prepare_databases returned non-zero; verifying the result anyway
The vConTACT3 database in .../vcontact3_db is unusable:
v232/RefSeq.232.0.3.mmseq_0.3_clu is not an mmseqs clustering database.

Resolution. curl is in docker/viroprofiler-vcontact3/pixi.toml and the lockfile is regenerated. The re-lock added curl and its dependency chain — libcurl, krb5, libnghttp2, libpsl, libedit, libev, icu, keyutils — and moved no package that was already pinned. Verified end to end with the rebuilt image: DB_VCONTACT3 downloads, verifies and publishes 14 GB. DB_GENOMAD, which the failure had stopped from running at all, publishes 1.4 GB in the same run.


I-53

A symlinked database that is not bind-mounted is silently re-downloaded. P2. Open.

Every DB_* process skips its work when the database is already there:

if [ ! -d ${params.db}/checkv ]; then
    checkv download_database ${params.db}
    ...

That test runs inside the container. A --db entry that is a symlink to a path outside it — which is how the databases on the development host are arranged, and what --container_binds exists for — is a dangling link in there, so -d is false and the download starts. DB_CHECKV re-fetched 6.4 GB this way, then failed when its mv landed on the symlink that was there all along.

The failure is loud, but it names the wrong thing: an mv collision rather than a missing bind. Passing --container_binds with the symlink targets avoids it, which is the documented requirement for a run and is just as necessary for --mode setup.

Passing --container_binds with the symlink targets makes every guard fire correctly, which is how the three already-built databases were skipped on the second attempt. That is the workaround, not the fix.

Worth fixing by testing the path on the host side instead — a when: clause reading file("${params.db}/checkv").exists() — so that the guard sees what the user sees.


I-54

docker.yml was invalid YAML for GitHub, so the lockfile gate never ran once. P1. Fixed.

PACKAGING.md, CLAUDE.md and HANDOFF.md all describe pixi lock --check --dry-run as the CI gate that keeps an image's lockfile honest. It had never executed. Two independent faults, either of which alone was enough:

  1. The push_to_registry job's matrix.include had every entry commented out. YAML parses that as include: null, and GitHub rejects an empty matrix at parse time — so the whole workflow file was invalid, check_locks included, even though nothing was wrong with it. The symptom is a run with conclusion: failure and an empty jobs array, which reads like an infrastructure hiccup rather than a syntax error.
  2. on: listed only push: tags: [v*, docker*]. The images and their lockfiles are edited on branches; by the time a tag exists, any drift is already released.

Found by pushing dev_ru and noticing the failed run: gh run list --workflow docker.yml returned exactly one entry, that push, failed. Every earlier commit that touched docker/ had gone through with no lockfile check at all.

The job was removed rather than repaired, because restoring it is a larger piece of work than making the file valid — see HANDOFF.md, "Not done yet" item 3 — and its commented-out matrix is stale in its own right: ten images listed, sixteen present under docker/. on: now fires on pushes and pull requests that touch docker/. All fifteen lockfiles pass.

The general shape is worth keeping in mind: a CI gate that cannot be parsed fails in the same direction as one that passes. Nothing goes red on the branch that broke it, the repository keeps documenting a guarantee it stopped providing, and the only way to notice is to ask when the check last actually ran.


I-55

Two of ABUNDANCE's six CoverM invocations feed no downstream process. P2. Closed — they are published for users, and stay.

ABUNDANCE runs coverm contig six times, once per method, each a separate pass over every BAM in the run:

Method Channel Consumers
count ab_count_ch RESULTS_TSE
tpm ab_tpm_ch RESULTS_TSE
trimmed_mean ab_trmean_ch RESULTS_TSE
covered_fraction ab_covfrac_ch RESULTS_TSE
rpkm ab_rpkm_ch none
reads_per_base ab_rpb_ch none

The last two are emitted and published, and no workflow references either channel. The TSE is built from the first four; bin/create_tse.r has no argument that could take the other two. vpfkit's rpb2bpb() was the one function that consumed a reads_per_base table, and nothing calls it either — it is deprecated as of vpfkit bfc3c9b, because CoverM measures exactly the depth it was estimating.

So a third of the abundance stage re-reads every BAM to produce two files that only ever reach publishDir. That is not free: each pass is I/O-bound over the full alignment set.

Both are kept, deliberately. They are published for users who want to do their own analysis with them — rpkm in particular is a normalization people ask for by name — and "no consumer inside the pipeline" is not the same as "no consumer". The cost is understood and accepted; what was wrong was that nothing recorded which outputs are for RESULTS_TSE and which are for the person reading publishDir.

The optimization that remains available, and is separate from this decision: coverm contig accepts several --methods in one invocation, so all six could come from one pass over the BAMs instead of six. It is not free to adopt — the output becomes one wide table rather than one table per method, which changes RESULTS_TSE's inputs and needs vpfkit's readers changed with it — so it is worth doing when something else is already touching that interface.


I-56

Neither PHAMB database could be built, and in both cases the guard turned the failure silent. P0. Fixed.

I-30 recorded that vog.hmm.tar.gz moved its profiles from the archive root into hmm/, and fixed the consequence for DRAM by making process_vogdb()'s glob recursive. The same upstream change broke a second consumer that was never touched:

cat VOG*.hmm > AllVOG.hmm

With the profiles at hmm/VOG00001.hmm, that glob matches nothing. Reproduced under the pipeline's own shell (process.shell = ['/bin/bash', '-euo', 'pipefail']):

cat cat: 'VOG*.hmm': No such file or directory, exit 1
Script Fails — but > AllVOG.hmm already created the file
Left behind A 0-byte AllVOG.hmm and the unpacked hmm/ directory

The failure is loud the first time. The second time it is silent, because the guard was [ ! -d ${params.db}/vogdb ] and the directory now exists: the process prints "VOGDB database already exists" and succeeds. VOGDB then runs hmmsearch against a 0-byte library, which returns no hits without erroring, and PHAMB's random forest is handed an empty VOG feature column — it still produces bin calls, from one fewer piece of evidence than it was fitted on.

Nothing caught it because --binning phamb has never been run: this database is only ever built for that path.

Fixed together with I-09. DB_VOGDB now finds profiles with find -name 'VOG*.hmm' regardless of archive layout, builds into the task work directory, and verifies the result by counting HMMER3/ magic lines — the header every model in a HMMER3 library begins with, so a count is a parse: a 0-byte concatenation, a truncated download and an HTML error page all give zero. It publishes only above 1000 profiles (vog236 has ~49000).

DB_MICOMPLETEDB turned out to be broken too, independently, and identically. Running the fixed VOGDB logic end to end surfaced it: the miComplete step produced a zero-byte Bact105.hmm. Bitbucket answers wget with 404 and curl with 200, for the same URL, with or without a browser User-Agent:

$ wget -O w.hmm  https://bitbucket.org/.../Bact105.hmm   ->  ERROR 404: Not Found,  0 bytes
$ curl -sSL -o c.hmm https://bitbucket.org/.../Bact105.hmm  ->  http=200, 6716478 bytes

So the same shape as VOGDB: the download fails, wget -O has already created the file, the directory the process made beforehand now exists, and the next run's [ ! -d ] guard skips it and reports success — leaving MICOMPLETEDB to hmmsearch an empty library.

Both PHAMB databases were therefore unbuildable, by two unrelated causes with one signature. DB_MICOMPLETEDB now uses curl -fsSL — -f being what turns an HTTP error into a non-zero exit rather than an error page written to the output file — and verifies that Bact105.hmm holds exactly 105 profiles before publishing. curl is now declared in docker/viroprofiler-base/pixi.toml rather than relied on transitively, which is I-17's lesson; pixi lock reported the lockfile already up to date, so no package moved.

Both databases have now been built with the fixed logic: vog236 gives 49116 profiles (4.4 GB), miComplete 105.

Three lessons this repository already knew, all of which applied here:

  • An exit status proves nothing about a database. The second run's status was 0.
  • [ -d ] is not a completeness check (I-53). Here it did not merely fail to detect a missing bind; it actively converted a hard failure into a silent one.
  • When an upstream layout change breaks one consumer, look for the others. I-30 found the cause and fixed one call site. The grep for the other one is cheap and was never done.

Where the same shape still is

Auditing every DB_* process after this, by whether it checks the content of what it produced rather than by which idiom it uses:

Verifies its output Does not
DB_GENOMAD, DB_CHECKAMG, DB_DRAM, DB_VIBRANT, DB_VCONTACT3, DB_VITAP, DB_VOGDB, DB_MICOMPLETEDB DB_CHECKV, DB_VIRSORTER2, DB_IPHOP, DB_EGGNOG, DB_KRAKEN2

Of the five, three — DB_IPHOP, DB_EGGNOG, DB_KRAKEN2 — also download straight into --db rather than into the task work directory, so a failure leaves exactly the residue that makes the next run skip.

DB_IPHOP has been run, on an HPC cluster, and it worked — so it is not in the position VOGDB was, despite looking similar here. It cannot run on this machine (iPHoP has no aarch64 build), which is why nothing in this repository's own verification record covers it.

That evidence still applies to the current code. DB_IPHOP was touched once during this modernization, by 88d9eb2, which replaced a hardcoded iPHoP_db_Sept21.tar.gz with a *.tar.gz glob — the archive is deleted after a successful download, so the change affects cleanup and not the download. The iphop download call is byte-identical to the one that ran.

What is not covered: it has no content check, so a partial download would leave a directory the guard then skips. It creates that directory before downloading into it, which is I-57's shape — though I-57 itself does not bite here, because it is specific to Docker creating a missing bind-mount source, and the cluster runs Apptainer.

DB_CHECKV and DB_VIRSORTER2 are lower risk only because they have run many times here and their outputs are in use; that is evidence about these particular downloads, not about the processes. DB_EGGNOG and DB_KRAKEN2 serve modules that are off by default.


I-57

A --db path that does not exist yet is created by Docker, as root, and then nothing can write to it. P1. Fixed.

--mode setup exists to build databases into --db, so pointing it at an empty path is the obvious first thing to do. Under -profile docker it failed:

mkdir: cannot create directory '/home/runner/work/.../phamb_db/vogdb': Permission denied

Docker creates a missing bind-mount source itself, and it creates it as root. The container, meanwhile, runs as the invoking user — docker.runOptions = '-u $(id -u):$(id -g)' in the docker profile — so the first DB_* process to reach a mkdir cannot write into the directory Docker just made for it.

Three things made this hard to read:

  • The message names neither Docker nor the mount. It looks like a filesystem permissions problem on the host, where the path is in fact writable.
  • It arrives late. DB_VOGDB downloads 554 MB, unpacks 49116 profiles and concatenates them before it ever tries to publish, so the failure is minutes in and after a great deal of apparently healthy output.
  • It does not reproduce under Apptainer, which does not create missing bind sources and would have refused up front. Every real run on the development machine uses Apptainer.

Fixed by creating the directory on the host, in SETUP, before any container starts — Nextflow evaluates that as the invoking user, so the directory exists and is owned correctly by the time anything is mounted. phamb_entry.nf does the same for its databases stage, which does not go through SETUP.

Found by running the PHAMB path in CI for the first time, which is also the first time this repository ran --mode setup under Docker rather than Apptainer.


I-58

The viewer image could not install vpfkit on amd64, because here was never declared. P0 on that platform. Fixed.

The first amd64 build of viroprofiler-viewer, in CI, died in the install_github() layer:

ERROR: dependency 'here' is not available for package 'vpfkit'

vpfkit imports here and has since 0.6.0. The image's pixi.toml never listed it, and the aarch64 builds on the development machine passed anyway: their solve picked r-golem 0.5.1, which depends on r-here, so the package arrived as a side effect. The linux-64 half of the same lockfile resolves r-golem 1.0.1, which dropped that dependency, and the side effect with it. One lockfile, two platforms, two different package sets — 381 packages against 352, R 4.5 against R 4.3 — and a dependency that one of them had only by accident.

Fixed by declaring r-here in docker/viroprofiler-viewer/pixi.toml. The re-lock adds exactly one package on linux-64 and changes nothing on linux-aarch64, so the local SIF built before the fix is the same image as the one the lock now describes.

The general rule this leaves behind: every package vpfkit imports must be declared in the manifest by name. install_github(dependencies = FALSE) installs nothing on its own, and a dependency that is present only through another package's dependency list is one solver choice away from disappearing.

I-59

viroprofiler-host could not be built by CI: the iPHoP fork it installs was private. P1. Fixed on 2026-09-20 by making the fork public; the image was published the same day.

The Dockerfile fetches iPHoP from https://github.com/rujinlong/iphop.git at a pinned ref, because the bioconda package carries three defects the fork repairs (I-27). That repository was private. An unauthenticated GitHub Actions runner got

fatal: could not read Username for 'https://github.com': No such device or address

on the first git fetch, before any layer of the image exists. The image on Docker Hub, denglab/viroprofiler-host:v0.2.6, dates from 2024-06 and predates the fork entirely.

While it stood, denglab/viroprofiler-host:v1.0.1 was the one tag conf/modules.config named that Docker Hub did not have, the host job was red on every Docker workflow run, and an amd64 run with the default --use_iphop true would have failed at VIRALHOST_IPHOP.

Closed by making rujinlong/iphop public, which also lets anyone rebuild the image a published tag was made from — what the org.opencontainers.image.source label on it promises. The alternative, a fine-grained token in a repository secret passed to the build as a secret mount, would have built the image while leaving it reproducible by nobody outside the organization. After the change a manual run of docker.yml on deng-lab/viroprofiler main with push ticked built and pushed the image.