Skip to content
LabShengLiPublic

About

Long read pipeline

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

LongVerse - one sentence to the allele-specific methylome: verified multi-agent AI for Nanopore and PacBio long-read sequencing

LongVerse is a unified Nextflow DSL2 pipeline for long-read sequencing (Oxford Nanopore Technologies, ONT, and PacBio HiFi), driven either by the command line or by an agent that composes the command from a plain-English request. It takes raw reads through basecalling to a modified-base BAM, extracts per-read and per-site methylation, and from there calls variants, phases reads into haplotypes, finds structural variants, and tests for differentially methylated cytosines.


1. Pipeline overview

LongVerse exposes three mutually exclusive branches, selected by --platform and --ecosystem:

Branch Flags Input What it runs
(1) ONT - Dorado (default) --platform ont --ecosystem dorado POD5 / FAST5 (or BAM via --input_bam) DORADO_UNTAR -> DORADO_CALL (modBAM) -> LV_QC -> LV_CALL_EXTRACT (per-read) -> UNIFY (per-site, NANOME format) -> optional MODKIT_BEDMETHYL, LV_SNIFFLES2, LV_CLAIR3 / LV_CLAIRS_TO, LV_PHASING (HP1/HP2) -> LV_DMC_METHYLKIT / LV_DMC_DSS -> MULTIQC_METHYLATION (site-level HTML)
(2) PacBio - Jasmine --platform pacbio --ecosystem jasmine HiFi unaligned BAM / tar PB_UNTAR -> PB_JASMINE (5mC) -> PB_PBMM2 (align) -> shared LV_* stack identical to branch (1)
(3) ONT - Guppy legacy --platform ont --ecosystem guppy FAST5 / POD5 UNTAR -> BASECALL (Guppy) -> ALIGNMENT -> QCEXPORT -> optional RESQUIGGLE and the legacy methylation tools: NANOPOLISH, MEGALODON, DEEPSIGNAL (v1 / v2), Guppy6, Tombo, DeepMod, METEORE, consensus NANOME (XGBoost), EVAL, REPORT, CLAIR3 / PHASING

See docs/pipeline.md for a stage-by-stage description and DAG.

High-level flow

raw reads ──▶ basecall ──▶ modBAM ──▶ per-read extract ──▶ per-site unify ──┐
   (POD5/FAST5/HiFi)      (Dorado /                                         │
                           Jasmine /                                        ├──▶ methylation HTML report
                           Guppy)                                           │     (MULTIQC_METHYLATION)
                                  ├──▶ modkit bedmethyl (optional)          │
                                  ├──▶ Sniffles2 SVs (optional)             │
                                  ├──▶ Clair3 / ClairS-TO variants ─────▶ phasing (HP1/HP2)
                                  └──▶ methylKit / DSS DMC (HP1 vs HP2)

2. Requirements

  • Nextflow ≥ 20.07.1 (DSL2, NF26 strict syntax).
  • Java 17+ (required by recent Nextflow versions).
  • One container engine: Docker, Singularity / Apptainer, or Conda.
  • GPU strongly recommended for DORADO_CALL, BASECALL (Guppy), MEGALODON, DEEPSIGNAL2, PB_JASMINE. CPU-only is supported for everything else.

LongVerse ships preconfigured container images (see nextflow.config, profiles.docker / profiles.singularity); you generally do not need to build anything by hand.


3. Agent: run it in plain sentences

LongVerse ships an MCP server, so the same pipeline can be driven from Claude Code, Claude Desktop, or any other stdio MCP host.

python3 -m venv ~/.longverse/venv
~/.longverse/venv/bin/pip install \
  "git+https://github.com/LabShengLi/longverse.git#subdirectory=agents/longverse-mcp"
claude mcp add --scope user longverse -- ~/.longverse/venv/bin/longverse-mcp

Any Python 3.10 or newer works, and nothing else has to be installed first. With uv the same thing is one line, which is what most MCP documentation shows:

claude mcp add --scope user longverse -- uvx --from \
  "git+https://github.com/LabShengLi/longverse.git#subdirectory=agents/longverse-mcp" \
  longverse-mcp

Use the first form unless you already have uv: claude mcp add records the command without checking that it exists, so a missing uvx is not reported until the host tries to start the server. Check the registration before relying on it:

claude mcp list
longverse: /home/you/.longverse/venv/bin/longverse-mcp - OK Connected

A server that was registered with a command that is not on PATH shows up here instead as Failed to connect - ENOENT: Executable not found in $PATH, which is the one failure the install itself will not tell you about.

Neither install line names a branch, so both take the repository's default. Nothing here needs editing when the development branch moves.

From the command line

Start Claude Code in whatever directory you want the run to happen in and type the sentence:

claude

Basecall and call CpG methylation for ONT data.

The first time a LongVerse tool is used Claude asks whether to allow it; after that the turn proceeds on its own. What comes back is a nextflow run command, already validated, with the demo data filled in. Nothing is submitted until you ask for it: composing a command and launching a run are separate tools, and only the first one answers a request that describes an analysis.

For scripting there is --print, which takes the sentence on stdin and returns one answer:

echo "Basecall and call CpG methylation for ONT data." \
  | claude --print --allowedTools "mcp__longverse__*"

--allowedTools is not optional here. Without it the tools are registered but not permitted, and the model answers from whatever else it can reach instead of asking the server; the reply looks reasonable and did not come from LongVerse. Some managed installations require the approval interactively whatever is passed, in which case the interactive form above is the one that works.

From the Claude Code desktop app

The desktop app reads the same user config claude mcp add writes, so if you have run the install above there is nothing else to do: open the app and type the sentence.

To configure it without the CLI, add the server to the app's JSON config:

{
  "mcpServers": {
    "longverse": {
      "command": "/absolute/path/to/.longverse/venv/bin/longverse-mcp",
      "args": []
    }
  }
}

The path has to be absolute: the app does not start from a shell and does not inherit your PATH, which is the usual reason a server that works in a terminal does not appear in the app. Restart the app after editing. Claude Desktop keeps its own config; on macOS and WSL claude mcp add-from-claude-desktop imports servers from it rather than retyping them.

Then type a sentence:

Basecall and call CpG methylation for ONT data.

Phase the ONT reads and give me methylation per haplotype.

On a cluster, say where the caches go, first. Container images are stored in $LONGVERSE_HOME, which defaults to ~/.longverse; a home directory under a quota is the wrong place for them.

export LONGVERSE_HOME=/path/to/scratch/$USER/longverse

Prerequisites it handles for you. Java and Nextflow are installed into $LONGVERSE_HOME/toolchain the first time they are needed, no root required; delete that one directory to undo it, or set LONGVERSE_AUTO_INSTALL=0 to be asked instead. longverse_check_environment reports what this host has, and longverse_setup_toolchain installs ahead of time. A container engine is the one thing the agent cannot install: Docker needs a daemon and root, Apptainer needs site configuration, so it is reported and left to you.

Your data, or the demo. Give a path and it uses your data. Name a genome in words (hg38, mm10, T2T/chm13) and it resolves it the way the pipeline does, otherwise hg38; a bare FASTA, a .fa.gz, a directory or a tar.gz all work, indexed or not. Name nothing and it runs the bundled chr20 GNAS demo and says so on its first line, so a demo result is never mistaken for your own.

Before it launches it inspects the input: POD5 or BAM, aligned or not, whether MM/ML tags are present and which platform wrote them, and chooses basecalling or extraction accordingly. It derives the container bind mounts the run needs and reports a bind failure as a bind failure rather than as a missing file. With no GPU it basecalls on CPU. Say HPC or SLURM in the sentence and it submits through the scheduler instead of running locally.

When it finishes it reports the results directory and the key files by role: aligned and per-haplotype BAM, QC statistics and report, per-site and per-read methylation, the Clair3 VCF, the differential methylation tables, and a coverage and methylation-frequency figure it draws automatically.

Platform notes: Linux and macOS run this directly. On Windows, Nextflow needs WSL2; install it with wsl --install and run everything inside the Linux distribution, with Docker Desktop on the WSL2 backend as the container engine.

Tool list and client configuration: agents/README.md.


4. Minimal test command (GitHub CI)

The same command the CI in .github/workflows/ci.yml runs on every push / PR. It downloads a small E. coli POD5/BAM test bundle from Zenodo and produces a complete methylation result set on a GitHub-hosted runner.

# 1. install Nextflow (one-liner)
curl -fsSL https://get.nextflow.io | bash
sudo mv nextflow /usr/local/bin/

# 2. run the bundled smoke test directly from GitHub (Docker)
nextflow run LabShengLi/longverse -profile test,docker

Equivalent forms:

# pin a branch or tag
nextflow run LabShengLi/longverse -r main -profile test,docker

# PacBio HiFi smoke test
nextflow run LabShengLi/longverse -profile test_pacbio,docker

# Singularity / Apptainer on HPC
nextflow run LabShengLi/longverse -profile test,singularity

All built-in test profiles are defined under conf/examples/ - see docs/profiles.md.

Results layout

After a successful test, outputs are written to results/ (override with --outdir):

results/
├── <dsname>-methylation-callings/
│   ├── Raw_Results-<dsname>/<dsname>.dorado_call/     # modBAM + bai, straight out of the caller
│   ├── Read_Level-<dsname>_{all,HP1,HP2}/             # per-read CpG calls
│   ├── Site_Level-<dsname>_{all,HP1,HP2}/             # per-site methylation, NANOME / methylKit / DSS
│   └── <dsname>_methylation_site_report.html          # site-level report (when runMultiqc=true)
├── <dsname>-vcall/                                    # with --phasing
│   ├── <dsname>_clair3_out/                           # variants
│   └── <dsname>_phased_bam/<dsname>_{HP1,HP2}/        # the haplotype BAMs
├── <dsname>-LVQC/                                     # read length, quality, coverage
├── <dsname>_DMC/                                      # differential methylation, HP1 vs HP2
└── <dsname>-run-log/                                  # per-step run logs

The HP1 and HP2 subdirectories appear only with `--phasing`, and `<dsname>_DMC/`
only when a DMC caller runs on them.

5. Running on real data

# ONT Dorado, raw POD5
nextflow run LabShengLi/longverse \
    -profile docker \
    --dsname  MySample \
    --input   '/path/to/pod5_or.tar.gz' \
    --genome  /path/to/hg38_dir          \
    --dorado_basecall_model dna_r10.4.1_e8.2_400bps_hac@v5.0.0 \
    --dorado_methcall_model dna_r10.4.1_e8.2_400bps_hac@v5.0.0_5mCG_5hmCG@v3 \
    --phasing --run_dss --sv_call --runModkit
# PacBio HiFi (unaligned BAM)
nextflow run LabShengLi/longverse \
    -profile docker \
    --platform pacbio --ecosystem jasmine \
    --dsname  HG002_HiFi \
    --input   '/path/to/hifi.tar.gz' \
    --genome  /path/to/hg38_dir
# USC CARC Slurm
nextflow run LabShengLi/longverse \
    -profile singularity,hpc \
    --account <your_slurm_account> --gpu_gresOptions 'gpu:1' \
    --dsname MySample --input ... --genome ...

The full list of options is in docs/parameters/ - see section 6.


6. Documentation

Detailed documentation for every CLI / config parameter lives in the docs/ directory:

You can also print the in-pipeline help message:

nextflow run LabShengLi/longverse --help

7. Data and figures

Test datasets the commands above download are archived on Zenodo at 10.5281/zenodo.20116126.

Figures in the LongVerse paper can be redrawn in a browser at LabShengLi/longverse-figures.


LongVerse is released under the MIT License.

About

Long read pipeline

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages