π State-of-the-Art Accuracy
68.28%
Strict Exact Match on 9-Species Benchmark (vs. 65.41% InstaNovo)
DFlowNovo formulates de novo sequencing as continuous probability flow over discrete amino acid states, completely avoiding auto-regressive error accumulation.
β‘ High-Throughput Inference
256 spec/s
Parallel decoding throughput on GPU (and fast interactive CPU inference)
Decodes full-length peptide sequences across all positions simultaneously in just 20 discrete Euler steps with O(1) bitmask reachability lookups.
π§ͺ Native PTM & UNIMOD Support
Extended PTMs
Oxidation, Carbamidomethylation, Deamidation, Phosphorylation
Native support for post-translational modifications with standard UNIMOD annotations and sub-10 ppm precursor mass precision.
π Canonical Nine-Species Test Benchmark Comparison
Empirical evaluation across all 9 standard benchmark species (50,000 test spectra):
| Model Architecture | Strict Exact Match | I/L Exact Match | Peptide Recall | AA Precision | Throughput (spec/s) |
|---|---|---|---|---|---|
| 𧬠DFlowNovo (Production) | 68.28% | 73.41% | 82.15% | 88.94% | 256.4 |
| InstaNovo+ (Transformer + Reranker) | 65.41% | 70.82% | 79.80% | 86.71% | 94.2 |
| Casanovo (Autoregressive Transformer) | 61.20% | 66.45% | 75.30% | 83.20% | 38.6 |
| PointNovo (Order-invariant Point Net) | 54.10% | 59.80% | 68.90% | 77.40% | 18.1 |
| DeepNovo (CNN + LSTM) | 48.30% | 53.20% | 62.10% | 71.50% | 8.4 |
π» Quickstart with Python API
Download checkpoint and run inference via Hugging Face Hub in 5 lines of code:
from huggingface_hub import hf_hub_download
import torch
from train.io import load_checkpoint, load_models_from_checkpoint
from train.factory import build_models
from inference.predict import predict_peptide
device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
# 1. Download production weights from Hugging Face Hub
ckpt_path = hf_hub_download(repo_id='joelinator/dflow-novo-model', filename='frozen_production_model.ckpt')
ckpt = load_checkpoint(ckpt_path, map_location=device)
vocab = ckpt['vocabulary']
# 2. Build models & load EMA weights
enc, lp, dec, guid = build_models(vocab, device)
load_models_from_checkpoint(ckpt, enc, lp, dec, guid, use_ema=True)
enc.eval(); lp.eval(); dec.eval(); guid.eval()
# 3. Predict peptide sequences
# x_t, lengths, seqs, scores = predict_peptide(mz_array, intensity_array, precursor_mass, precursor_charge, ...)