🧬 Generative Proteomics & Continuous Flow Matching

DFlowNovo

State-of-the-Art De Novo Peptide Sequencing via Continuous-Time Markov Chain Discrete Flow Matching with KnapsackDP Reachability Guidance

πŸ“¦ Model Repository (454 MB Production CKPT) πŸ† 68.28% Strict Exact Match (SOTA) ⚑ 256 spectra/sec GPU πŸ”— Direct Live App Link
DFlowNovo Interactive Gradio Studio

πŸ† State-of-the-Art Accuracy

68.28%
Strict Exact Match on 9-Species Benchmark (vs. 65.41% InstaNovo)

DFlowNovo formulates de novo sequencing as continuous probability flow over discrete amino acid states, completely avoiding auto-regressive error accumulation.

⚑ High-Throughput Inference

256 spec/s
Parallel decoding throughput on GPU (and fast interactive CPU inference)

Decodes full-length peptide sequences across all positions simultaneously in just 20 discrete Euler steps with O(1) bitmask reachability lookups.

πŸ§ͺ Native PTM & UNIMOD Support

Extended PTMs
Oxidation, Carbamidomethylation, Deamidation, Phosphorylation

Native support for post-translational modifications with standard UNIMOD annotations and sub-10 ppm precursor mass precision.

πŸ“Š Canonical Nine-Species Test Benchmark Comparison

Empirical evaluation across all 9 standard benchmark species (50,000 test spectra):

Model Architecture Strict Exact Match I/L Exact Match Peptide Recall AA Precision Throughput (spec/s)
🧬 DFlowNovo (Production) 68.28% 73.41% 82.15% 88.94% 256.4
InstaNovo+ (Transformer + Reranker) 65.41% 70.82% 79.80% 86.71% 94.2
Casanovo (Autoregressive Transformer) 61.20% 66.45% 75.30% 83.20% 38.6
PointNovo (Order-invariant Point Net) 54.10% 59.80% 68.90% 77.40% 18.1
DeepNovo (CNN + LSTM) 48.30% 53.20% 62.10% 71.50% 8.4

πŸ’» Quickstart with Python API

Download checkpoint and run inference via Hugging Face Hub in 5 lines of code:

from huggingface_hub import hf_hub_download import torch from train.io import load_checkpoint, load_models_from_checkpoint from train.factory import build_models from inference.predict import predict_peptide device = torch.device('cuda' if torch.cuda.is_available() else 'cpu') # 1. Download production weights from Hugging Face Hub ckpt_path = hf_hub_download(repo_id='joelinator/dflow-novo-model', filename='frozen_production_model.ckpt') ckpt = load_checkpoint(ckpt_path, map_location=device) vocab = ckpt['vocabulary'] # 2. Build models & load EMA weights enc, lp, dec, guid = build_models(vocab, device) load_models_from_checkpoint(ckpt, enc, lp, dec, guid, use_ema=True) enc.eval(); lp.eval(); dec.eval(); guid.eval() # 3. Predict peptide sequences # x_t, lengths, seqs, scores = predict_peptide(mz_array, intensity_array, precursor_mass, precursor_charge, ...)