Molecular & Computational Biologist

Krzysztof
Gizynski

Experience in molecular diagnostics, assay development, and computational biology. This portfolio brings together selected projects exploring diagnostic assay design, bioinformatics, AI-assisted scientific workflows, and computational methods for molecular biology.

Genomics, End to End

I am a molecular biologist with a microbiology background and industry experience in molecular diagnostics, assay development and biotechnology.

Alongside laboratory research and product development, I have increasingly worked in computational biology, developing reproducible bioinformatics workflows for microbial genomics and diagnostic applications. Recent work extends this to human clinical genomics: variant calling, ACMG/AMP classification, pharmacogenomics and somatic tumour–normal analysis.

This website brings together selected projects, case studies and technical notes from both areas.

833 Bacterial genomes processed across projects
4 Pathogens profiled
DSL2 Nextflow DSL2 — the default for pipeline orchestration
AMR Resistance surveillance focus
1.0/1.0 GIAB HG002 recall/precision (clinical pipeline)
8/8 Somatic PASS calls confirmed (COLO829 tumour–normal)
38 Real GeT-RM (CDC) reference samples validated (PGx, 5 genes)
3 Inheritance patterns handled (AR · X-linked · AD) in the ACMG/AMP classifier

Case Studies

Neisseria gonorrhoeae
AMR Diagnostic Pipeline
GitHub

End-to-end Nextflow DSL2 pipeline characterising antimicrobial resistance across 283 clinical strains. Integrates AMRFinderPlus resistance profiling with a custom alignment-based mutation caller (penA, gyrA, parC, mtrR, porB) and promoter-variant detection, core-genome phylogeny via Parsnp, and an interactive Dash dashboard linking resistance phenotypes to phylogenetic context for rapid clinical interpretation.

283 strains
Reference: FA1090
60% ciprofloxacin resistance (GyrA QRDR)
51% reduced cephalosporin susceptibility (mosaic penA)
Mean 10.6 AMR elements per strain
Dual-track: AMRFinderPlus + custom mutation caller
mtrR · porB promoter-variant analysis
Core-genome phylogeny
Interactive dashboard
Nextflow DSL2 AMRFinderPlus Parsnp BLAST+MAFFT Dash Docker Python
View Workflow Diagram
N. gonorrhoeae Diagnostic Pipeline Workflow Two-stream pipeline diagram: AMR and promoter analysis, and core genome phylogeny, converging into a Dash dashboard. N. gonorrhoeae Diagnostic Pipeline 283 clinical strains · FA1090 reference · Nextflow DSL2 · Docker 283 N. gonorrhoeae Genomes NCBI RefSeq · complete assemblies Prokka Annotation GFF · protein FASTA ① AMR Profiling ② Core Genome Phylogeny AMRFinderPlus Resistance genes · point mutations QRDR Mutation Calling penA · gyrA · parC · mtrR · porB BLAST+MAFFT alignment Promoter Variant Analysis mtrR · porB · IR detection Profile Aggregation Per-strain · cross-strain summary Resistance Profiles Per-strain TSV · 54 elements · 8 drug classes Parsnp Core-genome SNP alignment SNP Distance Matrix Pairwise distances · clustering ML Phylogeny FastTree · SNP-based tree Clade Annotation 4 major clades identified AMR Burden Overlay Clonality vs convergence Interactive Dash Dashboard 5 tabs · AMR overview · Mutations · Phylogeny · Promoters Analysis step Key output Dashboard
Streptococcus pyogenes
AMR & Typing Pipeline
GitHub

Scalable pipeline processing 375 publicly available S. pyogenes genomes with integrated resistance profiling, MLST, and emm typing. Identified 22.9% AMR prevalence across the cohort, with ST28 as the dominant sequence type (22.4%) and emm5 the most prevalent emm type (47.5%). Designed for rapid integration of newly sequenced clinical isolates into ongoing surveillance.

375 genomes
22.9% AMR prevalence
ST28 dominant (22.4%)
emm5 dominant (47.5%)
Nextflow DSL2 AMRFinderPlus MLST emmtyper Docker Python
View Workflow Diagram
S. pyogenes AMR and Typing Pipeline Workflow Three-stream pipeline: AMR profiling, MLST typing, and emm typing of 375 S. pyogenes genomes, converging into an integrated report. S. pyogenes AMR & Typing Pipeline 375 clinical strains · pubMLST spyogenes scheme · Nextflow DSL2 · Docker 375 S. pyogenes Genomes NCBI RefSeq · complete assemblies ① AMR Profiling ② MLST Typing ③ emm Typing AMRFinderPlus Resistance genes · Streptococcus_pyogenes Resistance Classification mef · erm macrolide · tet genes POINT Mutation Screening --plus flag · no beta-lactam resistance detected Profile Aggregation Per-strain · cross-strain summary AMR Prevalence: 22.9% Macrolide · tetracycline dominant MLST 7-locus gki · gtr · murI · mutS · recP · xpt · yqiL Allele Profiling PubMLST spyogenes database ST Assignment Sequence type per genome ST Frequency Distribution 133 unique STs · ranked by prevalence ST28 Dominant: 22.4% 133 distinct STs identified emmtyper In silico emm gene typing emm Type Assignment CDC emm database Prevalence Ranking Type frequency analysis AMR × emm Cross-tabulation AMR carriage rate per emm type emm5 Dominant: 47.5% 11 emm types identified Integrated AMR + Typing Report Per-strain TSV · cohort summary · 375 genomes · 47 min runtime Analysis step Key output
Chlamydia trachomatis · Assay Design
Validated Workflow for Diagnostic Target Selection
Multiplex qPCR/TaqMan In-silico Design

Reproducible target-selection workflow for a multiplex qPCR/TaqMan assay for C. trachomatis. Genome QC and tiering are performed before target discovery, ensuring that consensus sequences and inclusivity assessments are built only from high-confidence assemblies. A provenance audit traced 149 GenBank-only genomes to a single study of engineered laboratory strains; these were excluded from the diagnostic reference set, which was rebuilt from 125 complete RefSeq assemblies that were independently verified and tiered.

Four chromosomal loci (mutL, aroB, PSD2, and group_306) were independently evaluated as candidate targets alongside the multicopy diagnostic plasmid, which serves as the higher-sensitivity target. The proposed duplex pairs one chromosomal locus with the plasmid target, combining genomic redundancy with sensitivity.

The workflow integrates pan-genome analysis, sequence-conservation and inclusivity screening, exclusivity BLAST, quantitative Tm-gap assessment, and three-way structural QC of primer–probe sets, producing a computationally screened candidate panel for downstream experimental validation.

Chromosomal targets — mutL · aroB · PSD2 · group_306
125 genomes (97 primary / 10 review / 18 excluded)
149 GenBank-only engineered-strain genomes rejected
874 core genes (Panaroo pan-genome)
692/874 exclusivity-clean
4 targets · 7 oligo sets
Plasmid target — Primary + Reserve windows
34 plasmid-bearing genomes
Circular-alignment artifact caught & fixed (7,606 bp)
1 plasmid target · 2 candidate windows · 6 oligos
Tm gap 6.08–8.07°C across all 9 sets (≥5°C floor)
Panaroo MAFFT fastANI primer3-py NCBI BLAST Three-way dimer QC Python
View Workflow Diagram
C. trachomatis Multiplex Assay Design Workflow Two parallel streams from the same organism — chromosomal targets (125 genomes) and the diagnostic plasmid (34 genomes) — each processed through pan-genome/consensus analysis, exclusivity screening, Primer3 design, and three-way structural QC, converging into one locked multiplex panel. C. trachomatis Multiplex Assay Design Inclusivity · Exclusivity · Tm-Gap Rule · Three-Way Dimer QC · Primer3 Two genomic compartments, one QC standard ① Chromosomal Targets ② Plasmid Target 125 C. trachomatis Genomes RefSeq only · 149 engineered GenBank genomes rejected Genome QC & Tiering 97 primary · 10 quality-review · 18 excluded Panaroo Pan-genome 904 clusters → 874 core (≥99%) Exclusivity BLAST 692/874 core genes exclusivity-clean Primer3 Design Tm-gap ≥5°C rule · real-computed Tm Three-Way Dimer + Structure QC primer3-py hairpin/homo/heterodimer mutL · aroB · PSD2 · group_306 4 targets · 7 oligo sets 34 Plasmid-Bearing Genomes Circular contig, 7,415–7,510 bp Rotation-Corrected Alignment Fixed 2×-length MAFFT artifact Consensus Build 7,606 bp corrected consensus Primer3 Design Same Tm-gap ≥5°C standard Three-Way Dimer + Structure QC Same primer3-py standard BLAST Specificity Human genome + core_nt · pair-level rule Primary + Reserve Windows One plasmid · 2 candidate windows · 6 oligos Validated Diagnostic Candidate Targets 5 targets validated (4 chromosomal + 1 plasmid) · 9 oligo sets · 27 oligos Analysis step Key output
Klebsiella pneumoniae
Phylogenomics Pipeline
GitHub

Comprehensive phylogenomics pipeline spanning the full analysis chain for 50 K. pneumoniae genomes. Assembly QC (QUAST), AMR and resistance gene profiling (AMRFinderPlus), virulence and typing (Kleborate), SNP-based variant calling (Snippy), recombination masking (Gubbins), and maximum-likelihood phylogeny (IQ-TREE). Results delivered via a six-tab Dash dashboard. ST11 identified as the dominant lineage, consistent with global carbapenem-resistant K. pneumoniae trends.

50 genomes
ST11 dominant
6-module pipeline
6-tab dashboard
Full QUAST → IQ-TREE chain
Nextflow DSL2 QUAST AMRFinderPlus Kleborate Snippy Gubbins IQ-TREE Dash Docker
View Workflow Diagram
K. pneumoniae Phylogenomics Pipeline Workflow Three-stream pipeline: AMR profiling, molecular typing, and core genome phylogeny of 50 K. pneumoniae genomes, converging into a 6-tab Dash dashboard. K. pneumoniae Phylogenomics Pipeline 50 clinical genomes · Pasteur MLST scheme · Nextflow DSL2 · Docker 50 K. pneumoniae Genomes NCBI RefSeq · complete assemblies QUAST Quality Control N50 · contig count · assembly size ① AMR Profiling ② Molecular Typing ③ Core Genome Phylogeny AMRFinderPlus Resistance genes · Klebsiella_pneumoniae β-lactamase Detection ESBL · KPC · OXA · NDM classes Resistance Gene Matrix Per-strain binary presence/absence Drug Class Summary Cross-strain aggregation Resistance Profiles ESBL · KPC · carbapenem prevalence Kleborate MLST · virulence loci · capsule ST Assignment Pasteur scheme · 7 loci Virulence Scoring ybt · clb · iuc · iro · rmpA loci Clonal Group Analysis CG258 · high-risk lineages ST11 Dominant ST258 · ST307 · ST147 · ST16 Snippy SNP calling vs. reference genome Gubbins Recombination masking IQ-TREE Maximum likelihood · GTR+G model snp-dists Pairwise SNP distance matrix Phylogenetic Tree Clonal clusters · AMR burden overlay Interactive Dash Dashboard 6 tabs · QC · AMR · Typing · Virulence · Phylogeny · Comparison Analysis step Key output Dashboard
CAPN3 / DMD
Variant Calling Pipeline
GitHub

Reads-to-VCF pipeline scoped to the CAPN3 and DMD loci, independently benchmarked against Genome in a Bottle's HG002 reference sample. Raw reads are QC'd (FastQC/MultiQC) and adapter/quality-trimmed (fastp) before alignment. Two variant callers — GATK HaplotypeCaller and DeepVariant — are run independently and cross-checked against each other, then the concordant call set is annotated with Ensembl VEP transcript consequence and gnomAD population frequency, producing evidence ready for the downstream ACMG classifier.

1.0/1.0 SNP + INDEL recall/precision (GIAB HG002)
11-module Nextflow DSL2 pipeline
Two callers cross-checked: GATK + DeepVariant
2,267 concordant calls annotated (VEP + gnomAD v4.1)
fastp-trimmed reads, FastQC/MultiQC QC'd pre- and post-trim
Remote HTTP-range extraction — no full-genome download
CAPN3 (chr15) + DMD (chrX) scoped
Nextflow DSL2 FastQC fastp BWA-MEM2 GATK DeepVariant VEP gnomAD hap.py Docker
View Workflow Diagram
CAPN3/DMD Variant Calling Pipeline Workflow Reads-to-VCF pipeline: raw-read QC and adapter/quality trimming, then alignment and dual-caller variant calling converging into cross-check, then forking into GIAB benchmark and VEP/gnomAD annotation, feeding the ACMG classifier. CAPN3/DMD Variant Calling Pipeline CAPN3 (chr15) + DMD (chrX) · GIAB HG002 benchmark · Nextflow DSL2 · Docker HG002 GRCh38 BAM (remote) GIAB · HTTP range-extracted, no full download EXTRACT_REGION region FASTQ (R1/R2) FASTQ_QC (raw) FastQC + MultiQC · pre-trim TRIM_READS fastp · adapter/quality trim FASTQ_QC (trimmed) FastQC + MultiQC · post-trim BWA_ALIGN BWA-MEM2 → aligned BAM SORT_MARKDUP sort + mark duplicates (GATK) ① GATK HaplotypeCaller ② DeepVariant GATK_CALL HaplotypeCaller · 2,314 variants DEEPVARIANT_CALL WGS model · 2,906 variants CROSS_CHECK_VCFS bcftools isec · 2,267 concordant calls ③ Benchmark (hap.py) ④ Annotate (VEP + gnomAD) HAPPY_BENCHMARK vs. GIAB HG002 truth (NISTv4.2.1) ANNOTATE_CALLS VEP transcript + gnomAD v4.1 AF 1.0/1.0 recall/precision SNP + INDEL, both callers vs. GIAB 2,267 annotated calls concordant set · VEP + gnomAD → CAPN3-DMD-variant-classifier annotated VCF → VariantEvidenceBundle Analysis step Key output Feeds classifier
CAPN3 / DMD / BRCA1
ACMG/AMP Variant Classifier
GitHub

A scoped ACMG/AMP (Richards et al. 2015) variant classification engine, now spanning three genes and two inheritance mechanisms: CAPN3 (limb-girdle muscular dystrophy, autosomal recessive) and DMD (Duchenne/Becker muscular dystrophy, X-linked) form the original scope, extended to BRCA1 (hereditary breast/ovarian cancer syndrome, autosomal dominant with incomplete penetrance) as a deliberate generalization test against the real ClinGen ENIGMA BRCA1/BRCA2 VCEP v1.2 specification. Applies twelve ACMG criteria across population frequency, protein effect, computational evidence, functional-assay data, and structural variation — including a dedicated CNV deletion/duplication scoring module for DMD's largely exon-deletion mutational spectrum — combined by two independent systems (classic Table 5 and Tavtigian et al. 2020's Bayesian points) cross-checked against each other, plus a case-level layer reasoning about whether a patient's variant(s) actually explain their disease: biallelic X-linked (XX) inheritance scenarios for DMD, and a new "risk-conferring" case status reflecting BRCA1's dominant, incompletely penetrant mechanism rather than reusing the deterministic status the other two genes use.

35 point-mutation + 5 CNV + 14 case-level fixtures
259/259 automated tests passing
Three genes, two inheritance mechanisms (CAPN3 AR · DMD X-linked · BRCA1 AD)
Twelve ACMG/AMP criteria evaluators
Two combining systems (Table 5 + Bayesian, cross-checked)
CNV deletion + duplication scoring module (DMD)
Case-level reasoning (AR + X-linked incl. biallelic XX + BRCA1 risk-conferring)
Fed directly by the variant-calling pipeline's annotated output
Python dataclasses ACMG/AMP CNV PyYAML stdlib-only
View Workflow Diagram
CAPN3-DMD-variant-classifier Workflow ACMG/AMP evidence bundle evaluated against twelve criteria, with a separate CNV deletion/duplication module branching off in parallel, combined by two independent systems, cross-checked, then reasoned about at the case level. CAPN3-DMD-variant-classifier 35 point + 5 CNV + 14 case-level fixtures · 259/259 tests · twelve ACMG/AMP criteria · 3 genes VariantEvidenceBundle curated evidence, or fed by the calling pipeline Twelve ACMG/AMP Criteria Evaluators population frequency · protein effect · computational scores ① Classic (Table 5) ② Bayesian Points ③ CNV (DMD structural variants) TABLE5_COMBINE Richards et al. 2015 BAYESIAN_COMBINE Tavtigian et al. 2020 CNV_SCORING Riggs et al. 2020 Cross-checked classification Pathogenic → Benign · 2 systems agree CNV classification VUS → Pathogenic · Section 2 (dosage) Case-Level Reasoning AR (CAPN3) · X-linked (DMD) · AD risk-conferring (BRCA1) Clinical Interpretation variant + case-level · verified vs. 54 curated fixtures Analysis step Key output Final interpretation
TPMT / DPYD / SLCO1B1 / CYP2C19 / NUDT15
PGx Interpretation Pipeline
GitHub

A reproducible pharmacogenomics interpretation engine translating selected genomic variants into gene-specific allele/diplotype calls, functional phenotypes, and guideline-linked drug recommendations — built around explicit uncertainty rather than resolving genuinely ambiguous phasing by assumption. Five genes were chosen deliberately because they don't share one interpretation model: linked-variant dosage tables, activity-score summation, transporter biology, and direct multi-locus resolution. Deliberately scopes out CYP2D6 (structural-variant calling, unsolved even by specialist tools) rather than quietly shipping an incomplete answer. A standalone companion to the ACMG/AMP variant classifier below — same engineering discipline, applied to drug response instead of disease causation.

5 genes implemented end-to-end
233 automated tests, two independent runners
38 real GeT-RM (CDC) reference samples validated
5 report formats (JSON/TSV/HTML/MD/DOCX)
7 drug recommendations across 5 genes (Tier 2 evidence)
5-state explicit uncertainty model
CLI + Nextflow DSL2 batch entry points
Python stdlib-only dataclasses Nextflow DSL2 PharmVar CPIC/ClinPGx python-docx
View Workflow Diagram
PGx Interpretation Pipeline Workflow VCF input normalized, then five gene-specific diplotype/phenotype modules run in parallel, converging through a schema validation gate and versioned Tier 2 evidence lookup into a shared report-assembly layer, exposed via both a CLI and a Nextflow DSL2 batch pipeline. PGx Interpretation Pipeline 5 genes · 233 tests, two runners · 38 real GeT-RM samples validated · stdlib-only Python VCF (selected pharmacogene loci) TPMT · DPYD · SLCO1B1 · CYP2C19 · NUDT15 regions Variant Extraction & Normalization normalize.py · phased/unphased GT · multi-allelic ALT TPMT Linked 2-variant dosage table Phase ambiguity reported explicitly, not guessed DPYD Activity-score summation, 4 independent loci Different phenotype model SLCO1B1 Direct diplotype lookup Transporter, not an enzyme — zero schema change required CYP2C19 3 independent loci, resolved directly (not declined) Evidence-backed judgment call NUDT15 Single locus, no phasing Simplest case; first joint two-gene recommendation schema.validate() — Structural Uncertainty Gate Enforces 5 confidence states: SUPPORTED · AMBIGUOUS · INSUFFICIENT_DATA · UNSUPPORTED_ALLELE · UNRESOLVED Tier 2 Evidence Lookup — ClinPGx/CPIC (cached, versioned) 7 drug pairings across 5 genes · compound TPMT+NUDT15 · CYP2C19 dual-drug Report Assembly JSON · TSV · HTML · Markdown · DOCX — one shared, tested intermediate representation CLI — cli.py report single-sample entry point Nextflow DSL2 — main.nf fans CLI across a samplesheet Analysis step Uncertainty gate / key output Orchestration entry point
COLO829 / COLO829BL
Somatic Genomics Pipeline
GitHub

Tumour-normal somatic variant calling pipeline built around GATK Mutect2, benchmarked against a real published truth set for the COLO829/COLO829BL melanoma reference cell line pair. Reads are QC'd and aligned before contamination estimation and an interval-sharded Mutect2 call set — rebuilt as a genome-scale scatter/gather architecture after an unsharded run hit a real memory ceiling — is benchmarked with som.py and extended with CNVkit copy-number calling. A final module annotates PASS calls with SnpEff and queries CIViC's live clinical-evidence API, carrying a called variant through to real oncology evidence rather than stopping at a filtered VCF. Run against real ENA-extracted sequencing reads, the pipeline independently recovered three of COLO829's own documented driver mutations — BRAF V600E, a CDKN2A frameshift, and a UV-signature TERT promoter mutation — with zero false positives against the benchmark truth set.

8/8 PASS calls confirmed true positive (zero false positives, NYGC truth set)
3 real driver mutations independently matched: BRAF V600E · CDKN2A · TERT
132 real CIViC evidence rows retrieved live for BRAF V600E
Interval-sharded Mutect2 — genome-scale scatter/gather rebuilt after a real OOM
Real ENA-extracted tumour/normal reads (not synthetic), HiSeq X Ten WGS
som.py benchmarking + CNVkit copy-number calling alongside variant calling
Two Nextflow profiles: 8-gene melanoma driver panel + full genome
Nextflow DSL2 GATK Mutect2 CNVkit SnpEff CIViC som.py BWA-MEM2 Docker Python
View Workflow Diagram
COLO829 Somatic Genomics Pipeline Workflow Tumour-normal reads through QC, alignment and dedup, then contamination estimation and interval-sharded Mutect2 calling to a PASS-call gate, forking into copy-number calling, truth-set benchmarking, and SnpEff/CIViC clinical-evidence annotation, converging on the real-data milestone. COLO829 Somatic Genomics Pipeline BRAF · CDKN2A · TERT drivers confirmed on real data · Nextflow DSL2 · Docker Real tumour + normal FASTQ (ENA) COLO829 / COLO829BL · HiSeq X Ten WGS, gene-panel extracted FASTQC_MULTIQC raw QC — tumour + normal BWA_MEM2_ALIGN + SORT_MARKDUP dedup → analysis-ready BAMs GET_PILEUP_SUMMARIES + CALCULATE_CONTAMINATION cross-sample contamination estimate SPLIT_INTERVALS → N × MUTECT2 → MERGE interval scatter/gather, rebuilt after single-call OOM FILTER_MUTECT_CALLS — 8 PASS calls of 276 candidate records · gene-panel scope (dev profile) ① Copy-Number Calling ② Truth-Set Benchmarking ③ Oncology-Evidence Annotation CNVKIT_BATCH → CNVKIT_CALL tumour vs. normal coverage ratio Genome-wide CN segments panel-only coverage — disclosed scope limit SOMPY_BENCHMARK vs. NYGC published truth set 8/8 PASS calls = true positive 0 false positives · precision 1.0, all classes SNPEFF_ANNOTATE → CIVIC_ANNOTATE protein HGVS.p → live CIViC query 132 CIViC evidence rows BRAF V600E · real curated clinical evidence Real-data milestone BRAF V600E · CDKN2A · TERT — cross-checked against Cellosaurus & OncoKB Analysis step Key output Real-data milestone

Technical Stack

Pipeline & Workflow
Nextflow DSL2 bash Git
Containers
Docker Singularity
AMR & Typing
AMRFinderPlus MLST Kleborate emmtyper
Phylogenomics
Parsnp Snippy Gubbins IQ-TREE QUAST MAFFT
Clinical Genomics
FastQC fastp BWA-MEM2 GATK DeepVariant VEP gnomAD hap.py ACMG/AMP GIAB benchmarking PharmVar CPIC/ClinPGx GeT-RM benchmarking
Cancer / Somatic Genomics
GATK Mutect2 CNVkit SnpEff CIViC som.py
Data & Visualisation
Python Dash pandas Biopython
Regulatory Genomics
Inclusivity analysis Exclusivity analysis FDA 510(k) CE-IVDR Primer3 Panaroo NCBI BLAST fastANI

Get in Touch

If you'd like to discuss any of the projects featured here, exchange ideas, or connect regarding computational biology, molecular diagnostics or bioinformatics, I'd be pleased to hear from you.