Experience in molecular diagnostics, assay development, and computational biology. This portfolio brings together selected projects exploring diagnostic assay design, bioinformatics, AI-assisted scientific workflows, and computational methods for molecular biology.
I am a molecular biologist with a microbiology background and industry experience in molecular diagnostics, assay development and biotechnology.
Alongside laboratory research and product development, I have increasingly worked in computational biology, developing reproducible bioinformatics workflows for microbial genomics and diagnostic applications. Recent work extends this to human clinical genomics: variant calling, ACMG/AMP classification, pharmacogenomics and somatic tumour–normal analysis.
This website brings together selected projects, case studies and technical notes from both areas.
End-to-end Nextflow DSL2 pipeline characterising antimicrobial resistance across 283 clinical strains. Integrates AMRFinderPlus resistance profiling with a custom alignment-based mutation caller (penA, gyrA, parC, mtrR, porB) and promoter-variant detection, core-genome phylogeny via Parsnp, and an interactive Dash dashboard linking resistance phenotypes to phylogenetic context for rapid clinical interpretation.
Scalable pipeline processing 375 publicly available S. pyogenes genomes with integrated resistance profiling, MLST, and emm typing. Identified 22.9% AMR prevalence across the cohort, with ST28 as the dominant sequence type (22.4%) and emm5 the most prevalent emm type (47.5%). Designed for rapid integration of newly sequenced clinical isolates into ongoing surveillance.
Reproducible target-selection workflow for a multiplex qPCR/TaqMan assay for C. trachomatis. Genome QC and tiering are performed before target discovery, ensuring that consensus sequences and inclusivity assessments are built only from high-confidence assemblies. A provenance audit traced 149 GenBank-only genomes to a single study of engineered laboratory strains; these were excluded from the diagnostic reference set, which was rebuilt from 125 complete RefSeq assemblies that were independently verified and tiered.
Four chromosomal loci (mutL, aroB, PSD2, and group_306) were independently evaluated as candidate targets alongside the multicopy diagnostic plasmid, which serves as the higher-sensitivity target. The proposed duplex pairs one chromosomal locus with the plasmid target, combining genomic redundancy with sensitivity.
The workflow integrates pan-genome analysis, sequence-conservation and inclusivity screening, exclusivity BLAST, quantitative Tm-gap assessment, and three-way structural QC of primer–probe sets, producing a computationally screened candidate panel for downstream experimental validation.
Comprehensive phylogenomics pipeline spanning the full analysis chain for 50 K. pneumoniae genomes. Assembly QC (QUAST), AMR and resistance gene profiling (AMRFinderPlus), virulence and typing (Kleborate), SNP-based variant calling (Snippy), recombination masking (Gubbins), and maximum-likelihood phylogeny (IQ-TREE). Results delivered via a six-tab Dash dashboard. ST11 identified as the dominant lineage, consistent with global carbapenem-resistant K. pneumoniae trends.
Reads-to-VCF pipeline scoped to the CAPN3 and DMD loci, independently benchmarked against Genome in a Bottle's HG002 reference sample. Raw reads are QC'd (FastQC/MultiQC) and adapter/quality-trimmed (fastp) before alignment. Two variant callers — GATK HaplotypeCaller and DeepVariant — are run independently and cross-checked against each other, then the concordant call set is annotated with Ensembl VEP transcript consequence and gnomAD population frequency, producing evidence ready for the downstream ACMG classifier.
A scoped ACMG/AMP (Richards et al. 2015) variant classification engine, now spanning three genes and two inheritance mechanisms: CAPN3 (limb-girdle muscular dystrophy, autosomal recessive) and DMD (Duchenne/Becker muscular dystrophy, X-linked) form the original scope, extended to BRCA1 (hereditary breast/ovarian cancer syndrome, autosomal dominant with incomplete penetrance) as a deliberate generalization test against the real ClinGen ENIGMA BRCA1/BRCA2 VCEP v1.2 specification. Applies twelve ACMG criteria across population frequency, protein effect, computational evidence, functional-assay data, and structural variation — including a dedicated CNV deletion/duplication scoring module for DMD's largely exon-deletion mutational spectrum — combined by two independent systems (classic Table 5 and Tavtigian et al. 2020's Bayesian points) cross-checked against each other, plus a case-level layer reasoning about whether a patient's variant(s) actually explain their disease: biallelic X-linked (XX) inheritance scenarios for DMD, and a new "risk-conferring" case status reflecting BRCA1's dominant, incompletely penetrant mechanism rather than reusing the deterministic status the other two genes use.
A reproducible pharmacogenomics interpretation engine translating selected genomic variants into gene-specific allele/diplotype calls, functional phenotypes, and guideline-linked drug recommendations — built around explicit uncertainty rather than resolving genuinely ambiguous phasing by assumption. Five genes were chosen deliberately because they don't share one interpretation model: linked-variant dosage tables, activity-score summation, transporter biology, and direct multi-locus resolution. Deliberately scopes out CYP2D6 (structural-variant calling, unsolved even by specialist tools) rather than quietly shipping an incomplete answer. A standalone companion to the ACMG/AMP variant classifier below — same engineering discipline, applied to drug response instead of disease causation.
Tumour-normal somatic variant calling pipeline built around GATK Mutect2, benchmarked against a real published truth set for the COLO829/COLO829BL melanoma reference cell line pair. Reads are QC'd and aligned before contamination estimation and an interval-sharded Mutect2 call set — rebuilt as a genome-scale scatter/gather architecture after an unsharded run hit a real memory ceiling — is benchmarked with som.py and extended with CNVkit copy-number calling. A final module annotates PASS calls with SnpEff and queries CIViC's live clinical-evidence API, carrying a called variant through to real oncology evidence rather than stopping at a filtered VCF. Run against real ENA-extracted sequencing reads, the pipeline independently recovered three of COLO829's own documented driver mutations — BRAF V600E, a CDKN2A frameshift, and a UV-signature TERT promoter mutation — with zero false positives against the benchmark truth set.
If you'd like to discuss any of the projects featured here, exchange ideas, or connect regarding computational biology, molecular diagnostics or bioinformatics, I'd be pleased to hear from you.