FASTQ QC: Raw vs Trimmed — what fastp actually removed, and what changed downstream
FASTQ_QC ran twice: once on EXTRACT_REGION's raw output, once on
TRIM_READS's (fastp) output — same FastQC + MultiQC modules, same regions, so the
before/after is a direct comparison. The question trimming has to answer isn't "did quality go up"
(it should, trivially) but "did it throw away real signal along with the noise" — the alignment and
benchmark numbers below confirm it didn't.
Read yield and quality — before vs after fastp
96.65% of reads survived trimming — 673,098 in, 650,510 out. Of the 22,588 removed,
22,424 were low-quality reads and 164 had too many ambiguous (N) bases; none were dropped for being
too short, too long, or pure adapter dimer. Separately, 2,188 reads (0.33%) had a genuine Illumina
TruSeq adapter fragment trimmed off rather than the whole read discarded. Q20 rose from 96.81% to
97.99% and Q30 from 91.83% to 93.54% — the filtering targeted the low-quality tail, not a random slice
of the data.
FastQC module outcomes — before vs after
Both runs: 0% module failure rate across FastQC's full check suite. Trimming didn't introduce any new hard failures — only one new WARN.
"Sequence Length Distribution" WARN appears only post-trim — expected and benign:
fastp trims variable amounts per read rather than a fixed number of bases, so post-trim reads no
longer have a single uniform length, which is exactly what that FastQC check flags. It is not a
quality signal. "Per sequence GC content" WARNs in both runs for the same reason it did before
trimming — a scoped two-locus extraction is not expected to match FastQC's whole-genome background
GC model.
Downstream, nothing regressed. Post-trim: BWA_ALIGN mapping rose to 99.997% (properly
paired 99.82%), SORT_MARKDUP duplication rate dropped slightly to 0.92%, and HAPPY_BENCHMARK still
scored a clean 1.0 / 1.0 recall and precision for both callers and both variant types against the
GIAB HG002 truth set (612 truth variants, unchanged) — see the Benchmark Scorecard report.