Course Overview
The Variant Calling and Structural Genomics Pipeline Design Training Course for Bioinformaticians is an intensive program engineered to equip computational biologists, genomic researchers, and bioinformaticians with advanced skills in high-throughput sequencing data analysis. High-throughput next-generation sequencing (NGS) and third-generation long-read technologies generate vast quantities of genomic data requiring sophisticated computational pipelines to identify single nucleotide variants (SNVs), insertions/deletions (indels), and complex structural variants (SVs). This course addresses the critical demand for robust, reproducible, and scalable bioinformatics workflows capable of translating raw sequencing reads into actionable biological insights.
Throughout this comprehensive training, participants will dive deep into end-to-end genomic pipeline engineering, quality control, alignment algorithms, state-of-the-art variant calling tools, and structural variation discovery. The curriculum covers foundational and cutting-edge topics including Linux shell scripting for high-performance computing (HPC), GATK Best Practices workflows, long-read structural genomic analysis with PacBio and Oxford Nanopore data, annotation of single nucleotide polymorphisms (SNPs), copy number variation (CNV) detection, and automated workflow management using Nextflow and Snakemake. Participants will engage in extensive hands-on sessions designed to reinforce theoretical concepts with real-world genomic datasets.
Upon the successful completion of this Variant Calling and Structural Genomics Pipeline Design Training Course for Bioinformaticians, participants will be able to:
ü Design, construct, and execute end-to-end automated genomic pipelines for variant discovery and structural genomics.
ü Master quality control, read trimming, and high-accuracy reference alignment techniques for both short-read and long-read NGS data.
ü Implement GATK Best Practices for germline and somatic variant calling to accurately identify SNVs and small indels.
ü Detect, characterize, and visualize complex structural variants, copy number variations (CNVs), and genomic rearrangements using specialized bioinformatics algorithms.
ü Annotate identified variants using functional databases and assess their biological, clinical, or evolutionary significance.
ü Containerize genomic workflows using Docker and Singularity/Apptainer while deploying scalable pipelines with Nextflow and Snakemake on HPC clusters.
Training Methodology
The course is designed to be highly interactive, challenging and stimulating. It will be an instructor led training and will be delivered using a blended learning approach comprising of:
ü Interactive lectures delivered by senior bioinformatics pipeline engineers and genomic researchers.
ü Hands-on computational laboratory sessions utilizing high-performance Linux cloud instances.
ü Real-world case studies analyzing open-source human, plant, and microbial genomic datasets.
ü Guided troubleshooting sessions for debugging complex shell scripts and pipeline workflows.
ü Peer collaboration and code reviews using Git repositories to build production-grade code.
ü Interactive Q&A sessions ensuring tailored guidance for individual research projects and organizational workflows.
Our facilitators are seasoned industry professionals with years of expertise in their chosen fields. All facilitation and course materials will be offered in English.
Who Should Attend?
This Variant Calling and Structural Genomics Pipeline Design Training Course for Bioinformaticians would be suitable for, but not limited to:
ü Bioinformaticians and Computational Biologists seeking to automate and scale genomic analysis workflows
ü Genomic Researchers and Molecular Biologists aiming to gain hands-on computational expertise in NGS variant calling
ü Biostatisticians and Data Scientists working with high-throughput genomic and life sciences data
ü Core Facility Managers and Technical Leads managing NGS data infrastructure and analytical pipelines
ü Clinical Genomics Analysts involved in medical diagnostics, oncogenomics, and hereditary disease profiling
ü Software Engineers and Systems Administrators supporting bioinformatics cloud and HPC infrastructures
Benefits of the Training
Personal Benefits
ü Master in-demand skill sets in high-throughput genomic data processing, pipeline design, and long-read analysis.
ü Gain direct practical experience with industry-standard bioinformatic tools including GATK, SAMtools, BWA, Minimap2, and Snakemake.
ü Enhance career mobility and professional recognition as a proficient genomic pipeline engineer.
ü Develop the ability to independently troubleshoot complex computational errors in genomic workflows.
ü Build a robust portfolio of containerized, automated bioinformatics pipelines hosted on GitHub.
Organizational Benefits
ü Accelerate research and development timelines by deploying automated, high-throughput variant calling pipelines.
ü Ensure rigorous reproducibility, compliance, and standardized quality control across all internal genomic projects.
ü Reduce cloud computing and infrastructure costs through optimized, scalable pipeline architecture and resource allocation.
ü Empower internal teams to independently handle complex long-read and structural variant genomics without relying on external vendors.
ü Enhance overall research quality and diagnostic accuracy through state-of-the-art variant annotation and filtering methods.
ü Course Duration: 5 Days
ü Training Fee:
o Physical Training: USD 1,500
o Online / Virtual Training: USD 1,000
Module 1: Foundations of NGS Data, Quality Control, and Read Preprocessing
ü Overview of Short-Read (Illumina) and Long-Read (PacBio, Oxford Nanopore) Sequencing Technologies
ü Understanding Genomic File Formats: FASTQ, FASTA, and Quality Score Systems (Phred)
ü Raw Read Quality Control and Metrics Generation using FastQC and MultiQC
ü Adapter Trimming, Low-Quality Base Filtering, and Contamination Removal using Trimmomatic and Fastp
ü Practical Session: Executing quality control workflows on raw paired-end FASTQ data and generating unified MultiQC HTML reports
Module 2: High-Accuracy Read Alignment and Post-Alignment Processing
ü Reference Genome Architecture, Indexing, and Genome Build Comparisons (GRCh38 vs. CHM13 T2T)
ü Short-Read Alignment Algorithms: Burrows-Wheeler Transform, BWA-MEM, and Bowtie2
ü Long-Read Alignment Algorithms: Minimap2 and NGMLR
ü BAM/SAM File Manipulation, Sorting, Indexing, and Coverage Calculation with SAMtools and Sambamba
ü Marking PCR Duplicates, Base Quality Score Recalibration (BQSR), and Alignment QC with Picard Tools
ü Practical Session: Aligning raw reads to the human reference genome, sorting BAM files, marking duplicate reads, and computing depth of coverage
Module 3: Small Variant Calling Frameworks (SNVs and Indels)
ü Statistical Foundations of Variant Calling: Bayesian and Maximum Likelihood Models
ü GATK Best Practices Workflow for Germline Short Variant Discovery (HaplotypeCaller)
ü Joint Genotyping Strategies for Population-Scale Cohorts (GenotypeGVCFs)
ü Alternative Variant Callers: FreeBayes and BCFTools Call Pipeline
ü Hard Filtering vs. Variant Quality Score Recalibration (VQSR) for False-Positive Reduction
ü Practical Session: Implementing the GATK HaplotypeCaller workflow to identify germline SNVs and indels from aligned BAM files
Module 4: Somatic Variant Calling and Cancer Genomics
ü Distinguishing Somatic Mutations from Germline Variants: Paired Tumor-Normal Analysis Frameworks
ü Handling Tumor Purity, Subclonal Heterogeneity, and Copy Number Alterations in Cancer
ü Somatic SNV and Indel Discovery using Mutect2, Strelka2, and VarScan2
ü Filtering Artifacts using Panel of Normals (PoN) and Population Allele Frequencies
ü Visualizing Somatic Variants, Mutational Signatures, and Oncogenic Drivers
ü Practical Session: Running Mutect2 on paired tumor-normal BAM files, building a custom Panel of Normals, and applying artifact filtering
Module 5: Structural Variation (SV) Detection in Short-Read Data
ü Principles of Structural Variation: Deletions, Insertions, Inversions, Duplications, and Translocations
ü Short-Read SV Detection Signals: Split-Reads, Discordant Read-Pairs, and Read-Depth Alterations
ü Benchmarking Short-Read SV Callers: Manta, LUMPY, and DELLY
ü Integrating Ensemble SV Callers for Higher Precision and Sensitivity
ü Filtering and Genotyping Structural Variants using VCF tools and SURVIVOR
ü Practical Session: Detecting and genotyping structural variations in short-read data using Manta and evaluating call sets with SURVIVOR
Module 6: Long-Read Genomics for Structural Variant Discovery
ü Advantages of Long-Read Sequencing (PacBio HiFi, Oxford Nanopore) for Complex Genomic Regions
ü Long-Read Alignment Considerations and Base-Call Error Profiles
ü Long-Read SV Calling Frameworks: Sniffles2, CuteSV, and PBSV
ü Characterizing Repetitive Elements, Transposons, and Complex Structural Rearrangements
ü De Novo Genome Assembly Frameworks for Haplotype-Resolved SV Discovery
ü Practical Session: Processing Oxford Nanopore and PacBio HiFi data using Minimap2 and Sniffles2 to detect complex structural variants
Module 7: Copy Number Variation (CNV) Analysis and Detection
ü Principles of Copy Number Variation: Germline Copy Number Differences and Somatic Copy Number Alterations
ü Read-Depth Analysis for CNV Profiling: Normalization Methods and GC-Content Correction
ü CNV Calling Tools: CNVkit, GATK gCNV, and Control-FREEC
ü Integrating B-Allele Frequency (BAF) and Log2 Ratios for Allele-Specific CNV Profiling
ü Visualizing CNV Profiles Across Whole Genomes and Specific Chromosomal Arms
ü Practical Session: Executing CNVkit on target enrichment/WES data to profile copy number gains and losses and generating diagnostic scatter plots
Module 8: Variant Annotation, Filtering, and Functional Effect Prediction
ü Standardization of Variant Call Format (VCF) and Data Manipulation with BCFTools
ü Functional Variant Annotation Frameworks: SnpEff, VEP (Variant Effect Predictor), and ANNOVAR
ü Integrating Population Allele Frequencies (gnomAD, 1000 Genomes) and Pathogenicity Scores (CADD, REVEL)
ü Prioritizing Rare, Loss-of-Function, and Clinical Variants according to ACMG/AMP Guidelines
ü Querying and Filtering VCF Files using Slivar, SnpSift, and BCFTools
ü Practical Session: Annotating a multi-sample VCF using Ensembl VEP, applying custom Slivar filters, and prioritizing pathogenic candidate variants
Module 9: Automated Workflow Management with Nextflow and Smakemake
ü Principles of Declarative and Dataflow Programming for Computational Biology
ü Building Reproducible Pipelines with Snakemake: Rules, Inputs, Outputs, and Wildcards
ü Constructing Scalable Workflows with Nextflow: Processes, Channels, and Operators
ü Managing Pipeline Configuration, Execution Profiles, and Modular Architecture (nf-core)
ü Handling Error Recovery, Resume Mechanisms, and Dynamic Resource Allocation
ü Practical Session: Writing a modular Nextflow script to automate read alignment, BAM processing, and variant calling end-to-end
Module 10: Containerization, HPC Deployment, and Pipeline Reproducibility
ü Ensuring Software Reproducibility in Bioinformatics: Bioconda, Docker, and Singularity/Apptainer
ü Creating Custom Dockerfiles and Singularity Images for Genomic Toolkits
ü Integrating Containers with Workflow Managers (Nextflow/Snakemake)
ü Deploying Pipelines on HPC Schedulers (SLURM, PBS) and Cloud Infrastructures (AWS Batch, Google Cloud Life Sciences)
ü Best Practices for Open-Science Code Documentation, Version Control with Git, and GitHub Deployment
ü Practical Session: Containerizing a variant calling pipeline with Singularity and deploying it on a SLURM-managed HPC cluster using Nextflow
About Our Trainers
Our trainers are senior bioinformaticians, computational genomicists, and software engineers with over 15 years of practical experience in high-throughput sequencing data analysis, pipeline engineering, and academic/industrial R&D. They have led large-scale genomic projects across clinical diagnostics, cancer research, agricultural biotechnology, and evolutionary genomics. Having authored numerous peer-reviewed publications in high-impact scientific journals and contributed to open-source bioinformatics tools, our instruction team brings unmatched technical depth and practical insights directly into the classroom.
Quality Statement
Phoenix Center for Policy, Research and Training is committed to delivering world-class, practical, and impact-driven professional development programs. We adhere to rigorous instructional standards, continuously updating our curriculum to reflect the latest advancements in next-generation sequencing, long-read technologies, and computational frameworks. Our hands-on training environments, expert mentorship, and industry-aligned evaluation metrics ensure that every participant achieves actionable competency and immediate operational capability upon course completion.
Tailor-Made Courses
We understand that every organization has unique challenges and opportunities as well as unique training needs. Phoenix Training Center offers tailor-made courses designed to address specific requirements and challenges faced by your team or organization. Whether you need a customized curriculum, a specific duration, or on-site delivery, we can adapt our expertise to provide a training solution that perfectly aligns with your objectives. We can customize this Course to focus on your industry, specific risk profile, or internal stakeholder dynamics. Contact us to discuss how we can create a bespoke training program that maximizes value and impact for your team. For further inquiries, please contact us on Tel: +254720272325 / +254737296202 or Email: training@phoenixtrainingcenter.com.
Admission Criteria
ü Participants should be reasonably proficient in English.
ü Applicants must live up to Phoenix Center for Policy, Research and Training admission criteria.
Terms and Conditions
ü Discounts: Organizations sponsoring Four Participants will have the 5th attend Free
ü What is catered for by the Course Fees: Fees cater for all requirements for the training – Learning materials, Lunches, Teas, Snacks and Certification. All participants will additionally cater for their travel and accommodation expenses, visa application, insurance, and other personal expenses.
ü Certificate Awarded: Participants are awarded Certificates of Participation at the end of the training.
ü Course Improvement: The program content shown here is for guidance purposes only. Our continuous course improvement process may lead to changes in topics and course structure.
ü Approval of Course: Our Programs are NITA Approved. Participating organizations can therefore claim reimbursement on fees paid in accordance with NITA Rules.
Booking for Training
Kindly send an email to the Training Officer on training@phoenixtrainingcenter.com and we will send you a registration form. We advise you to book early to avoid missing a seat to this training. Or call us on +254720272325 / +254737296202
Payment Options
We provide 3 payment options, choose one for your convenience, and kindly make payments a week before the training starts (at least 5 to 7 days before the Training start date) to reserve your seat:
ü Groups of 5 People and Above – Cheque Payments to: Phoenix Center for Policy, Research and Training Limited should be paid in advance, a week before the training starts.
ü Invoice: We can send a bill directly to you or your company.
ü Deposit directly into Bank Account (Account details provided upon request)
Cancellation Policy
ü Payment for all courses includes a registration fee, which is non-refundable, and equals 15% of the total sum of the course fee.
ü Participants may cancel attendance 14 days or more prior to the training commencement date.
ü No refunds will be made 14 days or less before the training commencement date. However, participants who are unable to attend may opt to attend a similar training course at a later date or send a substitute participant provided the participation criteria have been met.
Accommodation and Airport Pick-up
For physical training attendees, we can assist with recommendations for accommodation near the training venue. Airport pick-up services can also be arranged upon request to ensure a smooth arrival. Please inform us of your travel details in advance if you require these services. For reservations contact the Training Officer on Email: training@phoenixtrainingcenter.com or on Tel: +254720272325 / +254737296202.
| Course Dates | Venue | Fees | Enroll |
|---|
Phoenix Training Center
Typically replies in minutes