MorPhiC experimental and analytical methods span null allele generation, molecular phenotype analysis, and scalable processing pipelines developed across consortium centers.
This page collects the consortium's current methods, protocols, and computational approaches in the same anchor structure already used across the site.
Experimental strategies for null allele generation
A proven strategy for understanding gene functions is to remove the gene or its protein product and perform a set of assays to study how its loss affects molecular and cellular phenotypes. The Data Production Centers have employed the following strategies for null allele generation:
1. Gene knock-out (KO)
KO - Premature termination codon (PTC)
Strategy used by: The Jackson Laboratory, view on protocols.io
The PTC+1 strategy involves precise CRISPR-Cas9 engineering of a premature stop codon and the insertion of a degenerate base in an early common exon, thereby truncating all isoforms of the protein and introducing a frame-shift mutation to ensure the production of a non-functional protein, essentially "knocking out" the normal function of the gene. Depending on the position of the PTC+1 mutation with respect to the genomic structure of the gene, the mutant mRNA may or may not be subject to nonsense mediated decay (NMD).
Strategy used by: Memorial Sloan Kettering Cancer Center
Frameshift indel mutations introduced by CRISPR-Cas9 typically result in the creation of a premature stop codon within the coding region, producing a truncated, nonfunctional protein and effectively "knocking out" the gene’s normal function. At MSK, they have generated knockout hPSC lines, and a small number of enhancer deletion lines. Then they individually barcoded these lines, and pooled them for differentiation experiments. Cells were collected at consecutive differentiation stages for scRNA-seq, followed by computational demultiplexing to analyze the transcriptomic phenotypes associated with each individual knockout condition.
KO - Critical exon deletion
Strategy used by: The Jackson Laboratory, view on protocols.io
A critical exon deletion refers to a genetic knockout (KO) where an early common, frame-shifting exon within a gene has been deleted, leading to a truncated, nonfunctional protein being produced; essentially "knocking out" the normal function of the gene. Depending on the position of the premature termination codon with respect to the genomic structure of the gene, the mutant mRNA may or may not be subject to nonsense mediated decay (NMD).
KO - Gene deletion
Strategy used by: The Jackson Laboratory, view on protocols.io
Gene deletion means the complete removal of most or all protein coding exons of the gene, eliminating the possibility of the protein production entirely. The non-coding RNA transcript that remains is not subject to nonsense mediated decay and can potentially be identified by RNA-seq.
KO - Insertion of a cassette (a KI-KO approach)
Strategy used by: Memorial Sloan Kettering Cancer Center
A cassette is knocked into the target gene coding region. This insertion is expected to disrupt transcription due to the large insertion size, introduce frameshift mutations, and create a premature termination codon, thus effectively "knocking out" the normal function of the gene.
2. Protein degradation
Auxin-inducible degron
Strategy used by: Northwestern University
Auxin-inducible degron (AID) is a chemical genetic tool that uses the plant hormone auxin to degrade specific proteins in mammalian cells, which allows researchers to study protein function in living tissues.
At NWU, cells are engineered to express the TIR1 auxin receptor and Crispr-Cas9 technology is used to add the AID protein degron tag to specific endogenous gene loci. Cells express the tagged protein normally until treated with Auxin, which induces rapid, reversible protein degradation.
3. Gene knockdown
CRISPR interference (CRISPRi)
Strategy used by: University of California San Francisco
CRISPRi (CRISPR interference) utilizes a catalytically dead Cas9 (dCas9) protein fused to a KRAB (Krüppel-associated box) domain, a transcriptional repressor. A guide RNA (gRNA) directs dCas9-KRAB to a specific DNA sequence, where it binds and blocks transcription by recruiting endogenous chromatin-modifying repressive complexes via the KRAB domain. This approach enables reversible, precise, and non-destructive repression of gene expression at the transcriptional level, making it ideal for functional genomics and studying essential genes.
Analysis Methods
Bulk RNAseq Differential Expression Analysis
Strategy used by: Fred Hutch Cancer Center, view on GitHub
To quantify KO effect using Bulk RNA-seq data, we mainly focus on Differential Expression (DE) analysis comparing knock out (KO) versus wild type (WT) samples, using method DESeq2, followed by gene set enrichment analysis.
The Jackson Laboratory produced bulk RNAseq data with gene knockout and WT involving different gene knockout strategies, different model systems, different cell lines, and different days. For certain gene, the datasets involve two oxygen conditions.
Each dataset is processed separately. To assess the differential expression (DE) effects of gene knockout, the steps taken include quality control, DE analysis, and functional category enrichment analysis. For data visualization, volcano plots are provided to highlight the significantly differentially expressed genes under each gene knockout, and bar plots are provided to show significantly enriched functional categories. Detailed DE results are provided in tables, including basic information for each gene and the corresponding log2(fold change), p-value, and adjusted p-value.
The gene differential expression analysis method used is DESeq2. The major type of DE testing is between knockout samples and wild type samples. The second type of DE testing is between the two oxygen conditions. Proper levels of sample groups are chosen to run DESeq2 on, balancing between borrowing information across samples and avoiding batch effect.
When running DESeq2 model, there are two covariates considered: sequencing run ID and read depth (computed as the 75% quantile of gene expression in each sample). The sequencing run ID is included if the samples involve more than one sequencing run. To decide whether to include read depth, a model including read depth is fit first, and the proportion of genes for which the read depth can explain a significant amount of variation is estimated. The read depth is included in the final model if the estimated proportion is above a threshold that is adjusted according to sample size.
To make the results more reliable, multiple filtering steps are carried out based on the expression level of the genes. Genes with low expression are excluded from DE testing or have adjusted p-value marked as NA.
scRNA-seq Cell Type Composition
Pseudobulk-based approach
Strategy used by: Fred Hutch Cancer Center, view on GitHub
To quantify perturbation-induced changes in cell-type composition, pseudobulk samples were generated for each clone at each time point. Linear regression was applied to log-transformed cell-type proportions, with genotype as the primary predictor and parental cell line, differentiation process, and data source included as covariates.
Cell-level approach
Strategy used by: Fred Hutch Cancer Center
If the design is to have many clones with a few cells per clone, we cannot calculate cell type composition for each clone with reasonable accuracy. Then we conduct cell level analysis. For each cell type, we asked whether cells carrying a given knockout were more or less likely to become that cell type than wild-type cells, using logistic regression. while accounting for the dependency of cells from the same clone by random effects. We further adopted a Bayesian version to stabilize estimates for rare knockout/cell-type combinations that would otherwise produce extreme p-values by standard logistic regression with random effects.
scRNA-seq Differential Expression Analysis
Pseudobulk-based approach
Strategy used by: Fred Hutch Cancer Center, view on GitHub
To identify differentially expressed genes (DEGs) between a KO genotype of interest and WT, pseudobulk samples were generated for each clone–cell type combination at each time point. Differential expression analysis was performed using the DESeq2 package in R, with read depth and cell background as covariates.
Cell-level approach
Strategy used by: Fred Hutch Cancer Center, view on GitHub
Alternatively, for each time point and cell type, cell-level differential expression analysis was performed using rank-sum test (FindMarkers function) from the Seurat R package with default parameters. If there multiple clones with mulitple cells per clone, we employ a count-based model (nebula) that includes clone random effects to account dependency of the cells within one clone.
This workflow quantifies the effects of genetic perturbations on chromatin accessibility using bulk ATAC-seq. It constructs a consensus set of accessible chromatin regions, identifies differentially accessible (DA) peaks between perturbation and control conditions, and characterizes DA peaks through genomic annotation, gene set enrichment analysis, and transcription factor motif enrichment. The workflow produces standardized peak sets, differential accessibility results, quality control metrics, and downstream functional annotations for each dataset.
To implement this workflow, alignment files (e.g. BAMs) across replicates of the same condition are merged and converted to paired-end tagAlign format to prioritize Tn5 insertion sites. Peaks for each condition are called using MACS3 with the default fragment-shifting model disabled and ATAC-specific parameters applied instead. Peaks overlapping regions with known anomalously high signals (e.g. repetitive elements) are removed.
All condition-level peaks are combined into one consensus set by merging locations within 10bp of each other. We then construct a sample by peak read-count matrix, normalize to counts per million (CPM), and retain peaks with CPM > 1 in at least 2 samples from any condition. Peaks are annotated to the nearest genomic feature.
In addition to these variable-width consensus peaks, we also generate paired fixed-width coordinates to allow for other downstream analyses like motif enrichment. For each consensus peak, we identify all per-condition summit locations. If all summits are within 10bp of each other, the final summit is set as the mean position; otherwise, we compute the per-bp read coverage across all samples and identify the position with the highest overall coverage. Each summit is then extended +/-250bp to a 500bp peak.
Differential accessibility testing is performed with DESeq2, contrasting each knockout condition against the wildtype. Gene set enrichment analysis is performed using a gene-level ranking score computed as the sum of log2FoldChange x -log10(padj) across all peaks annotated to that gene. Finally, for conditions with a sufficient number of DA peaks, motif enrichment is performed, comparing the 500bp fixed-width sequences of DA peaks against a random sample of non-DA peaks.
Outputs for each dataset include the consensus peak set (fixed- and variable-width), the peak x sample count matrix, DA results tables for each condition comparison, GSEA tables, motif enrichment (when thresholds are met), and an automatically rendered QC report covering per-sample library metrics (% mitochondrial reads, fraction of fragments in nucleosome-free vs. mono-nucleosome size ranges) and DESeq2 diagnostics.
This workflow applies ChromBPNet to identify the DNA sequence features underlying chromatin accessibility and how their predicted regulatory contributions change across experimental conditions. By training standardized, condition-specific models and comparing model-derived attribution scores, the workflow identifies transcription factor motif instances and regulatory sequence features predicted to drive differential chromatin accessibility following perturbation.
To implement this workflow, separate ChromBPNet models are trained for each experimental condition by merging replicate datasets and using identical chromosomal train/test splits. First, a bias model is trained to account for Tn5 sequence preferences before a ChromBPNet accessibility model is trained using the frozen bias model. The resulting neural network is interpreted using DeepLIFT to generate base-resolution contribution scores, which are aggregated across transcription factor motif instances to estimate their contribution to predicted chromatin accessibility. These motif-level contribution scores are then compared between condition-specific models to identify transcription factors and regulatory sequence features whose predicted regulatory contributions change following perturbation.
Gene programs are groups of genes with coordinated expression patterns that reflect shared cellular processes, cell identities, or regulatory activities. This workflow uses cNMF, a consensus-based matrix factorization approach, to decompose single-cell RNA-seq data into interpretable gene programs. cNMF applies NMF repeatedly with different random initializations, clusters the resulting factors, and aggregates them into consensus programs. This reduces stochastic variability and increases reproducibility compared to single NMF runs. The output consists of a program usage matrix (cells × programs), which describes the activity of each program in each cell, and a gene loading matrix (programs × genes), which identifies the genes that define each program.
To ensure both robustness and biological interpretability, cNMF results are systematically evaluated using multiple criteria:
Reconstruction error and stability, measuring how well the decomposition explains the data and how reproducible the programs are.
Co-regulation and coherence, testing whether program genes are functionally linked.
Biological enrichment, assessing whether programs correspond to known pathways, processes, or cell identities.
Cross-dataset reproducibility, ensuring that programs are consistent across experimental replicates and biological contexts.
Through this strategy, cNMF reliably identifies both identity programs (cell-type signatures) and activity programs (dynamic patterns such as cell cycle, stress responses, or signaling).
Uniform and Scalable Data Processing Methods
Analytical pipelines
To facilitate reproducible analyses across diverse data types, graphical and interactive analytical pipelines are developed in the open-source Biodepot platform with a training portal available at https://biodepot.github.io/training/ . Recipes for using the STAR Suite are provided at https://github.com/morphic-bio/morphic-recipes consisting of structured blueprints exposed through our MCP server to agents and through Launchpad to human users, generating reproducible scripts and pipelines. These analytical pipelines are developed by the DRACC with feedback solicited from members in the MorPhiC consortium. Pipelines currently supported include:
Data type
GitHub
Description
Bulk RNA-seq, Single cell (sc) RNA-seq, Perturb-seq, 10x Flex, and SLAM-seq
STAR Suite updates the original STAR aligner by integrating four modules — STAR-perturb, STAR-Flex, STAR-SLAM, and TranscriptVB — to provide complete internal C/C++ pipelines for bulk RNA-seq, scRNA-seq, Perturb-seq, 10x Flex, and SLAM-seq. The integration results in substantial speedups and a simplified toolchain that can be installed through pre-compiled binaries for researchers and agents. This serves as a drop-in replacement for the STAR aligner. Preprint: bioRxiv 2026.06.02.729736
Chromap Suite is an open-source C++ chromatin-accessibility platform that performs ATAC-seq alignment and narrow peak caller. This Suite also includes an embeddable callable-library API and an MCP server with a browser Launchpad. Preprint: bioRxiv 2026.06.02.729736
A workflow that utilizes MAGeCK to quantify gRNA abundance from CRISPR gene sequencing data, combine per-sample counts into a unified count table, and run gene-level enrichment/depletion testing across screens. Extraction of minimum p-values and FDR values per screen from the MAGeCK test rankings and volcano plots are included as workflow outputs. A pre-configured instance is available at https://gitpod.io/#https://github.com/morphic-bio/gRNA-Enrichment
A workflow for processing LC-MS metabolomics data, built around the asari/PCPFM (Python-Centric Pipeline For Metabolomics) framework for feature table generation (feature detection) and annotation from compound and MS/MS libraries across multiple chromatography modes (e.g., HILIC-negative, RP-positive).
Due to consent restrictions for the KOLF2 cell line and derivatives that require data mapping to the Y chromosome to be removed, a filtering step has been added to our analytical pipelines for all data generated using these cell lines. Specifically, all reads in the BAM and FASTQ files that align to the Y-chromosome are removed, but the counts tables include the Y chromosome data.
Scalable data processing
A scheduler based on the Temporal.io framework has been developed to enable optimizations of bioinformatics workflows. Specifically, users can transparently map workflow steps to diverse execution environments, including high-performance computing (HPC) resources managed by the SLURM resource manager through an easy-to-use graphical user interface. Asynchronous execution of workflows is supported to optimize resource utilization even when the scheduler cannot make use of a system’s full RAM and CPU resources. Pipelines are executed using a combination of UW compute resources and the NSF Bridges2 supercomputer.