help

Frequently asked questions

Frequently asked questions about the MorPhiC program, its goals, participants, and the science behind it.

General

MorPhiC stands for Molecular Phenotypes of Null Alleles in Cells. MorPhiC is a collaborative research project that aims to functionally characterize all protein-coding human genes. The project focuses on developing a catalog of molecular and cellular phenotypes for null alleles for every protein-coding human gene. The ultimate goal is to provide a comprehensive understanding of the biological function of each human gene, filling the knowledge gap for the majority of genes that are currently underrepresented in scientific literature.

To prioritize gene targets, the MorPhiC consortium considers a set of genes involved in critical cellular and organismal functions, including essential genes, transcription factors, developmental regulators, and disease-associated genes. Full list of genes to be studied under MorPhic.

The MorPhiC consortium consists of four Data Production Centers (DPCs), one Data Resource and Administrative Coordinating Center (DRACC), and three Data Analysis and Validation Centers (DAVs). The four DPCs are: The Jackson Laboratory (JAX), Memorial Sloan Kettering Cancer Center (MSK), Northwestern University (NWU), and University of California San Francisco (UCSF). The DRACC is composed of members from University of Miami (UM), European Bioinformatic Institute (EBI), University of Washington (UW), and Queen Mary University of London (QMUL). The three DAVs are: Fred Hutchinson Cancer Center (Fred-Hutch), The Jackson Laboratory (JAX), and Stanford University.

Data catalogue

The Data Catalogue provides access to MorPhiC studies, datasets, experimental metadata, protocols, and associated resources released by the consortium.

Depending on the study, datasets may include:

  • RNA sequencing (RNA-seq)
  • Single-cell transcriptomics
  • Chromatin accessibility data
  • CRISPR editing information
  • Computational analysis outputs
Additional dataset types may be added as new studies are released.

Experimental metadata describes the context in which a dataset was generated. This may include information such as the cell type, genome-editing strategy, sequencing platform, experimental protocol, quality control metrics, processing methods, and other details needed to interpret and reproduce the experiment.

A study is a collection of related experiments investigating one or more genes using specific cell models, genome-editing approaches, and molecular assays. Each study may contain multiple datasets and associated analyses.

A single gene may be investigated:

  • in different cell types
  • using different knockout strategies
  • with multiple molecular assays
  • at different experimental time points
Together, these complementary datasets provide a more complete understanding of gene function and the biological consequences of gene disruption.

Before public release, datasets undergo standardised validation, quality assessment, and consortium review to ensure they meet MorPhiC's data quality standards and are suitable for reuse by the research community.

Where available, the catalogue provides links to both raw and processed datasets, together with experimental metadata and documentation describing how the data were generated and analysed.

Processed datasets have undergone computational analysis, such as sequence alignment, quality filtering, normalisation, feature quantification, or other analytical workflows, making them easier to interpret and compare across studies.

The catalogue is expanded roughly 3 times a year as additional studies complete quality review and become available.

Each study and dataset includes citation information. Researchers should cite both the relevant dataset and any associated publications when using MorPhiC data in presentations or publications. More information on how to cite MorPhiC datasets can be found in the How to cite page.

Clonal cell lines

A clonal cell line is a population of genetically identical cells derived from a single genome-edited parent cell. Because all cells originate from the same clone, they carry the same engineered genetic modification, allowing experiments to be performed using a genetically uniform population.

Clonal cell lines help ensure that observed biological changes are caused by the intended gene knockout rather than by a population of cells carrying a mixture different genetic edits.

Following genome editing, typically using CRISPR-based methods, individual edited cells are isolated and expanded into separate colonies. Each clone is then validated to confirm that it carries the intended genetic modification before being used in downstream experiments.

Generating independent clones allows researchers to distinguish genuine biological effects of the targeted gene knockout from clone-specific variation or unintended editing events.

Clone validation may include sequencing of the edited genomic region, confirmation of the expected null allele, assessment of genomic integrity, and additional quality control assays before the clone is included in experimental studies.

A clonal cell line page may include:

  • Cell line identifier
  • Parent cell line
  • Target gene
  • Genome-editing strategy
  • Clone identifier
  • Validation results
  • Quality control information
  • Related datasets
  • Associated studies

The parental cell line is the original, unmodified cell line from which genome-edited knockout clones were generated. It serves as the experimental control when comparing the effects of gene disruption.

Yes. Although independent clones target the same gene, subtle biological differences may arise because of experimental variation or additional genetic changes acquired during cell culture. Comparing multiple validated clones increases confidence that observed phenotypes are caused by the intended gene knockout.

Some engineered cell lines may be distributed through consortium partners or external repositories, subject to availability, material transfer agreements, and the consortium's data and material sharing policies. Researchers can visit the Cell lines page to view available cell lines.

Genes

Protein-coding genes produce proteins that carry out most cellular functions. In Phase 1, MorPhiC focuses on approximately 1,000 protein-coding genes to evaluate scalable approaches for generating null alleles and characterising gene function before expanding to a larger catalogue.

In the context of the MorPhiC Consortium, we have operationally defined a null allele as one that reduces the amount of target protein or mRNA by 90% or more. MorPhiC creates null alleles to observe the molecular and cellular changes that occur when a gene is no longer functional, helping researchers understand its biological role.

Historically, genes associated with inherited diseases or common disorders have received the most attention. However, many human genes remain poorly characterised. MorPhiC aims to systematically study these understudied genes and generate a comprehensive resource for the research community.

Removing or inactivating a gene allows researchers to compare normal and knockout cells. Differences between them can reveal the biological processes, pathways, and cellular functions that depend on that gene.

Yes. Essential genes are included where possible, although they may require specialised experimental strategies because complete loss of function can prevent cells from surviving or dividing.

Yes. Many genes perform different functions depending on the cell type, developmental stage, or environmental conditions. As a result, disrupting a single gene may produce multiple molecular and cellular phenotypes.

Different cell types express different combinations of genes and proteins and perform specialised biological functions. Consequently, the same gene may have distinct roles in neurons, immune cells, epithelial cells, or stem cells, leading to different phenotypic outcomes after knockout.

The Phase 1 gene set was selected to represent a diverse range of biological pathways, disease relevance, gene families, and experimental challenges. This allows the consortium to evaluate multiple genome-editing and phenotyping strategies across a representative collection of genes.

Depending on the available data, a gene page may include gene annotations, knockout strategies, associated studies, molecular phenotypes, clonal cell lines, datasets, interactive visualisations, and links to external biological resources.

Molecular phenotypes are measurable changes at the molecular level that occur after a gene has been disrupted. These may include changes in gene expression, chromatin accessibility, protein abundance, signalling pathways, or other cellular processes.

Gene visualizations

The visualisations summarise experimental results for this gene across MorPhiC studies. Depending on the available data, they may display molecular phenotypes, gene expression changes, pathway enrichment, cell line information, datasets, and associated studies.

Different experiments measure different aspects of gene function. For example, one visualisation may show changes in gene expression, while another highlights affected biological pathways or available knockout cell lines. Together, they provide a more complete picture of the gene's role.

The available visualisations depend on the data generated for that gene. Some genes have been studied using multiple assays and cell types, while others may currently have fewer datasets available. As MorPhiC releases more data, additional visualisations may be added.

Not every experiment produces results for every gene, and some datasets may still be undergoing quality review or have not yet been released. Empty visualisations simply indicate that no data are currently available for that analysis.

Colours indicate different biological measurements depending on the figure. They may represent expression levels, statistical significance, fold changes, experimental groups, or different cell lines. Refer to the legend accompanying each visualisation for its specific meaning.

Fold change describes how much a measurement has changed relative to a control.

  • A positive fold change indicates an increase
  • A negative fold change indicates a decrease.
Fold change is commonly used to compare gene expression between knockout and control cells.

Differential gene expression analysis identifies genes whose expression changes significantly after the target gene has been knocked out. These changes help researchers understand which biological processes may be affected.

Visualisations often display only the most statistically significant or biologically relevant genes to improve readability. Complete datasets can usually be accessed through the associated study or downloadable files.

Many visualisations include statistical measures such as p-values, adjusted p-values (FDR), confidence intervals, or effect sizes. These help indicate whether observed differences are likely to reflect genuine biological changes rather than random variation.

Pathway enrichment analysis identifies biological pathways that contain more affected genes than would be expected by chance. This helps researchers understand which cellular processes may be disrupted by the gene knockout.

A molecular phenotype score summarises the magnitude or significance of molecular changes observed after gene disruption. The exact calculation depends on the assay and analysis pipeline used.

Network visualisations illustrate known or predicted relationships between genes or proteins. Connections may represent physical interactions, shared biological pathways, regulatory relationships, or functional associations.

Not necessarily. Some connections represent experimentally validated physical interactions, while others reflect shared pathways, co-expression, or computational predictions. The source of each interaction should be described in the accompanying metadata.

A gene can have different biological functions in different cell types. Studying multiple cell lines allows researchers to identify cell type-specific effects as well as functions that are conserved across tissues.

Different cell types express different genes, proteins, and regulatory networks. As a result, disrupting the same gene may produce different molecular phenotypes in different biological contexts.

Yes. Gene pages include links to related studies and datasets, allowing researchers to compare molecular phenotypes, available cell lines, and experimental results across multiple genes.

Differences may arise from variations in cell type, genome-editing strategy, experimental conditions, sequencing methods, or analytical pipelines. Examining multiple studies provides a more comprehensive understanding of gene function.

Most visualisations are linked to the underlying datasets or studies. Where available, users can download raw or processed data together with experimental metadata and documentation.

Yes. As MorPhiC releases additional datasets and analyses, existing gene pages may be updated with new visualisations, datasets, and biological insights.