Software ecosystem

Software and resources

Our software ecosystem supports the lab's research programme in cell identity, interpretable AI, single-cell and spatial omics, organoid systems, disease-state modelling, method benchmarking, and biological fidelity assessment.

The resources below are organised to match the five research directions on the Research page, with a cross-cutting Benchmark tab for standard-setting work. General-purpose tools and legacy projects are retained in one additional tab.

Computational Systems Biology lab logo

Organisation

  • Aligned to the five research directions.
  • Benchmarking, standard-setting and fidelity-assessment resources grouped separately.
  • Broad-use and legacy tools retained in one additional tab.
  • Links to code, documentation, and papers.

Research direction 1

Cell identity and cell-fate decisions

Tools for defining cell identity, improving annotation, integrating single-cell references, modelling multimodal identity signals and prioritising candidate regulators or compounds for cell-state conversion.

Single-cell integration

scMerge

Integrates and normalises single-cell RNA-seq data using stably expressed genes and pseudo-replicates.

Multimodal identity modelling

Matilda

Multi-task learning framework for multimodal single-cell omics, including simulation, dimension reduction, classification and feature selection.

Cell-state conversion

Refate

Prioritises candidate genes, networks and compounds for directed cellular conversion from starting to target cell states.

Research direction 2

Interpretable AI for single-cell systems biology

Tools and frameworks for learning from high-dimensional single-cell and multimodal data while exposing the features, programs and cell states that support each prediction.

Multi-task learning

Matilda

Jointly learns several multimodal single-cell tasks, including simulation, classification, feature selection and representation learning.

Feature selection benchmark

scDeepFeatures

Evaluation resource for deep learning-based feature selection in single-cell RNA-seq analysis.

Positive-unlabelled learning

AdaSampling

Adaptive sampling for positive-unlabelled and label-noise learning in bioinformatics applications.

Research direction 3

Single-cell, spatial and multimodal omics tools

Tools for data integration, annotation, multimodal analysis, phosphoproteomics and related machine-learning foundations.

CITE-seq analysis

CiteFuse

Supports preprocessing, integration, clustering, differential analysis and visualisation for CITE-seq data.

Phosphoproteomics

PhosR

Processes phosphoproteomic data and supports kinase, pathway, signalome and phosphorylation-site analyses.

Knowledge-guided clustering

CLUEY

Knowledge-guided framework for cell-type detection and clustering of unimodal and multimodal single-cell omics data.

Reference stable features

Stable genes and phosphosites

Reference stable genes and phosphorylation sites for single-cell and phosphoproteomic normalisation and integration.

Phosphoproteomic clustering

ClueR

Knowledge-based clustering for phosphoproteomic time-series data using kinase-substrate relationships.

Pathway and kinase analysis

directPA and KinasePA

Pathway and kinase perturbation analysis for transcriptomic, proteomic and phosphoproteomic datasets.

Kinase-substrate prediction

KSP-PUEL

Positive-unlabelled ensemble learning for kinase-substrate prediction from dynamic phosphoproteomics.

Cross-cutting impact

Benchmarking, standards and fidelity assessment

Method benchmarks and biological fidelity resources that help the field move from plausibility-based assessment towards quantitative, task-specific and biologically grounded standards.

Cell-type-number estimation

scCCESS

Consensus-clustering framework and benchmark resource for estimating the number of cell types in single-cell RNA-seq data.

Feature-selection evaluation

scDeepFeatures

Evaluation resource for deep learning-based feature selection in single-cell RNA-seq analysis.

Retinal organoid fidelity

Eikon

Atlas-linked resource for assessing retinal cell identity and the fidelity of human retinal organoids.

Brain culture fidelity

BrainSTEM

Single-cell fetal brain atlas resource for assessing transcriptomic fidelity of human midbrain cultures.

Pluripotency reference

Stem Cell Atlas

Multi-omic reference atlas for benchmarking phased progression of pluripotency and stem-cell state transitions.

Research direction 4

Stem cell and organoid systems

Resources for assessing cell-state fidelity, developmental progression and disease-relevant organoid systems.

Pluripotency atlas

Stem Cell Atlas

Multi-omic resource for phased progression of pluripotency.

Cerebral organoids

TransOmicsData

Data and code resource for trans-omic profiling of early human cerebral organoid formation.

Research direction 5

Disease-state modelling and therapeutic prioritisation

Tools for identifying disease-associated cell states, interpreting altered molecular features and prioritising candidate targets or compounds.

Sample-level modelling

scFeatures

Generates sample-level multi-view representations from single-cell and spatial data for disease-outcome modelling.

Candidate intervention prioritisation

Refate

Links regulatory-network modelling to candidate compound prioritisation for directed cell-state conversion and reprogramming hypotheses.

Signalling and phosphorylation

PhosR

Supports phosphoproteomic processing and kinase/signalling analyses relevant to disease-state interpretation.

General-purpose tools

General purpose machine learning tools

AdaSampling

Semi-supervised adaptive sampling for positive-unlabelled learning and learning from noisy class labels in bioinformatics applications.

Sample Subset Optimization

Evolutionary sample-subset selection for imbalanced data and ensemble learning problems in bioinformatics applications.

Legacy projects

Archived Google Code projects

Historical software projects are kept here as part of the lab's research and methods record.

Re-Fraction

Machine-learning algorithm in Java for protein inference.

Self-Boosted Percolator

Boosted learning algorithm in Java for peptide filtering.

Genetic Ensemble SNPX

Parallel genetic algorithm in Java for SNP interaction detection.

Ensemble of Filters

Ensemble algorithm in Perl for SNP interaction filtering.

OCAP

Open source mass spectrometry analysis pipeline in R.

DyWave

Dynamic wavelet package in C/C++ for mass spectrum modelling.

Imbalanced Data Sampling

Particle swarm optimisation algorithm in Java for imbalanced data sampling.