Skip to content

About

Audit of single-cell perturbation data and curated gene regulatory networks. Reliability spans 112-fold; curated networks predict almost none of it. Code, results and verification for the paper.

Resources

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

An audit of perturbation data and curated gene regulatory networks

DOI

Code, result files and figures for:

An audit of perturbation data and curated gene regulatory networks: reliability spans two orders of magnitude, and curated networks predict almost none of it. Ihor Kendiukhov. Institute of Medical Genetics and Applied Genomics, University of Tübingen.

Perturbation screens are used as ground truth and curated networks as prior knowledge. The quality of neither has been established. This repository contains everything needed to check the measurements that establish both.


Start here

git clone https://github.com/Biodyn-AI/perturbation-audit.git
cd perturbation-audit
pip install -r requirements.txt
python verify/verify_claims.py

That re-derives every headline number in the paper from the result files shipped in results/. It needs no raw data, no downloads and no configuration, and finishes in seconds. If it prints 25/25 claims verified, this repository and the manuscript agree.

Reproducing the results from raw data is a much larger job and is documented separately in docs/REPRODUCING.md.


What the paper found

Reliability of a perturbation response, measured with two independent reagents spans 112-fold across 17 datasets ($r = 0.008$–$0.876$)
The standard cell split used throughout the literature overstates that by a median of 2.26×, up to 29×
Best of 18 curated networks and 7 foundation models at predicting which genes respond +0.0087 AUROC over control
The same resources at transcription-factor activity inference 78–239× better than at membership
Co-expression from control cells, given equal numbers of predictions exceeds 7 of 9 curated resources
Pooling several screens vs. the best single screen $r = 0.321$ vs. $0.261$

The resources work in aggregate and fail per target. What is achievable is set by the dataset, not by the network.


Layout

├── verify/verify_claims.py     re-derives every headline number from results/ (start here)
├── results/                    140 result files -- every number in the paper comes from these
├── figures/                    all published figures, as PDF
├── perturbceiling/             the released tool: a dataset's ceiling before you benchmark on it
├── src/
│   ├── common/                 shared loaders: structure detection, network resources
│   ├── acquire/                dataset download with stall detection
│   ├── part_a_reliability/     how reproducible is a perturbation response?
│   ├── part_b_networks/        do curated networks and foundation models predict perturbation?
│   ├── part_c_probing/         interrogating the negative: calibration, transfer, combination
│   ├── part_d_filtering/       within-dataset quality, filtering, pooling
│   ├── controls/               controls added during revision
│   └── superseded/             analyses replaced by better ones -- kept, and labelled, for honesty
└── docs/
    ├── REPRODUCING.md          running the pipeline from raw data, with runtimes and memory
    ├── DATA_SOURCES.md         every input and where it came from
    └── CLAIM_INDEX.md          paper claim -> script -> result file

Every script begins with a docstring explaining what it measures, which confounds it controls, and where the design would go wrong without them. Those docstrings are the real documentation; this file is only a map.

perturbceiling

A single-file tool that reports a dataset's reagent-split reliability and attenuation ceiling before it is used as a benchmark. It detects gene and reagent columns automatically, applies a same-construct filter, and reports the mismatched floor alongside. Where no reagent replicate exists it says so rather than falling back to a cell split. See perturbceiling/.

A note on what is not here

The multi-gigabyte intermediate caches (p3_quarters, p2_coexpr, partB_preps, p2_reagent) are not included. They are large, and every one is regenerated deterministically from public data by the scripts in src/. docs/REPRODUCING.md gives the order and the cost.

Unpublished in-house model checkpoints that were scored during development are not distributed and their loader has been removed; the paper reports only the seven public models.

Citation and archive

This repository is archived at Zenodo: doi:10.5281/zenodo.22062254. That is the concept DOI and always resolves to the latest archived version; use it if you need a citable, immutable snapshot rather than the moving main branch.

Machine-readable metadata is in CITATION.cff. Please cite the paper rather than the software when referring to the findings.

License

MIT, see LICENSE. The datasets and network resources analysed here carry their own licences and are not redistributed.

About

Audit of single-cell perturbation data and curated gene regulatory networks. Reliability spans 112-fold; curated networks predict almost none of it. Code, results and verification for the paper.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages