A pure-Python port of Bioconductor edgeR (Robinson, McCarthy & Smyth, Bioinformatics 2010) — negative-binomial models for differential expression of count data.
- No
rpy2, no R install — the edgeR negative-binomial GLM workflow reimplemented in NumPy / SciPy, vectorised over genes - Matches Bioconductor edgeR 4.0 numerically (see Numerical parity)
- The canonical pipelines:
DGEList → calcNormFactors → estimateDisp → glmQLFit → glmQLFTestand the classicexactTest - TMM normalization, common / trended / tagwise dispersion, GLM and quasi-likelihood F-tests,
filterByExpr,cpm/aveLogCPM - Both Python-style (
glm_fit,estimate_disp,top_tags) and R-style (glmFit,estimateDisp,topTags) names exported
This is a standalone mirror of the implementation developed in
omicverse, where it powers the edgeR differential-expression backend ofov.bulk/pyDEG.
pip install pyedgerimport numpy as np
import pyedger
# counts: genes x samples raw count matrix; group: per-sample condition labels
dge = pyedger.DGEList(counts=counts, group=group)
dge = pyedger.calcNormFactors(dge) # TMM normalization
keep = pyedger.filterByExpr(dge, group=group)
dge = dge[keep]
# Quasi-likelihood F-test workflow (the recommended edgeR pipeline)
dge = pyedger.estimateDisp(dge, design)
fit = pyedger.glmQLFit(dge, design)
qlf = pyedger.glmQLFTest(fit, coef=1)
res = pyedger.topTags(qlf, n=np.inf)
res.head()dge = pyedger.estimateDisp(dge, design)
et = pyedger.exactTest(dge, pair=("control", "treated"))
pyedger.topTags(et)| Python | R counterpart |
|---|---|
DGEList |
DGEList |
calc_norm_factors / calcNormFactors |
calcNormFactors (TMM) |
filter_by_expr / filterByExpr |
filterByExpr |
estimate_disp / estimateDisp |
estimateDisp |
glm_fit / glmFit, glm_lrt / glmLRT |
glmFit, glmLRT |
glm_ql_fit / glmQLFit, glm_qlf_test / glmQLFTest |
glmQLFit, glmQLFTest |
exact_test / exactTest |
exactTest |
glm_treat / glmTreat |
glmTreat |
cpm, ave_log_cpm / aveLogCPM |
cpm, aveLogCPM |
top_tags / topTags |
topTags |
decide_tests_dge / decideTests |
decideTestsDGE |
DGEGLM, DGELRT, DGEExact, TestResults |
the corresponding S4 classes |
tests/test_r_parity.py checks pyedger against edgeR 4.0.16 (R 4.3.3, limma 3.58.1,
locfit 1.5-9) on a paired design with zero-count library pairs; the reference
values are regenerated by tests/data/make_reference.R. Since 0.1.1 the GLM
kernels follow edgeR's compiled code step by step and run vectorised over all
genes:
| Step | Ported from | Agreement with R |
|---|---|---|
glmFit (Levenberg, one-way shortcut, start values) |
glm_levenberg.cpp, glm_one_group.cpp |
~1e-9 |
estimateDisp with a design (Cox-Reid APL, zero-fit groups) and classic mode |
R_compute_apl.cpp, adj_coxreid.cpp, q2qnbinom |
~1e-10 |
Trend: locfitByCol |
locfit tree evaluation + interpolation | bit-identical |
maximizeInterpolant, aveLogCPM |
interpolator.cpp, R_ave_log_cpm.cpp |
~1e-13 |
squeezeVar / fitFDist with a covariate |
limma (natural-spline trend) | ~1e-14 |
glmQLFit(legacy=False) bias-adjusted deviance/df |
ql_glm.c, ql_weights.c, .computePriorS2 |
~1e-10 |
glmQLFit(legacy=False) and estimateGLMCommonDisp depend on edgeR's
optimize(tol=1e-5) over dispersion^(1/4), so R's own value is only defined to
~4e-5 relative; agreement there is at that level.
Defaults follow edgeR 4.0: glmQLFit(legacy=True, abundance_trend=True) and
estimateDisp(prior_df=None) (prior df estimated).
Robinson, M.D., McCarthy, D.J., Smyth, G.K. edgeR: a Bioconductor package for differential expression analysis of digital gene expression data. Bioinformatics 26(1), 139–140 (2010).
…and acknowledge omicverse / this repo for the Python port.
LGPL-3.0-or-later — matches the upstream Bioconductor package.