You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Brings the FSI residue re-curation onto the public framework and makes the
self-audit story a first-class part of the README — the rigor signal the
project should lead with.
Framework:
- map_uniprot_to_pdb_positions() amino-acid identity check (06) — silent
residue mismaps now fail loudly
- Cholera/Abrin catalytic_residues re-curated vs UniProt; SEB excluded
from FSI (exclude_from_fsi, honored by 06/12/13); Anthrax fixed
- functional_sites.json + all FSI/FSPE/SER/control result files and
framework figures regenerated on the corrected panel
Docs / public surfaces:
- New README "Evaluation Integrity" section: the audit, the loud-failure
check, the honest 5-to-3 significant-count correction
- docs/FSI_NUMBERING_AUDIT.md gains a Resolution section
- docs/DATA_CORRECTIONS.md gains the re-curation entry
- README + huggingface/README.md Key Results tables corrected (FSI, FSPE,
negative controls, realizability) — BoNT/A 2.87->2.24, Cholera 0.22->0.53,
SEB excluded, BoNT/A re-keyed P10844->P0DPI1; AUROC 0.994->0.981 sync
The paper manuscript is not part of this repo and remains private.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
ESM-2 embeddings nearly perfectly separate a toxin set from a benign homolog set (60 vs. 60 sequences) using a logistic regression classifier in the full 1280-dimensional embedding space. Dangerous queries retrieve other dangerous proteins with 91.7% precision at rank 1: without any fine-tuning or task-specific supervision.
> **Note**: The t-SNE projection (2D) shows partial visual overlap between classes. This does not contradict the AUROC = 0.994 result: logistic regression operates in the full 1280-dimensional space where the classes are nearly linearly separable. t-SNE is a dimensionality reduction for visualization only.
149
+
> **Note**: The t-SNE projection (2D) shows partial visual overlap between classes. This does not contradict the AUROC = 0.981 result: logistic regression operates in the full 1280-dimensional space where the classes are nearly linearly separable. t-SNE is a dimensionality reduction for visualization only.
**Mean FSI: 1.02** across the 7 FSI-scored structures (100 designs each). SEB is
167
+
excluded — a superantigen has no discrete catalytic site to recover (see
168
+
[Evaluation Integrity](#evaluation-integrity)). These values reflect the
169
+
2026-05 residue re-curation; two structures (Abrin, Ricin) sit close to 1.0 and
170
+
are not significant after Holm–Bonferroni correction.
134
171
135
172
The heterogeneity is scientifically informative, not a limitation:
136
173
137
-
-**BoNT-A (FSI = 2.87)**: The zinc-protease light chain imposes tight backbone constraints. In *every* one of 100 designs the model recovers functional residues beyond chance: backbone geometry unambiguously encodes dangerous function.
138
-
-**Tetanus toxin LC (FSI = 1.75)**: The zinc-dependent endopeptidase light chain shares mechanistic architecture with BoNT-A and shows similarly strong backbone-level specificity.
139
-
-**Abrin (FSI = 1.13)** and **Ricin (FSI = 1.10)**: Both ribosome-inactivating proteins show consistent functional recovery; the active-site Tyr–Glu–Arg triad is conserved across RIP-family designs.
140
-
-**SEB (FSI = 0.70)**: Superantigen activity arises from a distributed T-cell receptor interface, not enzymatic catalysis: backbone-level encoding is absent.
174
+
-**BoNT-A (FSI = 2.24)**: The zinc-protease light chain imposes tight backbone constraints. In 94 of 100 designs the model recovers functional residues beyond chance: backbone geometry unambiguously encodes dangerous function.
175
+
-**Tetanus toxin LC (FSI = 1.77)**: The zinc-dependent endopeptidase light chain shares mechanistic architecture with BoNT-A and shows similarly strong backbone-level specificity.
176
+
-**Abrin (FSI = 1.10)** and **Ricin (FSI = 1.07)**: Both ribosome-inactivating proteins recover the active-site Tyr–Tyr–Glu–Arg–Trp residues at a rate that is *directionally* above 1.0 but not significant after Holm–Bonferroni correction — a genuinely marginal signal, reported as such.
177
+
-**SEB**: Superantigen activity arises from a distributed T-cell receptor interface, not enzymatic catalysis. With no discrete catalytic site (and no UniProt-annotated functional residues), SEB is excluded from FSI rather than scored on an unverifiable residue set.
141
178
-**Streptolysin O (FSI = 0.45)**: Pore-forming activity requires ordered oligomerization on cholesterol-containing membranes; the monomeric backbone alone cannot encode this.
142
-
-**Cholera CTA1 (FSI = 0.22)**: Functional activity requires holotoxin assembly; the monomer backbone does not encode the relevant function.
179
+
-**Cholera CTA1 (FSI = 0.53)**: Functional activity requires holotoxin assembly; the monomer backbone only weakly encodes the relevant function.
143
180
-**Anthrax PA (FSI = 0.00)**: The phi-clamp phenylalanine (Krantz 2005) occupies a sterically unusual position that backbone geometry cannot constrain. Zero functional recovery across 100 designs is the most interpretable result in the dataset.
144
181
145
182

@@ -150,16 +187,16 @@ The heterogeneity is scientifically informative, not a limitation:
150
187
151
188
| Control | Mechanism match | Control FSI | Matched toxin FSI |*p* (Mann–Whitney, per-seq) |
The BoNT-A three-way comparison dissects fold geometry from zinc chemistry from dangerous function:
159
196
160
-
-**1AST (Astacin, FSI = 1.88)**: same HExxH zinc-binding fold as BoNT-A → elevated FSI confirms fold geometry contributes.
161
-
-**1LNF (Thermolysin, FSI = 1.66)**: same zinc chemistry (HExxH motif), but a *different fold* → elevated FSI persists, showing zinc chemistry alone also elevates specificity.
162
-
-**3BTA (BoNT-A, FSI = 2.87)**: significantly higher than both controls (p < 0.0001 vs both) → dangerous toxin function is encoded *beyond* what either shared fold geometry or shared zinc chemistry explains.
197
+
-**1AST (Astacin, FSI = 1.85)**: same HExxH zinc-binding fold as BoNT-A → elevated FSI confirms fold geometry contributes.
198
+
-**1LNF (Thermolysin, FSI = 1.69)**: same zinc chemistry (HExxH motif), but a *different fold* → elevated FSI persists, showing zinc chemistry alone also elevates specificity.
199
+
-**3BTA (BoNT-A, FSI = 2.24)**: significantly higher than both controls (p < 0.0001 vs both) → dangerous toxin function is encoded *beyond* what either shared fold geometry or shared zinc chemistry explains.
**Mean FSPE ratio: 0.836. Pooled meta-analysis: p = 0.073, r = 0.15.**
219
+
**Mean FSPE ratio: 0.66. Pooled meta-analysis: p = 2.6 × 10⁻⁸, r = 0.41** (n = 74 functional vs 300 background residues).
183
220
184
-
FSPE provides directional evidence (5/7 proteins, mean ratio 0.84) with Tetanus LC reaching significance (p < 0.0001, r = 1.00). Individual Mann–Whitney tests are structurally underpowered for most proteins given 3–9 annotated catalytic sites vs ~100 background residues. Tetanus LC is an exception: its 4 zinc-coordinating residues show near-perfect entropy discrimination (functional entropy 0.36 vs background 2.50). The embedding separability (AUROC = 0.994) confirms ESM-2 encodes functional information; FSPE localizes that encoding to specific residue positions with variable resolution depending on site density.
221
+
FSPE provides directional evidence (5/7 proteins show ratio < 1, mean 0.66), with Tetanus LC and BoNT-A reaching per-protein significance (both p < 0.0001, r = 1.00) and Cholera nominally significant (p = 0.014). Individual Mann–Whitney tests are structurally underpowered for proteins with few annotated catalytic sites; the pooled meta-analysis (p = 2.6 × 10⁻⁸) is the better-powered test and is now strongly significant. The embedding separability (AUROC = 0.981) confirms ESM-2 encodes functional information; FSPE localizes that encoding to specific residue positions. *(BoNT-A is now keyed to its correct accession P0DPI1; the prior P10844 entry was BoNT type B — see [`docs/DATA_CORRECTIONS.md`](docs/DATA_CORRECTIONS.md).)*
> **Note on the pooled distribution**: The functional sites histogram shows a bimodal shape: a heavy left tail at entropy ≈ 0 and a broad peak at entropy ≈ 2.0–2.8. The left tail is driven entirely by P04958 (Tetanus LC), whose 4 zinc-coordinating residues have near-zero prediction entropy. Removing P04958, the remaining 6 proteins show a unimodal distribution with a modest left-shift relative to background (mean 2.19 vs 2.37). This heterogeneity is reported in the per-protein breakdown above.
225
+
> **Note on the pooled distribution**: The functional-site entropy histogram has a heavy left tail at entropy ≈ 0, driven by the two strongest proteins — Tetanus LC and BoNT-A — whose zinc-coordinating residues have near-zero prediction entropy. The remaining proteins contribute a more modest left-shift relative to background. This heterogeneity is reported in the per-protein breakdown above.
| Anthrax PA (1ACC) | NONE (0.00) | 4 | Multi-component + heptamerization | very low |
204
240
205
241
The critical insight: **the two highest-FSI toxins (BoNT-A and Tetanus LC) also have the highest physical barrier (Tier 4)**. A framework measuring only computational risk would rank these most dangerous and potentially misdirect resources away from lower-FSI but more easily realizable threats.
0 commit comments