Skip to content

Repository files navigation

FIQ-1 — A Fragment-Grown, Structure-Validated Nrf2-Neh1 Binder from a One-Step Isoquinolinone Scaffold via an ML-Guided Workflow

A reproducible, end-to-end in silico workflow that couples a machine-learning docking surrogate with receptor-aware fragment growing to design an Nrf2-Neh1 binder, then carries it through structure-enabled optimization, re-docking against an experimental structure, explicit-solvent molecular dynamics, MM/GBSA, and a systems-pharmacology mapping of the target.

The product is not a validated drug candidate. It is a route to one: how a one-step isoquinolinone scaffold (FIQ-0) was found and then optimized into a lead-efficiency molecule (FIQ-1), what the synthesis would cost, and where the target sits in its disease network. All affinities are docking estimates within AutoDock Vina / Vinardo error and carry no experimental measurement of activity.

This repository accompanies the manuscript "FIQ-1: A Fragment-Grown, Structure-Validated Nrf2-Neh1 Binder Optimized from a One-Step Isoquinolinone Scaffold via an ML-Guided Workflow." Docking, MD, MM/GBSA, and systems-pharmacology protocols and input files are provided here.

Background

Nrf2 (nuclear factor erythroid 2-related factor 2) is the master transcription factor of the cellular antioxidant response. Under normal conditions its repressor Keap1 targets Nrf2 for proteasomal degradation, but oxidative or electrophilic stress stabilizes Nrf2, which enters the nucleus and activates cytoprotective genes through antioxidant response elements (AREs). In many tumors, mutations in KEAP1 or NFE2L2 lock Nrf2 in a constitutively active state; the resulting sustained antioxidant program helps cancer cells survive chemotherapy and radiotherapy, producing acquired drug resistance. Most reported Nrf2 inhibitors (e.g. ML385) target the Keap1–Nrf2 interface and came from resource-intensive high-throughput screens. The DNA-binding step is instead carried out by the Neh1 domain, a CNC-bZIP region that pairs Nrf2 with small Maf proteins and recognizes the ARE sequence — a structurally defined but under-explored site for a small molecule.

The two molecules

The workflow produces two molecules. FIQ-0 is the synthesizable intermediate that the fragment-growing step lands on; it is not proposed as an inhibitor. FIQ-1 is the structure-guided optimization of FIQ-0 and the molecule the paper carries forward.

FIQ-0 — one-step scaffold (intermediate)

Property Value
Name 3-(2-isopropylphenyl)isoquinolin-1(2H)-one
SMILES O=C1NC(=Cc2ccccc12)c1ccccc1C(C)C
Formula C₁₈H₁₇NO
InChIKey RTEXNGJLZKEKBT-UHFFFAOYSA-N
Molecular weight 263.34 g/mol
Heavy atoms 20
Computed logP 4.17
TPSA 29.1 Ų
H-bond donors / acceptors 1 / 1
Rotatable bonds 2
Lipinski violations 0
Vinardo affinity (7X5G, Pocket 1) −6.4 kcal/mol
Ligand efficiency 0.32
Predicted synthesis one step (Spaya, RScore 0.8)

FIQ-1 — optimized lead

Property Value
Name (E)-3-[4-(6-fluoro-1-oxo-1,2-dihydroisoquinolin-3-yl)-3-methylphenyl]prop-2-enoic acid
SMILES OC(=O)/C=C/c1ccc(-c2cc3cc(F)ccc3c(=O)[nH]2)c(C)c1
Formula C₁₉H₁₄FNO₃
InChIKey SNNGMVOJWHZXFV-XVNBXDOJSA-N
Molecular weight 323.3 g/mol
Heavy atoms 24
QED / cLogP 0.72 / 3.7
Vinardo affinity (7X5G, Pocket 1) −8.8 kcal/mol
Ligand efficiency 0.37
Predicted synthesis five-step convergent (Spaya)

FIQ-1 was obtained from FIQ-0 by three edits along the scaffold's open growth vector: the ortho-isopropyl group was truncated to a methyl, the isoquinolinone ring was fluorinated, and an (E)-cinnamic-acid arm was appended. On the experimental 7X5G structure this reaches the docking-score class of the established inhibitor ML385 (−9.1 kcal/mol) at 188 Da lower mass and 13 fewer heavy atoms, with the highest ligand efficiency of the three molecules (0.37 vs 0.32 for FIQ-0 and 0.25 for ML385). The efficiency advantage — not the raw score — is the substantive point.

Pipeline

The workflow is sequential: Phase 1 decides where to bind and builds a cheap way to search; Phase 2 uses that pocket and surrogate to build FIQ-0 by fragment growing; Phases 3–8 carry the molecule from a docked pose to a relaxed, parametrized, solvated structure and check that it can be synthesized. The experimental-structure re-docking, MD/MM-GBSA, and systems-pharmacology analyses (below) extend the work onto the CNC-bZIP crystal structure (PDB 7X5G) that became available afterward.

  1. ZINC screen and ML surrogate (target validation + screening engine) — a drug-like ZINC subset was reduced to 500 compounds by Lipinski/ADMET and PAINS filtering, then docked against three PrankWeb-predicted pockets of a predicted (AlphaFold) Nrf2-Neh1 model (AutoDock Vina, exhaustiveness 8). Pocket 1, in the CNC-bZIP region implicated in ARE–DNA recognition, gave the most favorable score (−3.99 kcal/mol) and was carried forward. A random-forest surrogate on Morgan ECFP4 fingerprints was expanded by three rounds of active learning to 12,964 labeled molecules (R² = 0.53, MAE = 0.23 kcal/mol on a 20% held-out set) and used to sweep ≈14.6 M ZINC molecules; a pharmacophore filter kept 11,530 hits, the top 15 of which re-docked as genuine (but only modest, −4.3 to −5.2 kcal/mol) binders. The takeaway that motivates Phase 2: the shallow, hydrophobic pocket is poorly matched to whole lead-sized molecules.

  2. Fragment-based design — about 1,000 ZINC fragments passing the Astex rule of three were docked into Pocket 1 (exhaustiveness 32) and ranked by ligand efficiency. Cross-pocket fragment linking was attempted but proved geometrically infeasible for all linker lengths from 3 to 11 Å, so receptor-aware fragment growing was used instead: growth vectors were computed on the best validated scaffold, 3-phenylisoquinolin-1(2H)-one (ZINC 3120460, −5.18 kcal/mol, LE ≈ 0.30), and medicinal-chemistry groups were grown along open vectors under a core constraint. Adding an ortho-isopropyl group improved the score by 0.31 kcal/mol while preserving the binding mode (core RMSD 0.10 Å), giving FIQ-0.

  3. Lead candidate docking — FIQ-0 docked into Pocket 1; the top pose was profiled with PLIP (two backbone H-bonds from the lactam, Lys506/Lys508; the isopropyl contacts Val509).

  4. Ligand parametrization — CGenFF atom typing and topology. A carbon-radical artifact caused by non-polar hydrogens absent from the docking output was corrected by all-atom reprotonation in PyMOL (h_add).

  5. System construction — solvation in a cubic TIP3P water box (≈80 Å edge), neutralization to 0.15 M KCl, CHARMM36m/CGenFF topology (CHARMM-GUI).

  6. Energy minimization — restrained steepest-descents minimization in GROMACS to Fmax < 1000 kJ mol⁻¹ nm⁻¹.

  7. Analysis — extraction and plotting of the potential-energy trajectory (and, on the MD side, ligand/backbone RMSD, contact occupancy, MM/GBSA decomposition).

  8. Retrosynthetic assessment — Spaya (Iktos) predicts a one-step route for FIQ-0 and a fully disclosed five-step convergent route for FIQ-1 from commercial building blocks. Neither molecule was synthesized in this work.

Key results

Screening and surrogate validation

The surrogate reproduced Vina scores with R² = 0.53 and MAE = 0.23 kcal/mol on a held-out set — adequate for enrichment triage over a large library, but it did not rank correctly within the high-scoring set (Pearson r ≈ 0). It is therefore a coarse pre-filter that lowers docking load, not a ranker; final selection rested on explicit re-docking of the shortlist. Its one durable advantage is reusability: once trained, it can pre-score new ZINC tranches without further docking.

FIQ-0 interaction profile (PLIP, AlphaFold model, chain A)

The docked pose is anchored by two backbone hydrogen bonds from the lactam (Lys506, 3.35 Å; Lys508, 4.03 Å) and six hydrophobic contacts (Arg503, Arg504, Lys508, Val509). The grown isopropyl group contacts Val509. Source data: 03_lead_candidate_docking/plip_report.xml.

Re-docking against the experimental structure (PDB 7X5G)

After the experimental Nrf2(A510Y)–MafG/ARE-DNA complex (PDB 7X5G) was released, every molecule was re-docked into the same experimental receptor under one consistent Vinardo protocol (Smina; box center (−47.79, −20.09, 10.35) Å, 22 × 22 × 22 ų, fixed seed). This puts FIQ-0, FIQ-1, and ML385 on a directly comparable footing.

Molecule Formula MW (g/mol) Heavy atoms Vinardo (kcal/mol) LE
FIQ-0 (starting scaffold) C₁₈H₁₇NO 263.3 20 −6.4 0.32
FIQ-1 (optimized lead) C₁₉H₁₄FNO₃ 323.3 24 −8.8 0.37
ML385 (reference inhibitor) C₂₉H₂₅N₃O₄S 511.6 37 −9.1 0.25

FIQ-1 reaches the docking-score class of ML385 (within 0.3 kcal/mol, i.e. inside docking error) at 188 Da lower mass, with the highest ligand efficiency of the three.

Molecular dynamics and MM/GBSA (matched, on 7X5G)

Backbone-restrained 20 ns MD (CHARMM36m/TIP3P, backbone harmonically restrained so the fold stays rigid at ≈0.33 Å) shows that neither FIQ-1 nor ML385 holds its docked pose in this shallow, solvent-exposed groove — an apparent property of the site rather than a flaw specific to FIQ-1. ML385 stays near 4 Å ligand-RMSD before a late excursion; FIQ-1 departs within the first nanosecond and drifts to 16–24 Å while remaining surface-associated in 96% of frames. End-point MM/GBSA (igb=8, 0.15 M salt, 76 frames) on the matched 7X5G footing:

Component FIQ-1 (kcal/mol) ML385 (kcal/mol)
van der Waals −13.00 −28.95
Electrostatic −149.41 −22.64
Polar solvation (E_GB) +154.43 +36.14
Nonpolar solvation −2.07 −3.36
Total ΔG_GB −10.04 ± 4.93 −19.02 ± 4.69

FIQ-1 reaches roughly half of ML385's end-point favorability while binding the groove more loosely and less specifically, at much lower mass. These are upper bounds on favorability (no entropy term) and should be read as a relative, same-protocol ranking. (The earlier FIQ-0 MM/GBSA on the AlphaFold model, −6.55 ± 0.40 kcal/mol, uses a different receptor and force field and is not commensurate with the table above.)

Energy minimization (system-construction check)

Restrained steepest descents converged in 659 steps (final Fmax = 962.8 kJ mol⁻¹ nm⁻¹), potential energy decreasing monotonically — a clash-free relaxation, not a binding measurement.

Quantity Value
Initial potential energy −764,617.81 kJ/mol
Final potential energy −796,748.69 kJ/mol
ΔE −32,130.88 kJ/mol
Steps to convergence 659

Potential energy vs. minimization step

Systems-pharmacology context of the target

The KEAP1–NFE2L2–ARE axis (21 core genes) was mapped onto public interactome, genetic, and expression resources to place the target (not FIQ-1) in its disease network, using hepatocellular carcinoma (HCC; MONDO:0007256) as the indication. Full metrics and sources: docs/SYSTEMS_PHARMACOLOGY.md. In brief: the axis is a tightly connected module (57-fold STRING enrichment, p ≈ 10⁻¹⁶); NFE2L2 ranks 18/15,470 Open Targets associations for HCC; the axis is altered in ~8% of TCGA-LIHC tumours with strictly mutually exclusive NFE2L2/KEAP1 lesions; and no Nrf2 inhibitor is in clinical development — a genetically validated, inhibitor-free node.

Synthetic accessibility

Spaya returned a one-step route for FIQ-0 (RScore 0.8) from N,2-dimethylbenzamide and 2-isopropylbenzonitrile, and a fully disclosed five-step convergent route for FIQ-1 built on 6-fluoroisoquinolin-1(2H)-one (CAS 214045-85-9; ≈$535/g at 1 g scale, all building blocks catalog compounds). See 08_retrosynthesis/.

Repository structure

.
├── 01_zinc_screen_ml_surrogate/    pocket selection + surrogate build: ADMET/PAINS filter, HTVS, RF surrogate, active learning
├── 02_fragment_based_design/       FIQ-0 construction: fragment screen, linker validation, growing, production surrogate
├── 03_lead_candidate_docking/      FIQ-0 Vina docking, PLIP interaction profiling
├── 04_ligand_parameterization/     CGenFF topology and hydrogen-reconstruction fix
├── 05_system_setup/                solvation, ionization, assembled topology
├── 06_energy_minimization/         GROMACS inputs, log, and minimized coordinates
├── 07_analysis/                    potential-energy extraction and figure
├── 08_retrosynthesis/              Spaya routes (FIQ-0 one-step; FIQ-1 five-step)
├── colab/                          Colab notebook + inputs for the restrained MD
├── kaggle/                         Kaggle notebooks: restrained MD, MM/GBSA, Gnina/Smina consensus, M1 enrichment, M2 ML385 comparator
├── receptor/                       Nrf2-Neh1 (AlphaFold) PDBQT receptor used in Phases 01–03
├── future_directions/              production-MD control files and driver/analysis/MM-PBSA scripts
├── docs/                           consolidated command log + systems-pharmacology metrics
├── CITATION.cff
├── LICENSE
└── requirements.txt

Reproduction

git clone https://github.com/ahmetoztekin/nrf2-inhibitor-md-pipeline.git
cd nrf2-inhibitor-md-pipeline
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python 07_analysis/plot_potential_energy.py \
    --xvg 07_analysis/potential_energy.xvg \
    --out 07_analysis/figures/em_potential_energy.png

The potential-energy series is regenerated from the GROMACS energy file with gmx energy -f 06_energy_minimization/em.edr -o 07_analysis/potential_energy.xvg. The restrained MD and MM/GBSA are run from the notebooks in colab/ and kaggle/; full commands are in docs/PIPELINE.md.

Scope and limitations

  • All binding numbers are docking scores within AutoDock Vina / Vinardo error (~2–3 kcal/mol); absolute values are a coarse ranking, not measurements. The FIQ-1–ML385 comparison rests on ligand efficiency, not a reliable affinity gap.
  • 7X5G is a protein–DNA complex, not a ligand-bound structure; the docked poses and the proposed CNC-bZIP interference remain structural hypotheses. No inhibitory activity has been measured for either molecule.
  • The MD is limited (single 20 ns backbone-restrained trajectory per molecule): neither ligand holds its crystallographic pose in this shallow groove, so the loose association reflects the site, not a FIQ-1-specific failure.
  • MM/GBSA omits entropy and is an upper bound on favorability, reported as a relative same-protocol ranking.
  • The systems-pharmacology analysis validates the target, not FIQ-1.
  • Physicochemical, ADMET, and synthetic-accessibility values are computational predictions; neither molecule was synthesized.

Software

Phases 01–02: RDKit; scikit-learn; Meeko; AutoDock Vina 1.2; NumPy; pandas; matplotlib; seaborn; joblib.

Phases 03–08 + MD/MM-GBSA: AutoDock Vina 1.2; Smina (Vinardo); Gnina; PyMOL; CGenFF; CHARMM-GUI; GROMACS 2026; PLIP; gmx_MMPBSA; Spaya (Iktos); SwissADME.

Versions and full citations are listed in docs/PIPELINE.md.

Citation

See CITATION.cff.

Author

Ahmet Öztekin — ahmetoztekin712@gmail.com

License

MIT (see LICENSE).

About

A reproducible, end-to-end computational pipeline that uses a machine-learning docking surrogate and receptor-aware fragment growing to design a new Nrf2-inhibitor candidate, followed by force-field parametrization, explicit-solvent energy minimization, and AI-assisted retrosynthetic assessment.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages