StructureMoa is a chemical-structure exploration tool. Structural similarity, clustering, and SAR relationships are intended for research exploration and do not establish shared potency, selectivity, safety, efficacy, or clinical performance.
Data sources
Only sources this project actually uses. There is no paper-DOI ingest and no live FDA/EMA API.
- PubChem — curated CIDs, compound-page links, and name-based SMILES lookup when a structure is missing
- ChEMBL — curated ChEMBL IDs and links. Bioactivity tables are not synced.
- Literature — public-summary seed notes and family descriptions. A citation is shown only when a field exists; DOIs are not invented.
- ClinicalTrials.gov — curated search queries that open trial listings. Trial records are not imported.
- FDA / EMA — curated agency strings on selected clinical-overlay rows. Label APIs are not called.
- Manual seed data — families, compounds, SMILES, tags, and relations. Default dataProvenance is seed.
Entity types
Nodes are classified from catalog role plus identifiers and tags. Types are not forced when they do not fit.
- Observed molecules — Verified drug, Known compound, Literature compound; payload/linker/ADC with external IDs
- Abstractions — Abstract scaffold; Abstract pharmacophore only when tags/name include pharmacophore or warhead
- Exploratory structures — Rule-generated SAR analog, Exploratory analog
A parent node is not automatically a literature MCS or shared chemical core. Family centers are labeled as exploration scaffolds.
Structure standardization
- SMILES are assigned from seed data, the known-SMILES map, PubChem lookup, or parent inheritance.
- Inherited SMILES are not independently assigned structures.
- InChI Key, PubChem CID, and ChEMBL ID are shown only when present.
Fingerprinting
- Type: Morgan fingerprint (RDKit circular fingerprint)
- Radius: 2
- Bits: 2048
Similarity
- Metric: Tanimoto (shared bits / union bits on packed uint8 fingerprints)
- Weekly mesh threshold: 0.45
Seed similarity edges, the weekly mesh, add-compound estimates, and family-compare matrices use this configuration. Edges labeled Manual are the exception.
Structural similarity does not predict biological equivalence.
Clustering
Used in CNS chemometric analysis only. There is no catalog-wide Butina clustering.
- Algorithm: Butina-style sphere exclusion
- Cutoff: 0.55 (Tanimoto)
- Medoid: the member with the highest average similarity to other cluster members
- Approved-neighbor band: 0.45–0.75
SAR relationships
Family-map derived_from edges are manually curated parent→child paths. They are not claimed literature reaction schemes.
similar_to is fingerprint-derived, related_drug is a catalog drug-class link, and conjugated_with is a payload–linker–ADC conjugate relationship.
CNS seed proposals only mutate SMILES with the fixed rules below, then validate with RDKit. This is not a generative model.
cl-to-f— Cl→F halogen walkf-to-cl— F→Cl halogen walkome-to-oh— OMe→OH demethylationoh-to-ome— OH→OMe methylationnme-expand— N-H→N-Me (amine methyl)aryl-h-to-f— Aryl-H→F (c1ccccc1 walk)
Exploration metrics
On-screen High / Medium / Low labels and numeric scores are an Exploration Score. They are not drug quality, efficacy, clinical probability, or investment potential.
Internal exploration metric based on catalog and structural signals. Not a prediction of efficacy, safety, potency, or development success.
Documented criteria
- Core: family-center flag, scaffold/payload/linker role, SMILES, smilesSource, InChIKey, approved/clinical/SAR tags, family size, derived-child count, relation count, notes
- Derivative: drug/derivative/adc role, SMILES, smilesSource, approved/clinical tags, Morgan similarity to center, family size, relation count
Normalization
- Additive rule score, then capped at 100
- Relative percentile in the tab pool: top 12% High (X-S), 12–62% Medium (X-A/X-B), remainder Low (X-C). Pools under 5 compounds stay Medium
Limitations
- Scores follow catalog completeness and tag coverage.
- The similarity term is a structural fingerprint, not a potency model.
- Approved/clinical tags are catalog tags, not a live regulatory API.
- Computed proximity (similarity, clusters, exploration scores) is not presented as a literature fact.