Multi-Target Breast Cancer Activity Predictor

Enter a SMILES string to predict pIC50 and activity class across 53 validated breast cancer targets using ensemble ML (RF · XGB · LGB).

Try: Tamoxifen Lapatinib analogue Aspirin Caffeine
Running predictions across 53 targets…
Active Targets
pIC50 ≥ 7.0  ·  IC50 ≤ 100 nM
Moderate Targets
pIC50 5.0–7.0  ·  100 nM–10 µM
Inactive Targets
pIC50 < 5.0  ·  IC50 > 10 µM
Query Molecule
Physicochemical Properties
Activity Distribution
Top 15 Targets by Predicted pIC50
Download full prediction results
Detailed Predictions by Target
Target pIC50 IC50 Activity Confidence Dataset

Batch Prediction

Predict activity for multiple compounds simultaneously — one SMILES per line.

Processing batch…

Target Library

53 validated breast cancer targets · Training set statistics · Best-performing models per target

Target Total Active Moderate Inactive % Active Distribution Mean pIC50 Best Reg. Best Clf.
AKT1 3571 1671 1639 261 46.8%
6.9 XGB LGB
AKT2 1061 500 512 49 47.1%
6.92 LGB XGB
AKT3 304 159 128 17 52.3%
7.03 XGB SVM
ANDR 2292 627 1530 135 27.4%
6.43 XGB LGB
AROMATASE 1980 490 1200 290 24.7%
6.22 XGB XGB
ASK1 1535 1113 374 48 72.5%
7.43 XGB LGB
AURA 3027 1338 1501 188 44.2%
6.88 LGB XGB
BCL2 2431 1898 454 79 78.1%
7.82 RF LGB
BRAF 5555 3729 1614 212 67.1%
7.47 LGB XGB
BRD4 7813 2298 4085 1430 29.4%
6.19 XGB XGB
CA9 249 65 151 33 26.1%
6.23 LGB RF
CCND1 369 207 147 15 56.1%
7.09 LGB XGB
CCNE1 58 35 23 0 60.3%
7.23 LGB LGB
CDK4 3820 2101 1327 392 55.0%
6.98 XGB RF
CDK6 1898 985 845 68 51.9%
7.0 XGB LGB
CHK2 431 179 186 66 41.5%
6.42 RF LGB
CXCR4 910 447 433 30 49.1%
6.86 XGB XGB
EGFR 12377 6493 4893 991 52.5%
7.03 XGB XGB
ERBB4 243 147 89 7 60.5%
7.22 XGB XGB
FGFR 3220 1897 1154 169 58.9%
7.2 LGB RF
HDAC1 9811 3606 5385 820 36.8%
6.59 LGB XGB
HDAC6 7307 4163 2824 320 57.0%
6.99 XGB XGB
HER2 4041 2185 1566 290 54.1%
6.9 XGB LGB
HER3 57 31 16 10 54.4%
6.86 XGB XGB
JAK2 11273 6002 4488 783 53.2%
7.02 XGB LGB
KRAS 10417 5155 4661 601 49.5%
6.97 LGB XGB
LDHA 717 89 383 245 12.4%
5.76 RF LGB
MAPK14 4842 2403 2222 217 49.6%
6.95 LGB LGB
MDM2 4444 3201 1011 232 72.0%
7.61 LGB XGB
MMP14 667 155 383 129 23.2%
6.03 LGB SVM
MMP2 3257 1326 1274 657 40.7%
6.51 XGB LGB
MMP9 2081 1081 692 308 51.9%
6.88 XGB XGB
MTOR 4147 2533 1500 114 61.1%
7.25 XGB XGB
NFKB1 137 18 104 15 13.1%
5.9 RF RF
NOTC1 112 82 26 4 73.2%
7.69 RF RF
OEST_A 3911 1970 1576 365 50.4%
7.01 XGB XGB
OEST_B 1780 731 656 393 41.1%
6.43 XGB RF
PARP_1 4107 2534 1359 214 61.7%
7.19 LGB XGB
PARP_10 197 3 160 34 1.5%
5.51 SVM SVM
PD1L1 3842 2390 1206 246 62.2%
7.51 LGB XGB
PK3CA 7635 2744 4224 667 35.9%
6.57 XGB XGB
PPAR_GAMMA 1558 587 832 139 37.7%
6.54 XGB XGB
PROGESTERONE 1452 580 854 18 39.9%
6.8 XGB XGB
RASN 1147 289 846 12 25.2%
6.26 RF RF
S1PR1 987 586 360 41 59.4%
7.27 XGB RF
STAT3 1050 68 622 360 6.5%
5.43 XGB RF
TGFR1 2692 1658 972 62 61.6%
7.24 XGB XGB
TGFR2 427 46 349 32 10.8%
5.92 XGB XGB
TLR4 122 7 74 41 5.7%
5.45 LGB MLP
TNF_ALPHA 1349 128 866 355 9.5%
5.66 XGB RF
TP53 283 208 67 8 73.5%
8.35 LGB MLP
VGFR2 8627 3850 4262 515 44.6%
6.81 XGB XGB
WEE1 1019 713 288 18 70.0%
7.48 XGB XGB

BreastCAR — multi-target QSAR for breast-cancer discovery

A machine-learning platform that scores any small molecule against 53 validated breast-cancer targets in seconds, delivering both a continuous potency estimate (pIC50) and an Active / Moderate / Inactive class call.

53
Targets
106
Trained models
233,975
Bioactivities
5
Algorithms
2048
ECFP4 bits
1
SMILES
User input molecule
2
Featurise
Morgan FP + RDKit descriptors
3
Score
53 regressors + 53 classifiers
4
Rank
Top targets by predicted pIC50
5
Report
Table, chart, properties

Methodology

  • FeaturesMorgan (ECFP4) fingerprint — 2048 bits, radius 2 — concatenated with 13 physicochemical RDKit descriptors (MolLogP, MolMR, TPSA, HBA/HBD, RotBonds, RingCount, AromaticRings, FractionCSP3, HeavyAtomCount, HeteroAtoms, BalabanJ, BertzCT).
  • AlgorithmsRandom Forest, XGBoost, LightGBM, SVM, MLP — benchmarked per target; best performer retained.
  • ValidationStratified 80/20 split, 5-fold cross-validation, y-randomisation, and Bemis–Murcko scaffold split.
  • ApplicabilityTanimoto-similarity threshold vs the training set; out-of-domain predictions are flagged.
  • LabelsActive if pIC50 ≥ 7, Moderate if ≥ 5, else Inactive.

Data source

BindingDB — curated IC50 and Ki measurements against human targets implicated in breast-cancer biology: hormone receptors, cyclin-dependent kinases, DNA-damage response, growth-factor signalling, and the tumour micro-environment.

Records were standardised, salt-stripped, de-duplicated, and converted to pIC50. Final dataset: 233,975 bioactivity measurements across 53 targets.

API

POST /predict · POST /predict/batch · GET /targets · GET /health

How to use

  • SinglePaste one SMILES → predictions for all 53 targets ranked by pIC50, with 2D structure and drug-likeness.
  • BatchUpload a SMILES list, download the full result table.
  • ExploreBrowse training-set composition per target with the winning algorithm.

Disclaimer

BreastCAR provides computational predictions for research triage and hypothesis generation. Results are not a substitute for experimental validation. Predictions on compounds outside the applicability domain of the training data should be interpreted with caution.

Draw molecule

Sketch a structure — click Use SMILES to send it to the predictor.