Pharmacotranscriptomic profiles, which capture drug-induced changes in gene expression, offer vast potential for computational drug discovery and are widely used in modern medicine. However, current computational approaches neglected the associations within gene‒gene functional networks and unrevealed the systematic relationship between drug efficacy and the reversal effect. Here, we developed a new genome-scale functional module (GSFM) transformation framework to quantitatively evaluate drug efficacy for in silico drug discovery. GSFM employs four biologically interpretable quantifiers: GSFM_Up, GSFM_Down, GSFM_ssGSEA, and GSFM_TF to comprehensively evaluate the multi-dimension activities of each functional module (FM) at gene-level, pathway-level, and transcriptional regulatory network-level. Through a data transformation strategy, GSFM effectively converts noisy and potentially unreliable gene expression data into a more dependable FM active matrix, significantly outperforming other methods in terms of both robustness and accuracy. Besides, we found a positive correlation between RSGSFM and drug efficacy, suggesting that RSGSFM could serve as representative measure of drug efficacy. Furthermore, we identified WYE-354, perhexiline, and NTNCB as candidate therapeutic agents for the treatment of breast-invasive carcinoma, lung adenocarcinoma, and castration-resistant prostate cancer, respectively. The results from in vitro and in vivo experiments have validated that all identified compounds exhibit potent anti-tumor effects, providing proof-of-concept for our computational approach.
| a | The inclusion of a gene within gene set FM can be demonstrated by Eq. (3): |
| b | The exclusion of a gene from set FM can be demonstrated by Eq. (4): |
| c | Calculate the ssGSEA for each gene and determine the ssGSEA for this gene set based on the maximum difference (selecting the ssGSEA with the highest absolute value) using Eq. (5): |
| a | If esup and esdown have the same algebraic sign then CMapGSFM = 0. |
| b | CMapGSFM = esup – esdown. |
| (1) | Given a ranked list of drug-related FMs L with n FM and a disease-related FMs S with m FMs. |
| (2) | Given a vector V indicating the position (1, 2, …, n) of each module in L and sort the FMs in S in ascending order such that V (i) is the position of FM i, where i = 1, 2, …, t. Compute the following two values: |
| 科 Family | 属数 Number of genus | 种数 Number of species | 占总种数比例 Percentage of total species (%) | 属 Genus | 种数 Number of species | 占总种数比例 Percentage of total species (%) |
|---|---|---|---|---|---|---|
| 鹅膏菌科Amanitaceae | 2 | 11 | 5.26 | 鹅膏菌属 Amanita | 10 | 4.78 |
| 小菇科 Mycenaceae | 2 | 12 | 5.74 | 丝盖伞属 Inocybe | 5 | 2.39 |
| 多孔菌科 Polyporaceae | 8 | 14 | 6.70 | 蜡蘑属 Laccaria | 5 | 2.39 |
| 红菇科 Russulaceae | 3 | 23 | 11.00 | 小皮伞属 Marasmius | 6 | 2.87 |
| 小菇属 Mycena | 11 | 5.26 | ||||
| 光柄菇属 Pluteus | 5 | 2.39 | ||||
| 红菇属 Russula | 17 | 8.13 | ||||
| 栓菌属 Trametes | 5 | 2.39 |