Skip Navigation
Skip to contents

Cancer Res Treat : Cancer Research and Treatment

OPEN ACCESS

Articles

Page Path
HOME > Cancer Res Treat > Volume 58(2); 2026 > Article
Original Article
Hematologic malignancy
Uncovering Putative Causal Non-coding RNAs in Acute and Chronic Myeloid Leukemia: A Genome-Wide Mendelian Randomization Study
Sunwoo Jung1orcid, Ji-Won Kim2,3orcid, Buhm Han1,4,5orcid
Cancer Research and Treatment : Official Journal of Korean Cancer Association 2026;58(2):664-676.
DOI: https://doi.org/10.4143/crt.2025.377
Published online: June 25, 2025

1Interdisciplinary Program in Bioengineering, Seoul National University, Seoul, Korea

2Division of Hematology and Medical Oncology, Department of Internal Medicine, Seoul National University Bundang Hospital, Seoul National University College of Medicine, Seongnam, Korea

3Department of Genomic Medicine, Seoul National University Bundang Hospital, Seoul National University College of Medicine, Seongnam, Korea

4Department of Biomedical Sciences, BK21 Plus Biomedical Science Project, Seoul National University College of Medicine, Seoul, Korea

5Convergence Dementia Research Center, Seoul National University Medical Research Center, Seoul, Korea

Correspondence: Ji-Won Kim, Division of Hematology and Medical Oncology, Department of Internal Medicine, Seoul National University Bundang Hospital, 82 Gumi-ro, 173 beon-gil, Bundang-gu, Seongnam 13620, Korea
Tel: 82-31-787-7084 E-mail: jiwonkim@snubh.org
Co-correspondence: Buhm Han, Department of Biomedical Sciences, Seoul National University College of Medicine, Room 615, Basic Science Building, 103 Daehak-ro, Jongro-gu, Seoul 03080, Korea
Tel: 82-2-3668-7618 E-mail: buhm.han@snu.ac.kr
• Received: April 3, 2025   • Accepted: June 24, 2025

Copyright © 2026 by the Korean Cancer Association

This is an Open Access article distributed under the terms of the Creative Commons Attribution Non-Commercial License (http://creativecommons.org/licenses/by-nc/4.0/) which permits unrestricted non-commercial use, distribution, and reproduction in any medium, provided the original work is properly cited.

  • 4,297 Views
  • 116 Download
  • 1 Web of Science
  • 1 Crossref
prev next
  • Purpose
    Although non-coding RNAs (ncRNAs) have often been implicated in various cancers, their causal roles in acute (AML) and chronic myeloid leukemia (CML) remain unclear. Here, we conducted a genome-wide two-sample Mendelian randomization (MR) study to investigate the causal effects of a comprehensive set of ncRNAs on AML and CML.
  • Materials and Methods
    We used summary statistics of blood expression quantitative trait loci (eQTL) from the eQTLGen consortium (31,684 European participants) as exposure data. Genome-wide association study summary statistics from the FinnGen study (AML, 322 cases; CML, 303 cases) and UK Biobank (UKBB) (AML, 318 cases; CML, 153 cases) served as outcome data. The generalized inverse-variance weighted (GIVW) method was used as the primary MR method. Two additional MR methods (generalized MR-Egger and weighted median), sensitivity analyses, and the HEIDI test were further employed to support our findings.
  • Results
    Upregulated HCG22 and RP11-42I10.1 were causally linked to increased AML risk, while the GMDS-AS1 locus was positively associated with increased CML risk. We highlight these ncRNAs for their consistent significance across all three MR methods, with no evidence of bias in sensitivity analyses (F-test, Cochran’s Q-test, MR-Egger intercept, MR-PRESSO global test) and no indication of confounding from the HEIDI test. These findings were primarily discovered in FinnGen (FDRGIVW < 0.05), their significance was validated in UKBB (pGIVW < 0.05). Upon validation with an independent linkage disequilibrium reference panel, they remained robust.
  • Conclusion
    This study provides evidence of causal relationships between ncRNAs and the subtypes of AML and CML, notably highlighting HCG22 and RP11-42I10.1 in AML and the GMDS-AS1 locus in CML.
Accumulating evidence indicates that only a small fraction (up to 2%) of human transcripts are translated into functional proteins, whereas the majority give rise to non-coding RNAs (ncRNAs) [1]. Although earlier studies have highlighted their crucial role as gene expression regulators in various cancer types [2,3], their impact on leukemia remains largely unexplored.
Leukemia is a blood cancer characterized by the uncontrolled proliferation of abnormal blood cells. Although the advent of stem cell transplant techniques and targeted agents has recently improved treatment outcomes for leukemia, achieving a cure remains challenging, and the prognosis for relapsed patients is still dismal [4,5]. Leukemia is broadly classified into myeloid leukemia (ML) and lymphoid leukemia (LL) based on the affected blood cell lineage. ML is further classified into acute myeloid leukemia (AML) and chronic myeloid leukemia (CML), two distinct disease entities that require independent investigative approaches [6].
Growing evidence highlights the crucial roles of ncRNAs in leukemogenesis. Their dysregulation has closely been linked to leukemia development and progression [7]. Increasing efforts have been made to identify ncRNAs as potential biomarkers for leukemia. According to the review by Dieter et al. [8], several ncRNAs have been consistently reported across multiple studies as dysregulated in human leukemia samples. However, many ncRNAs have yet to be investigated in the context of leukemia. Most previous studies have examined gene expression differences between cases and controls through experimental comparisons, focusing on a limited number of genes. Yet, these approaches are constrained in their ability to determine causality and to comprehensively explore the extensive ncRNA landscape.
Distinguishing causality from association is challenging, as it requires random perturbations to establish causal relationships. Random perturbation experiments are often expensive and can only deal with a few targets at a time. Fortunately, genome-wide causality analysis has recently become feasible thanks to the Mendelian randomization (MR) strategy [9]. MR is a methodology to look for the causal effect of an exposure variable on an outcome variable by perturbing instrumental variables (IVs) strongly associated with the exposure variable. In MR, single nucleotide polymorphisms (SNPs) are used as IVs. As genotypes are randomly assorted during meiosis by Mendel’s law of inheritance, SNPs can be considered perturbed retrospectively. Despite several limitations such as strict assumption requirements, MR is considered as an effective approach that expands causality analysis to a genome-wide scale [10].
Myeloid lineage cells remarkably outnumber lymphoid lineage cells in human peripheral blood [11], providing a biologically relevant context for focusing on two subtypes of ML, especially when the available data are limited to the tissue-level. The availability of large-scale blood-derived expression quantitative trait loci (eQTL) summary data now facilitates comprehensive two-sample MR analyses to identify causal ncRNAs associated with either AML or CML. To the best of our knowledge, no previous study has explored the causal effects of ncRNAs on AML or CML at a genome-wide scale. In this genome-wide two-sample MR study, we unraveled causal relationships between a comprehensive set of ncRNAs and two ML subtypes, aiming to provide evidence of the leukemogenic potential of ncRNAs.
1. Study design
We conducted a genome-wide two-sample MR analysis to identify ncRNAs (exposure) causally linked to either AML or CML risk (outcome) (Fig. 1). Using the GENCODE v47 gene annotation, we identified ncRNAs from large-scale blood eQTL summary statistics accessed through the eQTLGen consortium. Genome-wide association study (GWAS) summary statistics for AML and CML were obtained from two large consortia: the FinnGen study and UK Biobank (UKBB), both serving as discovery and validation datasets. The generalized inverse-variance weighted (GIVW) method was used as the primary MR method, with GTEx v8 whole genome sequencing (WGS) data as the primary linkage disequilibrium (LD) reference panel. Sensitivity analyses and the heterogeneity in dependent instruments (HEIDI) test filtered out findings with potential bias or confounding. Two additional MR methods were utilized to strengthen our results: the generalized MR-Egger (GEgger) and weighted median (WM) methods. Those ncRNAs robustly identified in one cohort (either FinnGen or UKBB) were validated in the other. Ultimately, we highlighted ncRNAs only if they remained significant with the distinct LD reference panel (1000 Genomes European). Colocalization analysis and multivariable MR (MVMR) were also employed to provide additional supporting evidence. The study was conducted in accordance with the STROBE-MR reporting guidelines [12].
2. Exposure summary data
We obtained eQTL summary statistics for blood-expressed ncRNAs from the eQTLGen consortium, which includes data from 31,684 samples (25,482 whole blood samples and 6,202 peripheral blood mononuclear cell [PBMC] samples) of healthy European ancestry (S1 Table) [13]. To identify ncRNAs, we utilized comprehensive gene annotations provided by the GENCODE project (S2 Fig.) [14]. The eQTLGen consortium does not include eQTLs for genes on the X and Y chromosomes.
3. Outcome summary data
We obtained GWAS summary statistics for AML and CML from the FinnGen study and UKBB [15,16], both comprising participants of European ancestry (S1 Table). Specifically, the FinnGen study, confined to the Finnish population, includes 322 cases and 378,747 controls for AML and 303 cases and 378,748 controls for CML. Similarly, the UKBB, which focuses on the British population, includes 318 cases and 404,466 controls for AML and 153 cases and 404,466 controls for CML. The imbalance between the number of cases and controls was accounted for in each GWAS [17]. Confounding variables such as age and sex were also adjusted for during the association tests in each GWAS.
4. Instrument selection
For each ncRNA, we first retrieved all available cis-eQTLs (1 Mb upstream of the transcription start site to 1 Mb downstream of the end site) from the eQTLGen dataset. We then selected those with a minor allele frequency ≥ 1%. To ensure a certain level of independence among IVs, we applied LD clumping with an R2 threshold of < 0.1. We selected GTEx v8 WGS as our primary LD reference panel, since both the exposure and outcome populations in our study are of European ancestry, and GTEx v8 WGS includes a substantial number of European individuals (n=866) [18]. We additionally employed 1000 Genomes European WGS (n=503) as a secondary LD reference panel for validation [19]. Finally, we excluded eQTLs with an association p-value (PeQTL) greater than 1×10-6 (S3 Fig., S4 Table).
5. MR analysis
We utilized a two-sample MR approach, leveraging summary statistics from two independent datasets: one for the exposure (ncRNA expression) and another for the outcome (AML or CML status). From each dataset, we obtain genetic effect sizes for IV-exposure associations and IV-outcome associations. By applying a statistical method to those genetic associations, we can infer the existence and strength of causal effects between the exposure and the outcome. Two-sample MR offers key advantages, including flexibility in study design, computational efficiency for large-scale testing, and enhanced statistical power.
Before causal effect estimation, we harmonized our datasets to ensure that the genetic effects of each IV on the exposure and the outcome corresponded to the same allele. Additionally, we applied Steiger filtering to minimize the risk of reverse causation and enhance the validity of IVs (S3 Fig.) [20]. We discarded any IV that failed harmonization or Steiger filtering.
To robustly estimate the causal effect of each ncRNA on two ML subtypes, we adopted the GIVW method as our main statistical method [21]. This method was chosen for two key reasons: (1) to ensure sufficient statistical power and (2) to account for correlated IVs. Given the limited number of AML and CML cases in current GWAS studies from large consortia, an approach that maximizes statistical efficiency is essential. GIVW is well-suited for this purpose, as it enhances statistical power by aggregating information across multiple IVs while effectively accounting for correlations among the IVs via an LD matrix. According to Sadler et al. [22], incorporating an LD matrix within the MR framework helps safeguard MR estimates against potential biases due to a lenient clumping threshold (R2 < 0.1). We performed GIVW using the “MendelianRandomization v0.10.0” R package [23]. GTEx v8 WGS served as the primary LD reference panel, with 1000 Genomes European WGS as a secondary panel for validation. To reinforce our findings, we employed two additional MR methods: GEgger and WM. GEgger extends MR-Egger regression by estimating the causal effect while accounting for both horizontal pleiotropy and correlated IVs. WM is a robust estimator that provides a consistent causal effect estimate even when up to 50% of IVs are invalid. Both methods were implemented using the MendelianRandomization v0.10.0 R package. For the exposures with only a single IV, we used the Wald ratio estimator. We computed the Wald ratio estimator as implemented in the “TwoSampleMR v0.6.6” R package [24].
After applying GIVW with GTEx v8 WGS as the LD reference, ncRNAs with expression levels showing a significant causal effect (false discovery rate [FDR] corrected p-value (FDRGIVW) smaller than 0.05) on either AML or CML were primarily identified as putative causal.
6. Sensitivity analysis
We conducted sensitivity analyses to assess potential biases in our MR estimates. In particular, we employed the F-test, Cochran’s Q-test, MR-Egger intercept test, and MR-PRESSO global test.
To examine weak instrument bias within our causal estimates, we performed the F-test as implemented in the “MendelianRandomization v0.10.0” R package. Sanderson et al. [25] suggest that causal estimates could be biased if an F-statistic falls below 10, indicating the potential influence of weak instruments.
To assess heterogeneity among IVs, we conducted Cochran’s Q-test as implemented in the “MendelianRandomization v0.10.0” R package. The test statistic in Cochran’s Q-test quantifies the extent to which the MR estimates of individual IVs deviate from the overall causal estimate derived from all IVs.
To examine average pleiotropic effect across IVs, we performed the MR-Egger intercept test as implemented in the “MendelianRandomization v0.10.0” R package. In addition, we performed the MR-PRESSO global test using the “TwoSampleMR v0.6.6” R package (NbDistribution=2,000) to check for evidence of horizontal pleiotropy potentially driven by IVs acting as outliers.
Among the primarily identified ncRNAs, only those that met the criteria in all sensitivity analyses were considered in this study.
7. HEIDI test
We performed the HEIDI test as implemented in summary-data-based MR (SMR) software [26]. It distinguishes true causal effects from associations driven by LD. For the HEIDI test, we utilized GTEx v8 WGS and 1000 Genomes European WGS as reference panels in separate analyses.
Among the candidate ncRNAs, those that did not satisfy the HEIDI test were ignored.
8. Validation in an independent cohort
For ncRNAs robustly identified as causal in one cohort (either FinnGen or UKBB; FDRGIVW < 0.05), we validated them by assessing whether they maintained statistical significance (pGIVW < 0.05) and exhibited consistent effect directions in the other cohort.
9. Validation with an independent LD reference
We repeated the MR analysis, including LD clumping and LD matrix generation for GIVW, with 1000 Genomes European WGS as the LD reference panel. This allowed us to confirm whether candidate ncRNAs remained statistically significant (pGIVW < 0.05 in both FinnGen and UKBB) with consistent effect directions. Given our active use of LD-adjusted MR methods, we deemed it essential to confirm robustness using an independent LD reference panel, which prompted this additional validation step.
10. Colocalization analysis
To investigate whether the genetic associations for the exposure and outcome share a common causal variant at a given genetic locus, we conducted colocalization analysis using the “coloc v5.2.3” R package [27]. It evaluates five hypotheses (H0, H1, H2, H3, and H4). H0 assumes that no genetic variants at the locus are associated with both the exposure (ncRNA expression) and outcome (AML or CML). H1 suggests that variants at the locus are associated only with the exposure, while H2 suggests that the variants are associated only with the outcome. H3 assumes that distinct variants at the locus are associated with either the exposure or outcome, implying LD. H4 assumes that a common variant at the locus is associated with both the exposure and outcome, supporting colocalization.
Conventionally, a posterior probability for H4 (PP.H4) greater than 75%-80% is considered strong evidence supporting colocalization between the exposure and outcome. However, relying solely on PP.H4 may provide insufficient evidence for colocalization when there is a lack of strong genetic associations with the outcome at a specific locus. Such scenarios, characterized by high PP.H1 and low PP.H2, often reflect limited statistical power in the outcome data to detect colocalization. In these cases, calculating the colocalization probability conditional on the presence of a causal variant for the outcome, defined as PP.H4/(PP.H3+PP.H4), can provide a clearer indication of whether the exposure and outcome share the same causal variant, especially when IV-outcome associations are weak [28].
11. Multivariable MR
We performed MVMR using the “TwoSampleMR v0.6.6” R package to disentangle potential confounding by neighboring genes. For each candidate ncRNA, MVMR was conducted by including its four nearest neighboring genes (based on chromosomal coordinates) as additional exposures.
1. Identification of causal ncRNAs for AML
The MR pipeline is outlined in Fig. 1. Using the FinnGen dataset for discovery (Fig. 2A), eight ncRNAs were primarily identified as causal for AML (FDRGIVW < 0.05) (Fig. 2B and C). Of these, six (CTD-2517M22.14, HCG22, KRT10-AS1, LINC01505, RP11-42I10.1, and SPEN-AS1) satisfied all sensitivity analysis criteria (F ≥ 10, pQ-test > 0.05, pEgger_intercept > 0.05, and pPRESSO_global > 0.05) and met the HEIDI test threshold (pHEIDI > 0.05) (Table 1, S5 and S6 Tables). Notably, HCG22 and RP11-42I10.1 (also called LOC105371240) maintained PMR < 0.05 with both GEgger and WM (Table 1), and their significance was further validated in the UKBB dataset (Table 2, S7 Table). HCG22 consistently showed a risk effect on AML (elevated expression increases AML risk) across the two cohorts (odds ratio [OR] [95% confidence interval (CI)], 1.47 [1.33 to 1.62] and 1.26 [1.13 to 1.40] in FinnGen and UKBB, respectively) (Table 2). Likewise, RP11-42I10.1 consistently exhibited a risk effect on AML across the two cohorts (OR [95% CI], 3.87 [2.07 to 7.22] and 3.12 [1.39 to 6.98] in FinnGen and UKBB, respectively). Upon repeating the MR analysis with the independent LD reference (1000 Genomes EUR), both HCG22 and RP11-42I10.1 retained their significance (Table 2, S8 Table).
Using the UKBB dataset for discovery, nine ncRNAs were primarily identified as causal for AML (FDRGIVW < 0.05) (Fig. 2B and D). Of these, seven (AC007278.2, AC007278.3, HOTAIRM1, LINC02470, RP11-1398P2.1, RP11-799D4.4, and TWF2-DT) demonstrated no evidence of bias across all sensitivity analyses and no indication of confounding in the HEIDI test (Table 1, S5 and S6 Tables). Of note, this group included HOTAIRM1, a well-established cancer-relevant long ncRNA (lncRNA) [29,30]. RP11-1398P2.1 maintained pMR < 0.05 with two additional MR methods (Table 1). Unfortunately, its significance was not further validated in the FinnGen dataset (Table 2, S7 Table).
In summary, our results indicated that upregulation of HCG22 and RP11-42I10.1 was causally associated with increased AML risk. They remained robust following the primary MR method, sensitivity analyses, two additional MR methods, and validation using an independent cohort. Even in the repeated MR analysis using the independent LD reference, they were consistently significant.
2. Identification of causal ncRNAs for CML
The same MR pipeline was applied to the CML study (Fig. 1). Using the FinnGen dataset for discovery (Fig. 3A), eight ncRNAs were primarily identified as causal for CML (FDRGIVW < 0.05) (Fig. 3B and C). Among these, seven (AF131215.9, GATD1-DT, GMDS-AS1, RP11-173B14.4, RP11-362F19.1, RP11-830F9.6, and Y_RNA) remained robust, meeting all sensitivity analysis criteria and the HEIDI test threshold (Table 3, S5 and S6 Tables). Importantly, AF131215.9 and GMDS-AS1 (also called GMDS-DT) maintained pMR < 0.05 with both GEgger and WM (Table 3), and their significance was further validated in the UKBB dataset (Table 4, S7 Table). AF131215.9 consistently showed a protective effect on CML across the two cohorts (OR [95% CI], 0.77 [0.70 to 0.84] and 0.71 [0.61 to 0.82] in FinnGen and UKBB, respectively) (Table 4). In contrast, GMDS-AS1 consistently displayed a risk effect on CML across the two cohorts (OR [95% CI], 2.81 [1.75 to 4.51] and 1.85 [1.03 to 3.32] in FinnGen and UKBB, respectively). After repeating the MR analysis with the independent LD reference, only GMDS-AS1 retained its significance (Table 4, S8 Table).
Using the UKBB dataset for discovery, eight ncRNAs were primarily identified as causal for CML (FDRGIVW < 0.05) (Fig. 3B and D). Among these, five (CD27-AS1, FLVCR1-DT, HOTAIRM1, RBFADN, and RP11-514P8.2) showed no evidence of bias across all sensitivity analyses and no indication of confounding in the HEIDI test (Table 3, S5 and S6 Tables). Of note, this group included CD27-AS1, a lncRNA previously implicated in promoting AML progression [31]. HOTAIRM1, identified as causal for AML in this study, was also found to be causal for CML, consistently demonstrating a protective effect on both ML subtypes in the UKBB dataset. CD27-AS1, HOTAIRM1, and RBFADN maintained pMR < 0.05 with two additional MR methods (Table 3), but their significance was not further validated in the FinnGen dataset (Table 4, S7 Table).
In summary, our results suggested that upregulated GMDS-AS1 was causally linked to increased CML risk. It remained significant following the primary MR method, sensitivity analyses, two additional MR methods, and validation in an independent cohort. Even in the repeated MR analysis using the independent LD reference, it showed consistent significance.
3. Colocalization analysis
We conducted a colocalization analysis to determine whether each identified ncRNA and either AML or CML share common causal variants. Overall, our colocalization analysis based solely on PP.H4 did not yield strong evidence (S9 Table). This may be attributed to the small sample size for AML and CML cases, as well as limitations inherent in ncRNA study [32,33]. The remarkably high PP.H1 and low PP.H2 (close to zero) for most ncRNAs suggest that our outcome data are underpowered for colocalization analysis (S10 Fig.). To enhance interpretability of colocalization evidence, we instead used PP.H4/(PP.H3+PP.H4) ≥ 70% as an indicator of a likely colocalization of IV-exposure and IV-outcome associations [28,34,35]. In AML, PP.H4/(PP.H3+PP.H4) exceeded 70% for HCG22 (71.9% in FinnGen) and RP11-42I10.1 (85.9% in FinnGen) (S9 Table). In CML, PP.H4/(PP.H3+PP.H4) exceeded 70% for CD27-AS1 (90.1% in UKBB) (S9 Table).
In summary, HCG22 and RP11-42I10.1 showed colocalization evidence based on PP.H4/(PP.H3+PP.H4), whereas GMDS-AS1 did not. While the lack of colocalization evidence does not necessarily invalidate the causal association suggested by MR, it suggests caution in interpreting GMDS-AS1 as the causal gene. This may indicate that the GWAS associations for CML were markedly weak or unstable (the sample size for CML was even smaller than that of AML), or that the broader locus surrounding GMDS-AS1, rather than the gene itself, contributes to CML risk.
4. Multivariable Mendelian randomization
For the strong candidate ncRNAs (HCG22 and RP11-42I10.1 in AML, and GMDS-AS1 in CML) (S11 Table), we performed MVMR to disentangle the effects of nearby genes, particularly for HCG22 located in the complex major histocompatibility complex (MHC) region (S12 Fig.). In the MVMR analysis, we included four closest neighboring genes (based on chromosomal coordinates) and assessed whether the candidate gene retained its statistical significance. HCG22, RP11-42I10.1, and GMDS-AS1 remained highly significant in the MVMR analyses (pMVMR=4.38E-20, 3.06E-05, and 1.99E-08, respectively), with consistent effect directions (S13 Table).
In this MR study, we investigated the causal relationships between a comprehensive set of ncRNAs and two subtypes of ML. Using GWAS summary statistics from two independent cohorts, FinnGen and UKBB, we identified a set of putative causal ncRNAs for AML and CML. In AML, HCG22 and RP11-42I10.1 were ultimately highlighted as promising risk factors, demonstrating robust MR evidence and resistance to potential biases. They were primarily discovered in the FinnGen dataset and subsequently validated in the UKBB dataset. Their significance was further confirmed using the independent LD reference. They also showed colocalization evidence and remained significant in the MVMR analysis. In CML, the GMDS-AS1 locus was ultimately highlighted as an intriguing risk-associated region, consistently showing MR significance and robustness against potential biases. It was discovered in the FinnGen dataset and validated in the UKBB dataset. Its significance was further corroborated using the independent LD reference panel. GMDS-AS1 retained significance in the MVMR analysis but lacked colocalization evidence, warranting caution in pinpointing the ncRNA itself as the putative causal factor.
HCG22 (HLA complex group 22) is a lncRNA located within the human MHC on chromosome 6. Previous studies reported that HCG22 overexpression may exhibit an inhibitory effect in oral squamous cell carcinoma [36,37], while another study indicated that it may act as a tumor promoter in papillary thyroid cancer [38]. GMDS-AS1 is a lncRNA also located on chromosome 6 but outside the MHC region. An earlier study described that GMDS-AS1 overexpression may promote tumorigenesis in colorectal cancer [39], while another revealed that it may function as a tumor suppressor in lung adenocarcinoma [40]. For RP11-42I10.1, evidence of its involvement in cancers remains highly limited. We hope this study serves as an initial step toward uncovering the functional relevance of these ncRNAs.
To date, efforts have been made to explore the leukemogenic potential and clinical implications of ncRNAs in AML and CML. Most have focused on the differential expression of target ncRNAs between leukemia patients and healthy controls, using bone marrow, plasma, or serum samples [41,42]. Expression levels were quantified using quantitative reverse transcription–polymerase chain reaction or RNA-seq profiling to determine upregulation or downregulation in leukemia cases [42,43]. Some studies have also identified mutations linked to the dysregulated expression of target ncRNAs and investigated their functional characteristics, especially their roles in epigenetic modulation [44,45]. Additionally, predictive modeling and survival analysis were often performed to assess the prognostic significance of candidate ncRNAs [46,47]. The aforementioned analytical approaches, however, are limited in their capacity to investigate a broad range of ncRNAs simultaneously. While they provide valuable insights through functional analysis, they fall short of providing evidence to establish causality.
The main strength of this study lies in uncovering causal relationships between all available ncRNAs (according to the large-scale eQTL summary data) and two subtypes of ML through genome-wide MR. We focused our investigation on the two subtypes of ML, as myeloid lineage cells dominate circulating leukocytes [11]. Considering the high tissue specificity of ncRNAs and the relevance of our approach to leukemia research, we employed eQTL summary data derived from blood samples. GWAS summary data for AML and CML were obtained from the two large consortia, yet the number of AML and CML cases remains limited. To overcome this limitation, we used the GIVW method to ensure sufficient statistical power while mitigating bias from correlated IVs. MR results for ncRNAs are more straightforward to interpret compared to protein-coding genes when gene expression is used as the exposure, since they do not go through translation.
Therefore, beyond identifying ncRNAs as putative causal factors in AML and CML, our study carries several clinical implications. First, their consistent MR evidence across cohorts supports their potential as transcriptomic biomarkers, particularly given their detectability in peripheral blood. Second, if future studies elucidate their roles in leukemogenesis, these ncRNAs may serve as therapeutic targets—for example, via antisense oligonucleotides or RNA-based therapies. Finally, this study presents a scalable framework for prioritizing ncRNAs using genetic data, applicable to other malignancies for uncovering overlooked contributors from the non-coding transcriptome. As interest in the regulatory roles of ncRNAs grows, our results offer a focused entry point for translating non-coding genetic signals into meaningful biological insight and clinical utility.
This study has limitations that were beyond our ability to address. The ncRNA annotation remains incomplete, indicating that many ncRNAs still could not be explored in our analyses. Most ncRNAs analyzed in this study were lncRNAs, largely due to the availability of more comprehensive annotations for lncRNAs compared to other ncRNA classes (S14 Table). Compared to protein-coding genes, eQTL signals are generally weaker for ncRNAs due to their low expression levels, small regulatory networks, and biological noise (e.g. context-specific roles) [32]. We believe that these limitations can be mitigated by advancements in sequencing technology and improvements in the ncRNA annotation. Despite its severity and clinical significance, leukemia is a rare cancer type. Given the highly constrained sample size for AML and CML, we conducted the most appropriate MR analyses to the best of our ability. We believe that increasing the number of participants in future AML and CML GWAS studies would improve the power and precision of this MR analysis. Most importantly, accumulating evidence indicates that leukemia originates from leukemic stem cells (LSCs), rather than from differentiated cells [48,49]. Since LSCs primarily reside in the bone marrow, using bone marrow-specific gene expression data will be the most relevant for leukemia research. However, due to ethical and technical challenges, obtaining bone marrow expression paired with genotype information for a sufficiently large number of individuals is currently unfeasible. We therefore utilized eQTL data from whole blood and PBMCs as the closest available proxy. Notably, Sakhinia et al. [50] showed that gene expression profiles of peripheral blood and bone marrow were similar for 10 out of 15 AML indicator genes, suggesting that peripheral blood can partially reflect bone marrow expression. As more bone marrow-derived expression data accumulate in the future, more precise analyses will become feasible. Lastly, from a statistical perspective, the ncRNAs identified in this study demonstrated strong potential to contribute to leukemogenesis. However, their specific mechanisms remain unclear. Further sophisticated functional and clinical studies are necessary to clarify their underlying biological mechanisms.
In conclusion, we identified a set of candidate ncRNAs causally associated with either AML or CML risk. Notably, upregulated HCG22 and RP11-42I10.1 were strongly linked to increased AML risk, while the GMDS-AS1 locus was positively associated with increased CML risk.
Supplementary materials are available at Cancer Research and Treatment website (https://www.e-crt.org).

Ethical Statement

All summary statistics used in our MR analyses were sourced from previously published studies, each of which obtained the required ethical approvals and informed consent. The relevant studies are cited within this article.

Author Contributions

Conceived and designed the analysis: Jung S.

Collected the data: Jung S, Han B.

Contributed data or analysis tools: Jung S, Han B.

Performed the analysis: Jung S.

Wrote the paper: Jung S, Kim JW, Han B.

Interpreted the results: Jung S, Kim JW, Han B.

Conflict of Interest

Buhm Han is the CEO of SpintoAI Inc. Ji-Won Kim reports research support paid to his institution from Debiopharm, MitoImmune, Pyramid Biosciences, Adlai Nortye, Merck Sharp and Dohme, AstraZeneca, Lilly, and Aslan; participation in a Data Safety Monitoring Board or Advisory Board for MedPacto and Lilly.

Funding

This work was supported by the National Research Foundation of Korea (NRF) (Grant number 2022R1A2B5B02001897) funded by the Korean government, Ministry of Science, and ICT. This work was also supported by the Creative-Pioneering Researchers Program funded by Seoul National University and by the AI-Bio Research Grant through Seoul National University. Lastly, this work was supported by a grant from the Seoul National University Bundang Hospital Research Fund (No. 02-2020-017).

Fig. 1.
Graphical abstract of the study design. The contents within the dashed rectangle represent external validation components. eQTL, expression quantitative trait loci; FDR, false discovery rate; GIVW, Generalized Inverse-variance Weighted; HEIDI, heterogeneity in dependent instruments; LD, linkage disequilibrium; lncRNA, long non-coding RNA; MAF, minor allele frequency; miRNA, microRNA; MR, Mendelian Randomization; MR-Egger, Mendelian Randomization–Egger regression; scaRNA, small cajal body-specific RNA; snoRNA, small nucleolar RNA; snRNA, small nuclear RNA; UKBB, UK Biobank; WGS, whole genome sequencing.
crt-2025-377f1.jpg
Fig. 2.
Primary Mendelian randomization (MR) analysis between non-coding RNA (ncRNA) expression and acute myeloid leukemia (AML). (A) The total number of ncRNAs assessed for causal relationships with AML in each cohort. (B) The number of ncRNAs primarily identified as causal for AML (FDRGIVW < 0.05) in the FinnGen (blue) and UK Biobank (UKBB) (red) cohorts. (C) Volcano plot of MR results for AML in the FinnGen dataset. The horizontal dashed line (grey) indicates the significance threshold for MR estimates (FDRGIVW < 0.05). The vertical dashed line (grey) represents an odds ratio of 1. (D) Volcano plot of MR results for AML in the UKBB dataset. FDR, false discovery rate.
crt-2025-377f2.jpg
Fig. 3.
Primary Mendelian randomization (MR) analysis between non-coding RNA (ncRNA) expression and chronic myeloid leukemia (CML). (A) The total number of ncRNAs assessed for causal relationships with CML in each cohort. (B) The number of ncRNAs primarily identified as causal for CML (FDRGIVW < 0.05) in the FinnGen (blue) and UK Biobank (UKBB) (red) cohorts. (C) Volcano plot of MR results for CML in the FinnGen dataset. The horizontal dashed line (grey) indicates the significance threshold for MR estimates (FDRGIVW < 0.05). The vertical dashed line (grey) represents an odds ratio of 1. (D) Volcano plot of MR results for CML in the UKBB dataset. FDR, false discovery rate.
crt-2025-377f3.jpg
Table 1.
MR results for putative causal ncRNAs in AML
Discovery Gene name CHRa) #IVb) Method OR (95% CI) p-valuec)
FinnGen CTD-2517M22.14 8 19 GIVW 0.71 (0.60-0.84) 8.24E-05
Weighted median 0.88 (0.68-1.13) 3.04E-01
GEgger 0.59 (0.39-0.90) 1.31E-02
HCG22 6 22 GIVW 1.47 (1.33-1.62) 1.46E-14
Weighted median 1.46 (1.17-1.83) 7.42E-04
GEgger 1.38 (1.04-1.85) 2.82E-02
KRT10-AS1 17 38 GIVW 0.64 (0.52-0.79) 4.18E-05
Weighted median 0.62 (0.47-0.81) 5.94E-04
GEgger 0.83 (0.57-1.20) 3.16E-01
LINC01505 9 17 GIVW 0.68 (0.56-0.83) 1.21E-04
Weighted median 0.70 (0.52-0.94) 1.75E-02
GEgger 0.84 (0.53-1.34) 4.68E-01
RP11-42I10.1 16 3 GIVW 3.87 (2.07-7.22) 2.11E-05
Weighted median 4.15 (1.88-9.14) 4.16E-04
GEgger 4.60 (1.22-17.28) 2.39E-02
SPEN-AS1 1 24 GIVW 1.53 (1.26-1.86) 2.25E-05
Weighted median 1.52 (1.03-2.25) 3.72E-02
GEgger 1.48 (0.97-2.26) 6.67E-02
UKBB AC007278.2 2 33 GIVW 0.83 (0.77-0.89) 8.59E-07
Weighted median 0.82 (0.66-1.02) 7.69E-02
GEgger 0.76 (0.66-0.87) 4.56E-05
AC007278.3 2 68 GIVW 0.92 (0.90-0.95) 1.88E-07
Weighted median 0.90 (0.79-1.02) 9.23E-02
GEgger 0.98 (0.90-1.07) 6.99E-01
HOTAIRM1 7 62 GIVW 0.92 (0.89-0.96) 6.72E-05
Weighted median 0.92 (0.82-1.04) 1.88E-01
GEgger 0.90 (0.83-0.98) 1.81E-02
LINC02470 12 64 GIVW 0.79 (0.70-0.89) 5.89E-05
Weighted median 0.89 (0.76-1.03) 1.27E-01
GEgger 0.80 (0.70-0.92) 1.16E-03
RP11-1398P2.1 4 46 GIVW 0.73 (0.64-0.83) 2.60E-06
Weighted median 0.74 (0.58-0.94) 1.50E-02
GEgger 0.76 (0.60-0.97) 3.05E-02
RP11-799D4.4 17 10 GIVW 1.51 (1.27-1.81) 4.67E-06
Weighted median 1.49 (1.07-2.07) 1.88E-02
GEgger 1.19 (0.64-2.20) 5.81E-01
TWF2-DT 3 21 GIVW 1.42 (1.22-1.64) 2.46E-06
Weighted median 1.50 (0.94-2.40) 9.07E-02
GEgger 0.86 (0.50-1.45) 5.62E-01

The listed ncRNAs were causally linked to AML risk, with FDRGIVW < 0.05 and no evidence of bias across all sensitivity analyses. AML, acute myeloid leukemia; CI, confidence interval; FDR, false discovery rate; GEgger, generalized MR-Egger; GIVW, generalized inverse-variance weighted; MR, Mendelian randomization; ncRNA, non-coding RNA; OR, odds ratio; UKBB, UK Biobank.

a) CHR, the chromosome on which the corresponding ncRNA is located,

b) #IV, the number of instrumental variables used for a corresponding ncRNA,

c) p-values represent nominal p-values before applying FDR correction.

Table 2.
External validation for AML associated causal ncRNAs
Gene name Dataset GTEx v8 (main analysis)
1000 Genomes EUR
ORa) (95% CI) p-valueb) OR (95% CI) p-value
HCG22 FinnGen 1.47 (1.33-1.62) 1.46E-14 1.47 (1.30-1.65) 2.60E-10
UKBB 1.26 (1.13-1.40) 5.07E-05 1.13 (1.02-1.26) 0.017
RP11-42I10.1 FinnGen 3.87 (2.07-7.22) 2.11E-05 4.00 (1.52-10.5) 4.91E-03
UKBB 3.12 (1.39-6.98) 5.65E-03 3.50 (1.31-9.34) 0.012
RP11-1398P2.1 FinnGen 1.02 (0.91-1.14) 7.35E-01 0.98 (0.89-1.10) 7.80E-01
UKBB 0.73 (0.64-0.83) 2.60E-06 0.74 (0.67-0.82) 1.88E-08

Each ncRNA showing a consistent causal relationship with AML across all MR methods in one cohort (either FinnGen or UKBB) was validated in the other. Also, MR analyses were repeated using the 1000 Genomes European LD reference panel for ncRNAs identified as promising candidates in the main analyses. AML, acute myeloid leukemia; CI, confidence interval; FDR, false discovery rate; GIVW, generalized inverse-variance weighted; LD, linkage disequilibrium; MR, Mendelian randomization; ncRNA, non-coding RNA; UKBB, UK Biobank.

a) OR, odds ratio estimated using GIVW,

b) p-values represent nominal p-values (pGIVW) prior to FDR correction.

Table 3.
MR results for putative causal ncRNAs in CML
Discovery Gene name CHRa) #IVb) Method OR (95% CI) p-valuec)
FinnGen AF131215.9 8 157 GIVW 0.77 (0.70-0.84) 2.44E-08
Weighted median 0.86 (0.78-0.95) 2.22E-03
GEgger 0.75 (0.63-0.91) 3.19E-03
GATD1-DT 11 57 GIVW 0.88 (0.83-0.93) 1.64E-05
Weighted median 0.93 (0.79-1.10) 3.97E-01
GEgger 0.94 (0.83-1.06) 3.17E-01
GMDS-AS1 6 9 GIVW 2.81 (1.75-4.51) 1.80E-05
Weighted median 2.68 (1.66-4.32) 5.23E-05
GEgger 2.16 (1.07-4.37) 3.11E-02
RP11-173B14.4 13 5 GIVW 1.99 (1.42-2.79) 5.57E-05
Weighted median 2.10 (1.31-3.36) 2.01E-03
GEgger 1.25 (0.17-9.24) 8.30E-01
RP11-362F19.1 4 53 GIVW 1.16 (1.09-1.24) 1.08E-06
Weighted median 1.17 (1.00-1.36) 5.01E-02
GEgger 1.19 (1.06-1.34) 4.11E-03
RP11-830F9.6 16 8 GIVW 0.36 (0.23-0.57) 1.30E-05
Weighted median 0.46 (0.26-0.81) 7.58E-03
GEgger 0.90 (0.22-3.65) 8.85E-01
Y_RNA 9 11 GIVW 0.70 (0.59-0.82) 1.49E-05
Weighted median 0.81 (0.57-1.13) 2.15E-01
GEgger 0.97 (0.48-1.93) 9.25E-01
UKBB CD27-AS1 12 35 GIVW 1.71 (1.29-2.25) 1.74E-04
Weighted median 1.86 (1.41-2.45) 1.24E-05
GEgger 2.00 (1.45-2.77) 2.82E-05
FLVCR1-DT 1 94 GIVW 1.20 (1.14-1.26) 8.71E-13
Weighted median 1.20 (0.97-1.49) 8.77E-02
GEgger 1.27 (1.16-1.39) 2.93E-07
HOTAIRM1 7 62 GIVW 0.87 (0.82-0.92) 2.52E-06
Weighted median 0.83 (0.70-0.99) 3.61E-02
GEgger 0.81 (0.71-0.92) 8.73E-04
RBFADN 18 19 GIVW 0.69 (0.60-0.80) 6.19E-07
Weighted median 0.72 (0.55-0.93) 1.32E-02
GEgger 0.68 (0.50-0.94) 1.93E-02
RP11-514P8.2 7 22 GIVW 1.44 (1.19-1.75) 1.71E-04
Weighted median 1.33 (0.93-1.91) 1.19E-01
GEgger 1.54 (1.01-2.36) 4.54E-02

The listed ncRNAs were causally linked to CML risk, with FDRGIVW < 0.05 and no evidence of bias across all sensitivity analyses. CI, confidence interval; CML, chronic myeloid leukemia; FDR, false discovery rate; GIVW, generalized inverse-variance weighted; MR, Mendelian randomization; ncRNA, non-coding RNA; OR, odds ratio; UKBB, UK Biobank.

a) CHR, the chromosome on which the corresponding ncRNA is located,

b) #IV, the number of instrumental variables used for a corresponding ncRNA,

c) p-values represent nominal p-values before applying FDR correction.

Table 4.
External validation for CML associated causal ncRNAs
Gene name Dataset GTEx v8 (main analysis)
1000 Genomes EUR
ORa) (95% CI) p-valueb) OR (95% CI) p-value
AF131215.9 FinnGen 0.77 (0.70-0.84) 2.44E-08 0.92 (0.85-1.00) 4.86E-02
UKBB 0.71 (0.61-0.82) 7.03E-06 0.94 (0.83-1.06) 0.29
GMDS-AS1 FinnGen 2.81 (1.75-4.51) 1.80E-05 2.22 (1.42-3.46) 4.68E-04
UKBB 1.85 (1.03-3.32) 4.02E-02 2.06 (1.15-3.69) 0.014
CD27-AS1 FinnGen 0.95 (0.78-1.15) 5.85E-01 0.97 (0.79-1.19) 0.75
UKBB 1.71 (1.29-2.25) 1.74E-04 1.47 (1.04-2.06) 2.81E-02
HOTAIRM1 FinnGen 1.01 (0.97-1.05) 7.91E-01 1.00 (0.95-1.04) 0.84
UKBB 0.87 (0.82-0.92) 2.52E-06 0.85 (0.79-0.92) 7.05E-05
RBFADN FinnGen 0.97 (0.87-1.07) 5.25E-01 1.01 (0.83-1.23) 0.92
UKBB 0.69 (0.60-0.80) 6.19E-07 0.82 (0.62-1.08) 1.63E-01

Each ncRNA showing a consistent causal relationship with CML across all MR methods in one cohort (either FinnGen or UKBB) was validated in the other. Also, MR analyses were repeated using the 1000 Genomes European LD reference panel for ncRNAs identified as promising candidates in the main analyses. CI, confidence interval; CML, chronic myeloid leukemia; FDR, false discovery rate; LD, linkage disequilibrium; MR, Mendelian randomization; ncRNA, non-coding RNA; UKBB, UK Biobank.

a) OR, odds ratio estimated using GIVW,

b) p-values represent nominal p-values (pGIVW) prior to FDR correction.

  • 1. Kaikkonen MU, Lam MT, Glass CK. Non-coding RNAs as regulators of gene expression and epigenetics. Cardiovasc Res. 2011;90:430–40. ArticlePubMedPMC
  • 2. Liu Y, Liu X, Lin C, Jia X, Zhu H, Song J, et al. Noncoding RNAs regulate alternative splicing in cancer. J Exp Clin Cancer Res. 2021;40:11.ArticlePubMedPMCPDF
  • 3. Xu Y, Qiu M, Shen M, Dong S, Ye G, Shi X, et al. The emerging regulatory roles of long non-coding RNAs implicated in cancer metabolism. Mol Ther. 2021;29:2209–18. ArticlePubMedPMC
  • 4. Osman AE, Deininger MW. Chronic myeloid leukemia: modern therapies, current challenges and future directions. Blood Rev. 2021;49:100825.ArticlePubMedPMC
  • 5. Nadiminti KV, Sahasrabudhe KD, Liu H. Menin inhibitors for the treatment of acute myeloid leukemia: challenges and opportunities ahead. J Hematol Oncol. 2024;17:113.ArticlePubMedPMCPDF
  • 6. Houshmand M, Simonetti G, Circosta P, Gaidano V, Cignetti A, Martinelli G, et al. Chronic myeloid leukemia stem cells. Leukemia. 2019;33:1543–56. ArticlePubMedPMCPDF
  • 7. Bhat AA, Younes SN, Raza SS, Zarif L, Nisar S, Ahmed I, et al. Role of non-coding RNA networks in leukemia progression, metastasis and drug resistance. Mol Cancer. 2020;19:57.ArticlePubMedPMCPDF
  • 8. Dieter C, Lourenco ED, Lemos NE. Association of long non-coding RNA and leukemia: a systematic review. Gene. 2020;735:144405.ArticlePubMed
  • 9. Lawlor DA, Harbord RM, Sterne JA, Timpson N, Davey Smith G. Mendelian randomization: using genes as instruments for making causal inferences in epidemiology. Stat Med. 2008;27:1133–63. ArticlePubMed
  • 10. Smith GD, Ebrahim S. Mendelian randomization: prospects, potentials, and limitations. Int J Epidemiol. 2004;33:30–42. ArticlePubMed
  • 11. Gabrilovich DI, Ostrand-Rosenberg S, Bronte V. Coordinated regulation of myeloid cells by tumours. Nat Rev Immunol. 2012;12:253–68. ArticlePubMedPMCPDF
  • 12. Skrivankova VW, Richmond RC, Woolf BA, Davies NM, Swanson SA, VanderWeele TJ, et al. Strengthening the reporting of observational studies in epidemiology using mendelian randomisation (STROBE-MR): explanation and elaboration. BMJ. 2021;375:n2233.ArticlePubMedPMC
  • 13. Vosa U, Claringbould A, Westra HJ, Bonder MJ, Deelen P, Zeng B, et al. Large-scale cis- and trans-eQTL analyses identify thousands of genetic loci and polygenic scores that regulate blood gene expression. Nat Genet. 2021;53:1300–10. PubMedPMC
  • 14. Frankish A, Carbonell-Sala S, Diekhans M, Jungreis I, Loveland JE, Mudge JM, et al. GENCODE: reference annotation for the human and mouse genomes in 2023. Nucleic Acids Res. 2023;51:D942–d9. ArticlePubMedPMCPDF
  • 15. Kurki MI, Karjalainen J, Palta P, Sipila TP, Kristiansson K, Donner KM, et al. FinnGen provides genetic insights from a well-phenotyped isolated population. Nature. 2023;613:508–18. PubMedPMC
  • 16. Bycroft C, Freeman C, Petkova D, Band G, Elliott LT, Sharp K, et al. The UK Biobank resource with deep phenotyping and genomic data. Nature. 2018;562:203–9. ArticlePubMedPMCPDF
  • 17. Zhou W, Nielsen JB, Fritsche LG, Dey R, Gabrielsen ME, Wolford BN, et al. Efficiently controlling for case-control imbalance and sample relatedness in large-scale genetic association studies. Nat Genet. 2018;50:1335–41. ArticlePubMedPMCPDF
  • 18. GTEx Consortium. The Genotype-Tissue Expression (GTEx) project. Nat Genet. 2013;45:580–5. PubMedPMC
  • 19. Auton A, Brooks LD, Durbin RM, Garrison EP, Kang HM, Korbel JO, et al. A global reference for human genetic variation. Nature. 2015;526:68–74. PubMedPMC
  • 20. Hemani G, Tilling K, Davey Smith G. Orienting the causal relationship between imprecisely measured traits using GWAS summary data. PLoS Genet. 2017;13:e1007081ArticlePubMedPMC
  • 21. Burgess S, Dudbridge F, Thompson SG. Combining information on multiple instrumental variables in Mendelian randomization: comparison of allele score and summarized data methods. Stat Med. 2016;35:1880–906. ArticlePubMedPDF
  • 22. Sadler MC, Auwerx C, Lepik K, Porcu E, Kutalik Z. Quantifying the role of transcript levels in mediating DNA methylation effects on complex traits and diseases. Nat Commun. 2022;13:7559.ArticlePubMedPMCPDF
  • 23. Yavorska OO, Burgess S. MendelianRandomization: an R package for performing Mendelian randomization analyses using summarized data. Int J Epidemiol. 2017;46:1734–9. ArticlePubMedPMC
  • 24. Hemani G, Zheng J, Elsworth B, Wade KH, Haberland V, Baird D, et al. The MR-Base platform supports systematic causal inference across the human phenome. Elife. 2018;7:e34408PubMedPMC
  • 25. Sanderson E, Spiller W, Bowden J. Testing and correcting for weak and pleiotropic instruments in two-sample multivariable Mendelian randomization. Stat Med. 2021;40:5434–52. ArticlePubMedPMCPDF
  • 26. Zhu Z, Zhang F, Hu H, Bakshi A, Robinson MR, Powell JE, et al. Integration of summary data from GWAS and eQTL studies predicts complex trait gene targets. Nat Genet. 2016;48:481–7. ArticlePubMedPDF
  • 27. Wallace C. A more accurate method for colocalisation ana-lysis allowing for multiple causal variants. PLoS Genet. 2021;17:e1009440ArticlePubMedPMC
  • 28. Zuber V, Grinberg NF, Gill D, Manipur I, Slob EA, Patel A, et al. Combining evidence from Mendelian randomization and colocalization: review and comparison of approaches. Am J Hum Genet. 2022;109:767–82. ArticlePubMedPMC
  • 29. Ren T, Hou J, Liu C, Shan F, Xiong X, Qin A, et al. The long non-coding RNA HOTAIRM1 suppresses cell progression via sponging endogenous miR-17-5p/B-cell translocation gene 3 (BTG3) axis in 5-fluorouracil resistant colorectal cancer cells. Biomed Pharmacother. 2019;117:109171.ArticlePubMed
  • 30. Ahmadov U, Picard D, Bartl J, Silginer M, Trajkovic-Arsic M, Qin N, et al. The long non-coding RNA HOTAIRM1 promotes tumor aggressiveness and radiotherapy resistance in glioblastoma. Cell Death Dis. 2021;12:885.ArticlePubMedPMCPDF
  • 31. Tao Y, Zhang J, Chen L, Liu X, Yao M, Zhang H. LncRNA CD27-AS1 promotes acute myeloid leukemia progression through the miR-224-5p/PBX3 signaling circuit. Cell Death Dis. 2021;12:510.ArticlePubMedPMCPDF
  • 32. Mattick JS, Amaral PP, Carninci P, Carpenter S, Chang HY, Chen LL, et al. Long non-coding RNAs: definitions, functions, challenges and recommendations. Nat Rev Mol Cell Biol. 2023;24:430–47. ArticlePubMedPMCPDF
  • 33. Gawronski KA, Kim J. Single cell transcriptomics of noncoding RNAs and their cell-specificity. Wiley Interdiscip Rev RNA. 2017;8:e1433
  • 34. Huang A, Wu X, Lin J, Wei C, Xu W. Genetic insights into repurposing statins for hyperthyroidism prevention: a drug-target Mendelian randomization study. Front Endocrinol (Lausanne). 2024;15:1331031.ArticlePubMedPMC
  • 35. Ying H, Wu X, Jia X, Yang Q, Liu H, Zhao H, et al. Single-cell transcriptome-wide Mendelian randomization and colocalization reveals immune-mediated regulatory mechanisms and drug targets for COVID-19. EBioMedicine. 2025;113:105596.ArticlePubMedPMC
  • 36. Yin J, Zeng X, Ai Z, Yu M, Wu Y, Li S. Construction and analysis of a lncRNA-miRNA-mRNA network based on competitive endogenous RNA reveal functional lncRNAs in oral cancer. BMC Med Genomics. 2020;13:84.ArticlePubMedPMCPDF
  • 37. Wang M, Feng Z, Li X, Sun S, Lu L. Assessment of multiple pathways involved in the inhibitory effect of HCG22 on oral squamous cell carcinoma progression. Mol Cell Biochem. 2021;476:2561–71. ArticlePubMedPDF
  • 38. Cao X, Ma C, Wu Y, Huang J. lncRNA HCG22 regulated cell growth and metastasis of papillary thyroid cancer via negatively modulating miR-425-5p. Endokrynol Pol. 2024;75:20–6. ArticlePubMed
  • 39. Ye D, Liu H, Zhao G, Chen A, Jiang Y, Hu Y, et al. LncGMDS-AS1 promotes the tumorigenesis of colorectal cancer through HuR-STAT3/Wnt axis. Cell Death Dis. 2023;14:165.ArticlePubMedPMCPDF
  • 40. Zhao M, Xin XF, Zhang JY, Dai W, Lv TF, Song Y. LncRNA GMDS-AS1 inhibits lung adenocarcinoma development by regulating miR-96-5p/CYLD signaling. Cancer Med. 2020;9:1196–208. ArticlePubMedPDF
  • 41. Priya, Garg M, Talwar R, Bharadwaj M, Ruwali M, Pandey AK. Clinical relevance of long non-coding RNA in acute myeloid leukemia: a systematic review with meta-analysis. Leuk Res. 2024;147:107595.ArticlePubMed
  • 42. Solomon Y, Berhan A, Almaw A, Ersino T, Damtie S, Kiros T, et al. Long non-coding RNA as potential diagnostic markers for acute myeloid leukemia: a systematic review and meta-analysis. Cancer Med. 2024;13:e7376PubMedPMC
  • 43. Mishra S, Liu J, Chai L, Tenen DG. Diverse functions of long noncoding RNAs in acute myeloid leukemia: emerging roles in pathophysiology, prognosis, and treatment resistance. Curr Opin Hematol. 2022;29:34–43. ArticlePubMedPMC
  • 44. Gao S, Zhou B, Li H, Huang X, Wu Y, Xing C, et al. Long noncoding RNA HOTAIR promotes the self-renewal of leukemia stem cells through epigenetic silencing of p15. Exp Hematol. 2018;67:32–40. ArticlePubMed
  • 45. Roy L, Chatterjee O, Bose D, Roy A, Chatterjee S. Noncoding RNA as an influential epigenetic modulator with promising roles in cancer therapeutics. Drug Discov Today. 2023;28:103690.ArticlePubMed
  • 46. Tsai CH, Yao CY, Tien FM, Tang JL, Kuo YY, Chiu YC, et al. Incorporation of long non-coding RNA expression profile in the 2017 ELN risk classification can improve prognostic prediction of acute myeloid leukemia patients. EBioMedicine. 2019;40:240–50. ArticlePubMedPMC
  • 47. Wang Y, Li Y, Song HQ, Sun GW. Long non-coding RNA LINC00899 as a novel serum biomarker for diagnosis and prognosis prediction of acute myeloid leukemia. Eur Rev Med Pharmacol Sci. 2018;22:7364–70. PubMed
  • 48. Shlush LI, Zandi S, Mitchell A, Chen WC, Brandwein JM, Gupta V, et al. Identification of pre-leukaemic haematopoietic stem cells in acute leukaemia. Nature. 2014;506:328–33. PubMedPMC
  • 49. Chen Y, Mobius S, Riege K, Hoffmann S, Hochhaus A, Ernst T, et al. Genetic separation of chronic myeloid leukemia stem cells from normal hematopoietic stem cells at single-cell resolution. Leukemia. 2023;37:1561–6. ArticlePubMedPMCPDF
  • 50. Sakhinia E, Farahangpour M, Tholouli E, Liu Yin JA, Hoyland JA, Byers RJ. Comparison of gene-expression profiles in parallel bone marrow and peripheral blood samples in acute myeloid leukaemia by real-time polymerase chain reaction. J Clin Pathol. 2006;59:1059–65. ArticlePubMedPMC

Figure & Data

REFERENCES

    Citations

    Citations to this article as recorded by  
    • Day + 30 detection of minimal residual FLT3-ITD by high-sensitivity PCR-NGS predicts relapse risk and guides post-transplant maintenance in AML
      Shan Jiang, Dan Feng, Li Wan, Nan Yang, Jiao Ma, Jiaxin Cao, Pan Pan, Yawei Zheng, Yigeng Cao, Wenbin Cao, Chen Liang, Xin Chen, Rongli Zhang, Qiaoling Ma, Jialin Wei, Weihua Zhai, Donglin Yang, Yi He, Sizhou Feng, Mingzhe Han, Yao Yao, Aiming Pang, Erlie
      BMC Medicine.2026;[Epub]     CrossRef

    • PubReader PubReader
    • ePub LinkePub Link
    • Cite
      CITE
      export Copy Download
      Close
      Download Citation
      Download a citation file in RIS format that can be imported by all major citation management software, including EndNote, ProCite, RefWorks, and Reference Manager.

      Format:
      • RIS — For EndNote, ProCite, RefWorks, and most other reference management software
      • BibTeX — For JabRef, BibDesk, and other BibTeX-specific software
      Include:
      • Citation for the content below
      Uncovering Putative Causal Non-coding RNAs in Acute and Chronic Myeloid Leukemia: A Genome-Wide Mendelian Randomization Study
      Cancer Res Treat. 2026;58(2):664-676.   Published online June 25, 2025
      Close
    • XML DownloadXML Download
    Uncovering Putative Causal Non-coding RNAs in Acute and Chronic Myeloid Leukemia: A Genome-Wide Mendelian Randomization Study
    Image Image Image
    Fig. 1. Graphical abstract of the study design. The contents within the dashed rectangle represent external validation components. eQTL, expression quantitative trait loci; FDR, false discovery rate; GIVW, Generalized Inverse-variance Weighted; HEIDI, heterogeneity in dependent instruments; LD, linkage disequilibrium; lncRNA, long non-coding RNA; MAF, minor allele frequency; miRNA, microRNA; MR, Mendelian Randomization; MR-Egger, Mendelian Randomization–Egger regression; scaRNA, small cajal body-specific RNA; snoRNA, small nucleolar RNA; snRNA, small nuclear RNA; UKBB, UK Biobank; WGS, whole genome sequencing.
    Fig. 2. Primary Mendelian randomization (MR) analysis between non-coding RNA (ncRNA) expression and acute myeloid leukemia (AML). (A) The total number of ncRNAs assessed for causal relationships with AML in each cohort. (B) The number of ncRNAs primarily identified as causal for AML (FDRGIVW < 0.05) in the FinnGen (blue) and UK Biobank (UKBB) (red) cohorts. (C) Volcano plot of MR results for AML in the FinnGen dataset. The horizontal dashed line (grey) indicates the significance threshold for MR estimates (FDRGIVW < 0.05). The vertical dashed line (grey) represents an odds ratio of 1. (D) Volcano plot of MR results for AML in the UKBB dataset. FDR, false discovery rate.
    Fig. 3. Primary Mendelian randomization (MR) analysis between non-coding RNA (ncRNA) expression and chronic myeloid leukemia (CML). (A) The total number of ncRNAs assessed for causal relationships with CML in each cohort. (B) The number of ncRNAs primarily identified as causal for CML (FDRGIVW < 0.05) in the FinnGen (blue) and UK Biobank (UKBB) (red) cohorts. (C) Volcano plot of MR results for CML in the FinnGen dataset. The horizontal dashed line (grey) indicates the significance threshold for MR estimates (FDRGIVW < 0.05). The vertical dashed line (grey) represents an odds ratio of 1. (D) Volcano plot of MR results for CML in the UKBB dataset. FDR, false discovery rate.
    Uncovering Putative Causal Non-coding RNAs in Acute and Chronic Myeloid Leukemia: A Genome-Wide Mendelian Randomization Study
    Discovery Gene name CHRa) #IVb) Method OR (95% CI) p-valuec)
    FinnGen CTD-2517M22.14 8 19 GIVW 0.71 (0.60-0.84) 8.24E-05
    Weighted median 0.88 (0.68-1.13) 3.04E-01
    GEgger 0.59 (0.39-0.90) 1.31E-02
    HCG22 6 22 GIVW 1.47 (1.33-1.62) 1.46E-14
    Weighted median 1.46 (1.17-1.83) 7.42E-04
    GEgger 1.38 (1.04-1.85) 2.82E-02
    KRT10-AS1 17 38 GIVW 0.64 (0.52-0.79) 4.18E-05
    Weighted median 0.62 (0.47-0.81) 5.94E-04
    GEgger 0.83 (0.57-1.20) 3.16E-01
    LINC01505 9 17 GIVW 0.68 (0.56-0.83) 1.21E-04
    Weighted median 0.70 (0.52-0.94) 1.75E-02
    GEgger 0.84 (0.53-1.34) 4.68E-01
    RP11-42I10.1 16 3 GIVW 3.87 (2.07-7.22) 2.11E-05
    Weighted median 4.15 (1.88-9.14) 4.16E-04
    GEgger 4.60 (1.22-17.28) 2.39E-02
    SPEN-AS1 1 24 GIVW 1.53 (1.26-1.86) 2.25E-05
    Weighted median 1.52 (1.03-2.25) 3.72E-02
    GEgger 1.48 (0.97-2.26) 6.67E-02
    UKBB AC007278.2 2 33 GIVW 0.83 (0.77-0.89) 8.59E-07
    Weighted median 0.82 (0.66-1.02) 7.69E-02
    GEgger 0.76 (0.66-0.87) 4.56E-05
    AC007278.3 2 68 GIVW 0.92 (0.90-0.95) 1.88E-07
    Weighted median 0.90 (0.79-1.02) 9.23E-02
    GEgger 0.98 (0.90-1.07) 6.99E-01
    HOTAIRM1 7 62 GIVW 0.92 (0.89-0.96) 6.72E-05
    Weighted median 0.92 (0.82-1.04) 1.88E-01
    GEgger 0.90 (0.83-0.98) 1.81E-02
    LINC02470 12 64 GIVW 0.79 (0.70-0.89) 5.89E-05
    Weighted median 0.89 (0.76-1.03) 1.27E-01
    GEgger 0.80 (0.70-0.92) 1.16E-03
    RP11-1398P2.1 4 46 GIVW 0.73 (0.64-0.83) 2.60E-06
    Weighted median 0.74 (0.58-0.94) 1.50E-02
    GEgger 0.76 (0.60-0.97) 3.05E-02
    RP11-799D4.4 17 10 GIVW 1.51 (1.27-1.81) 4.67E-06
    Weighted median 1.49 (1.07-2.07) 1.88E-02
    GEgger 1.19 (0.64-2.20) 5.81E-01
    TWF2-DT 3 21 GIVW 1.42 (1.22-1.64) 2.46E-06
    Weighted median 1.50 (0.94-2.40) 9.07E-02
    GEgger 0.86 (0.50-1.45) 5.62E-01
    Gene name Dataset GTEx v8 (main analysis)
    1000 Genomes EUR
    ORa) (95% CI) p-valueb) OR (95% CI) p-value
    HCG22 FinnGen 1.47 (1.33-1.62) 1.46E-14 1.47 (1.30-1.65) 2.60E-10
    UKBB 1.26 (1.13-1.40) 5.07E-05 1.13 (1.02-1.26) 0.017
    RP11-42I10.1 FinnGen 3.87 (2.07-7.22) 2.11E-05 4.00 (1.52-10.5) 4.91E-03
    UKBB 3.12 (1.39-6.98) 5.65E-03 3.50 (1.31-9.34) 0.012
    RP11-1398P2.1 FinnGen 1.02 (0.91-1.14) 7.35E-01 0.98 (0.89-1.10) 7.80E-01
    UKBB 0.73 (0.64-0.83) 2.60E-06 0.74 (0.67-0.82) 1.88E-08
    Discovery Gene name CHRa) #IVb) Method OR (95% CI) p-valuec)
    FinnGen AF131215.9 8 157 GIVW 0.77 (0.70-0.84) 2.44E-08
    Weighted median 0.86 (0.78-0.95) 2.22E-03
    GEgger 0.75 (0.63-0.91) 3.19E-03
    GATD1-DT 11 57 GIVW 0.88 (0.83-0.93) 1.64E-05
    Weighted median 0.93 (0.79-1.10) 3.97E-01
    GEgger 0.94 (0.83-1.06) 3.17E-01
    GMDS-AS1 6 9 GIVW 2.81 (1.75-4.51) 1.80E-05
    Weighted median 2.68 (1.66-4.32) 5.23E-05
    GEgger 2.16 (1.07-4.37) 3.11E-02
    RP11-173B14.4 13 5 GIVW 1.99 (1.42-2.79) 5.57E-05
    Weighted median 2.10 (1.31-3.36) 2.01E-03
    GEgger 1.25 (0.17-9.24) 8.30E-01
    RP11-362F19.1 4 53 GIVW 1.16 (1.09-1.24) 1.08E-06
    Weighted median 1.17 (1.00-1.36) 5.01E-02
    GEgger 1.19 (1.06-1.34) 4.11E-03
    RP11-830F9.6 16 8 GIVW 0.36 (0.23-0.57) 1.30E-05
    Weighted median 0.46 (0.26-0.81) 7.58E-03
    GEgger 0.90 (0.22-3.65) 8.85E-01
    Y_RNA 9 11 GIVW 0.70 (0.59-0.82) 1.49E-05
    Weighted median 0.81 (0.57-1.13) 2.15E-01
    GEgger 0.97 (0.48-1.93) 9.25E-01
    UKBB CD27-AS1 12 35 GIVW 1.71 (1.29-2.25) 1.74E-04
    Weighted median 1.86 (1.41-2.45) 1.24E-05
    GEgger 2.00 (1.45-2.77) 2.82E-05
    FLVCR1-DT 1 94 GIVW 1.20 (1.14-1.26) 8.71E-13
    Weighted median 1.20 (0.97-1.49) 8.77E-02
    GEgger 1.27 (1.16-1.39) 2.93E-07
    HOTAIRM1 7 62 GIVW 0.87 (0.82-0.92) 2.52E-06
    Weighted median 0.83 (0.70-0.99) 3.61E-02
    GEgger 0.81 (0.71-0.92) 8.73E-04
    RBFADN 18 19 GIVW 0.69 (0.60-0.80) 6.19E-07
    Weighted median 0.72 (0.55-0.93) 1.32E-02
    GEgger 0.68 (0.50-0.94) 1.93E-02
    RP11-514P8.2 7 22 GIVW 1.44 (1.19-1.75) 1.71E-04
    Weighted median 1.33 (0.93-1.91) 1.19E-01
    GEgger 1.54 (1.01-2.36) 4.54E-02
    Gene name Dataset GTEx v8 (main analysis)
    1000 Genomes EUR
    ORa) (95% CI) p-valueb) OR (95% CI) p-value
    AF131215.9 FinnGen 0.77 (0.70-0.84) 2.44E-08 0.92 (0.85-1.00) 4.86E-02
    UKBB 0.71 (0.61-0.82) 7.03E-06 0.94 (0.83-1.06) 0.29
    GMDS-AS1 FinnGen 2.81 (1.75-4.51) 1.80E-05 2.22 (1.42-3.46) 4.68E-04
    UKBB 1.85 (1.03-3.32) 4.02E-02 2.06 (1.15-3.69) 0.014
    CD27-AS1 FinnGen 0.95 (0.78-1.15) 5.85E-01 0.97 (0.79-1.19) 0.75
    UKBB 1.71 (1.29-2.25) 1.74E-04 1.47 (1.04-2.06) 2.81E-02
    HOTAIRM1 FinnGen 1.01 (0.97-1.05) 7.91E-01 1.00 (0.95-1.04) 0.84
    UKBB 0.87 (0.82-0.92) 2.52E-06 0.85 (0.79-0.92) 7.05E-05
    RBFADN FinnGen 0.97 (0.87-1.07) 5.25E-01 1.01 (0.83-1.23) 0.92
    UKBB 0.69 (0.60-0.80) 6.19E-07 0.82 (0.62-1.08) 1.63E-01
    Table 1. MR results for putative causal ncRNAs in AML

    The listed ncRNAs were causally linked to AML risk, with FDRGIVW < 0.05 and no evidence of bias across all sensitivity analyses. AML, acute myeloid leukemia; CI, confidence interval; FDR, false discovery rate; GEgger, generalized MR-Egger; GIVW, generalized inverse-variance weighted; MR, Mendelian randomization; ncRNA, non-coding RNA; OR, odds ratio; UKBB, UK Biobank.

    CHR, the chromosome on which the corresponding ncRNA is located,

    #IV, the number of instrumental variables used for a corresponding ncRNA,

    p-values represent nominal p-values before applying FDR correction.

    Table 2. External validation for AML associated causal ncRNAs

    Each ncRNA showing a consistent causal relationship with AML across all MR methods in one cohort (either FinnGen or UKBB) was validated in the other. Also, MR analyses were repeated using the 1000 Genomes European LD reference panel for ncRNAs identified as promising candidates in the main analyses. AML, acute myeloid leukemia; CI, confidence interval; FDR, false discovery rate; GIVW, generalized inverse-variance weighted; LD, linkage disequilibrium; MR, Mendelian randomization; ncRNA, non-coding RNA; UKBB, UK Biobank.

    OR, odds ratio estimated using GIVW,

    p-values represent nominal p-values (pGIVW) prior to FDR correction.

    Table 3. MR results for putative causal ncRNAs in CML

    The listed ncRNAs were causally linked to CML risk, with FDRGIVW < 0.05 and no evidence of bias across all sensitivity analyses. CI, confidence interval; CML, chronic myeloid leukemia; FDR, false discovery rate; GIVW, generalized inverse-variance weighted; MR, Mendelian randomization; ncRNA, non-coding RNA; OR, odds ratio; UKBB, UK Biobank.

    CHR, the chromosome on which the corresponding ncRNA is located,

    #IV, the number of instrumental variables used for a corresponding ncRNA,

    p-values represent nominal p-values before applying FDR correction.

    Table 4. External validation for CML associated causal ncRNAs

    Each ncRNA showing a consistent causal relationship with CML across all MR methods in one cohort (either FinnGen or UKBB) was validated in the other. Also, MR analyses were repeated using the 1000 Genomes European LD reference panel for ncRNAs identified as promising candidates in the main analyses. CI, confidence interval; CML, chronic myeloid leukemia; FDR, false discovery rate; LD, linkage disequilibrium; MR, Mendelian randomization; ncRNA, non-coding RNA; UKBB, UK Biobank.

    OR, odds ratio estimated using GIVW,

    p-values represent nominal p-values (pGIVW) prior to FDR correction.


    Cancer Res Treat : Cancer Research and Treatment
    Close layer
    TOP