* denotes equal contribution; ** denotes (co-)corresponding authors
Featured Research
A cross-population compendium of gene–environment interactions. Nature 688–697 (2026).
Environmental differences in genetic effect sizes, namely, gene-environment interactions, may uncover the genetic encoding of phenotypic plasticity1-3. We provide a cross-population atlas of gene-environment interactions comprising 440,210 individuals from European and Japanese populations, with replication in 539,794 individuals from diverse populations. By decomposing the contributions from age, sex and lifestyles, we delineate the aetiology of these gene-environment interactions, including a reverse-causality from a disease-related dietary change. Genome-wide analyses uncovered missing heritability and trait-trait relationships connected by the synergistic effects of genome and environments, which systematically affected polygenic prediction accuracy and cross-population portability. Single-cell projection revealed aging shift of pathways and cell types responsible for genetic regulation. Omics-level gene-environment analyses identified multiple sex-discordant genetic effects in lipid metabolism, informing clinical trial failures for genetically supported drug development. Our comprehensive gene-environment study decodes the dynamics of genetic associations, offering insights into complex trait biology, personalized medicine and drug development.
A gene–environment interaction atlas across phenotypes, environments, and populations, showing the importance of gene–environment interactions for understanding the biology underlying dynamic genetic effects and informing precision medicine and drug development.
Inconsistent embryo selection across polygenic score methods. Nature Human Behaviour 2264–2267 (2024).
Private enterprises offer preimplantation genetic testing with polygenic scores to select embryos with ‘desirable’ potential. In silico simulations using biobank resources show that the selected embryo would rely substantially on the choice of polygenic score method and randomness in score construction, which raises ethical concerns.
Focusing on embryo selection based on genetically predicted future traits in in vitro fertilization, we showed that embryo rankings varied substantially across scoring methods and due to random fluctuations in score construction.
Common germline risk variants impact somatic alterations and clinical features across cancers. Cancer Research 83, 20–27 (2022).
Aggregation of genome-wide common risk variants, such as polygenic risk score (PRS), can measure genetic susceptibility to cancer. A better understanding of how common germline variants associate with somatic alterations and clinical features could facilitate personalized cancer prevention and early detection. We constructed PRSs from 14 genome-wide association studies (median n = 64,905) for 12 cancer types by multiple methods and calibrated them using the UK Biobank resources (n = 335,048). Meta-analyses across cancer types in The Cancer Genome Atlas (n = 7,965) revealed that higher PRS values were associated with earlier cancer onset and lower burden of somatic alterations, including total mutations, chromosome/arm somatic copy-number alterations (SCNA), and focal SCNAs. This contrasts with rare germline pathogenic variants (e.g., BRCA1/2 variants), showing heterogeneous associations with somatic alterations. Our results suggest that common germline cancer risk variants allow early tumor development before the accumulation of many somatic alterations characteristic of later stages of carcinogenesis. Significance:: Meta-analyses across cancers show that common germline risk variants affect not only cancer predisposition but the age of cancer onset and burden of somatic alterations, including total mutations and copy-number alterations.
A pan-cancer analysis of germline–somatic associations, which showed that patients with a higher genetic risk of cancer tend to develop cancer at a younger age and with fewer somatic mutations and somatic copy-number alterations.
*Namba, S., *Konuma, T., Wu, K.-H., Zhou, W. & Okada, Y. A practical guideline of genomics-driven drug discovery in the era of global biobank meta-analysis. Cell Genomics 2, 100190 (2022).
Genomics-driven drug discovery is indispensable for accelerating the development of novel therapeutic targets. However, the drug discovery framework based on evidence from genome-wide association studies (GWASs) has not been established, especially for cross-population GWAS meta-analysis. Here, we introduce a practical guideline for genomics-driven drug discovery for cross-population meta-analysis, as lessons from the Global Biobank Meta-analysis Initiative (GBMI). Our drug discovery framework encompassed three methodologies and was applied to the 13 common diseases targeted by GBMI (N mean = 1,329,242). Individual methodologies complementarily prioritized drugs and drug targets, which were systematically validated by referring previously known drug-disease relationships. Integration of the three methodologies provided a comprehensive catalog of candidate drugs for repositioning, nominating promising drug candidates targeting the genes involved in the coagulation process for venous thromboembolism and the interleukin-4 and interleukin-13 signaling pathway for gout. Our study highlighted key factors for successful genomics-driven drug discovery using cross-population meta-analyses.
Developing and validating a practical framework for genomics-driven drug discovery in the era of cross-biobank GWAS meta-analyses that integrates association results with multi-omics data to identify therapeutic targets and candidate drugs.
Preprints
Beyond point estimates: quantifying predictive uncertainty reveals hidden dimensions of biological age acceleration and improves risk interpretation. bioRxiv (2026).
Abstract: Biological age estimates are increasingly used to study aging, disease risk, and mortality, yet their predictive uncertainty is rarely quantified. Consequently, conventional age-gap measures can treat deviations as equally informative even when the underlying biological age predictions differ substantially in reliability. We developed a framework for uncertainty-aware biological aging that generates calibrated prediction intervals and individualized probabilities of accelerated or decelerated aging alongside point estimates. We applied this framework to the UK Biobank Pharma Proteomics Project, evaluating three composite and eleven organ-specific biological age clocks. Predictive uncertainty varied substantially both within and across clocks, revealing that apparently extreme age gaps can differ markedly in the strength of evidence supporting accelerated or decelerated aging. In particular, low-accuracy clocks, including many organ-specific clocks, provided little evidence for confidently accelerated or decelerated aging. Beyond biological age gaps, prediction-interval width was independently associated with disease risk and mortality, particularly for composite, brain, and immune clocks, suggesting that predictive uncertainty captures an additional dimension of biological aging that may reflect increased molecular heterogeneity and dysregulation associated with aging and disease. We replicated these findings in Biobank Japan and an independent clinical cohort from Stanford. By incorporating individual-specific predictive uncertainty, our framework provides a more informative characterization of biological aging and enables improved individual-level risk stratification for disease prevention and longitudinal monitoring.The Biobank Rare Variant consortium powers the discovery of rare genetic associations through global collaboration. medRxiv (2026).
Rare coding variants can have large effects on disease risk and provide direct routes from human genetics to disease mechanisms and therapeutic targets, but their discovery is constrained by sample size, particularly for low-prevalence diseases. Here we establish the Biobank Rare Variant Analysis (BRaVa) consortium, a global rare variant association resource that integrates sequencing and linked health-record data from ten biobanks and cohorts comprising over 1.2 million individuals across diverse ancestries. We performed gene-based meta-analyses of rare coding variation across 33 clinical endpoints and 11 quantitative traits. Aggregating evidence across biobanks and ancestries identified 514 gene-trait associations, including 31 not previously reported in prior studies or curated association resources following systematic literature review. Notably, 36.1% of gene-level associations were undetectable in any individual biobank, and 91 emerged only through cross-ancestry meta-analysis, demonstrating that federated integration enables discovery beyond the reach of single cohorts. Similar gains were observed at the variant level, where 25.0% of phenotype-locus associations were detectable only through meta-analysis. Effect size estimates were correlated across ancestries with concordant directions of effect, supporting the generalizability of rare variant associations. The identified signals implicate pathways involved in transcriptional and epigenetic regulation, metabolism, vascular and epithelial biology, and immune function, highlighting rare coding variation as an engine for biological discovery across medical record phenotypes. For example, damaging variation in ANKRD12 implicates inflammatory transcriptional dysregulation in asthma and chronic obstructive pulmonary disease, and ultra-rare predicted loss-of-function variants in NAA15 link protein acetylation processes to type 2 diabetes risk. BRaVa establishes a scalable framework and freely available community resource for rare variant meta-analysis across global biobanks. Public release of gene- and variant-level association summary statistics provides a reference map of rare coding variant associations to support disease gene discovery, biological interpretation, and therapeutic target prioritization as sequencing-linked health-record resources continue to expand.Genome-wide association and Mendelian randomization analyses link Helicobacter pylori infection to Human Leukocyte Antigen polymorphisms and autoimmune diseases. medRxiv (2026).
Abstract: Helicobacter pylori ( H. pylori ) infects the gastric epithelium of approximately half of the global population, and is a well-known risk factor for developing gastric cancer. Despite the clinical significance of H. pylori infection, many genetic factors that contribute to susceptibility remain unidentified. While it is well-established that H. pylori infection can result in gastritis and peptic ulcers, which may progress to gastric cancer, its causal link to other diseases remains unclear. We performed the genome-wide association study (GWAS) for anti- H. pylori IgG antibody titers, which were validated as a surrogate marker for H. pylori infection by the correlation with clinical traits, followed by gene-based and pathway analyses, involving up to 140,863 individuals. This included 56,967 in the discovery phase, and 68,211 in the replication phase from Japanese cohorts, and an additional 15,685 from European populations in a cross-ancestry meta-analysis. We reveal significant associations between H. pylori infection and polymorphisms in the Human Leukocyte Antigen (HLA) class II region within the Major Histocompatibility Complex (MHC), as well as genes related to innate immunity, including CCDC80 , NFKBIZ , TIFA , PSCA , and TRAF3 . Mendelian randomization (MR) analysis revealed that genetic liability to H. pylori infection has both positive and negative causal relationships with a variety of diseases, including autoimmune-related diseases such as Type 1 diabetes, Hashimoto’s disease, atopic dermatitis, as well as traits like body height and weight. These genetic findings strongly support the notion that genetic liability to H. pylori infection influences not only gastrointestinal diseases, but also a broader spectrum of health issues, thereby providing valuable insights for public health strategies and personalized medicine approaches.Expanding the genetic landscape of endometriosis: Integrative -omics analyses implicate key genes and pathways in a multi-ancestry study of over one million women. Res Sq (2025).
Abstract We report the findings of a genome-wide association study (GWAS) meta-analysis of endometriosis across 14 biobanks worldwide, including 32% non-European patient participants, as part of the Global Biobank Meta-Analysis Initiative (GBMI). Out of 58 total loci (29 previously unreported), the largest meta-analysis accounted for 46 (20 previously unreported). We detected the first genome-wide significant loci (2p13.3 and 20q13.2) uniquely driven by the African-ancestry meta-analysis. Our imaging- and surgery-confirmed phenotypes yielded six additional previously unreported loci. Leveraging our large and diverse study population, we observed SNP heritability estimates of 9-13% for all ancestry groups, and 13 loci had at least one variant in the credible set after fine-mapping. Investigating the complex array of endometriosis comorbidities and risk factors revealed 135 genetically correlated phenotypes and 95 with evidence of vertical pleiotropy, including triglycerides and anxiety disorders. We prioritized 35 disease-relevant cellular contexts from the endometrial cell atlas and found 322 examples of differentially expressed genes in cells from donors with endometriosis. Further high-throughput multi-omic analyses implicated a total of 282 genes in endometriosis pathogenesis. Our diverse, comprehensive GWASs, with downstream analyses spanning molecular to phenotypic scales, provide detailed evidence for aspects of endometriosis including the role of immune cell types, Wnt signaling, and cellular proliferation. These interconnected pathways and risk factors underscore the complex, multi-faceted etiology of endometriosis, suggesting multiple targets for precise and effective therapeutic interventions.Yang, C., Namba, S., Matsuda, K., Okada, Y., Moran, L., Vincent, A. & Marques, F. Z. Acetate, a fibre-derived gut metabolite, is associated with reduced cardiovascular disease risk in females with early menopause. medRxiv (2026).
Abstract: Background and Aims: Estrogen deficiency and testosterone excess substantially increase cardiovascular disease risk in females. Dietary fibre and its microbial by-products, short-chain fatty acids, have cardioprotective effects. We aim to investigate whether plasma acetate – the most abundant short-chain fatty acid – is associated with improved cardiovascular outcomes in females with altered sex hormone profiles. Methods: This cohort study included 105,563 female participants from the UK Biobank and Biobank Japan with up to 10 years of follow-up. The primary outcome was major adverse cardiovascular event (MACE) in relation to early menopause (surrogate for estrogen insufficiency) and plasma free testosterone. Proteomics profiling explored underlying molecular pathways. Results: Higher plasma acetate was associated with lower 10-year MACE incidence (HR=0.887, p =0.002) in the UK Biobank. Acetate levels above the median attenuated the high MACE risk associated with early menopause (HR=1.155, p =0.075) compared with lower acetate levels (HR=1.431, p<0.001), further supported by a negative early menopause × acetate interaction on MACE incidence (HR=0.894, p =0.037). This mitigation pattern was replicated in Biobank Japan (above median: HR=1.350, p =0.067; below median: HR=1.423, p =0.028). The top acetate quartile (>27.1g fibre/day) attenuated the early menopause-associated MACE risk. Proteomics implicated the underrepresentation of pro-inflammatory pathways. High acetate also attenuated MACE risk associated with elevated free testosterone, although the interaction term was not significant. Conclusions: Higher plasma acetate mitigated cardiovascular risk in females, particularly in those with early menopause, potentially by modulating pro-inflammatory pathways. These findings support recommendations for higher dietary fibre intake as a CVD prevention strategy for females with early menopause. Translational Perspective: Plasma acetate is the key product of gut microbial fermentation of fibre. Using two cohort studies, totalling 105,563 female participants followed for 10-years from two independent biobanks with different ethnicities, higher plasma acetate was significantly associated with lower cardiovascular disease risk, especially for those with early menopause. Higher microbial acetate protects against cardiovascular disease in females, with greater benefits in those with hormone related cardiovascular vulnerability. These findings may inform the design of clinical trials and guidelines, supporting a personalised nutrition strategy based on hormonal profile.Genomic analyses reveal new insights into Alzheimer’s disease. medRxiv (2025).
Alzheimer’s disease (AD) is the most common cause of dementia, with global case numbers projected to reach 153 million in 2050. AD is highly heritable, with twin-based heritability estimates of 60-80%. While 1,200 causal loci are predicted to exist for AD, approximately 80 have been associated with AD in two recent studies, suggesting that many loci remain to be discovered. Here, we analyzed data from 109,479 cases, 74,141 proxy cases, 2,131,799 controls, and 499,708 proxy controls from diverse ancestries, identifying 118 loci in a multi-ancestry analysis and 9 additional loci in ancestry-specific analyses, 48 of which are new. We identified new AD risk genes, prioritized potential drug targets, and identified microglia and, for the first time, several neuronal cell types enriched for AD-associated genetic risk. Moreover, we improved polygenic prediction and estimated a single-nucleotide polymorphism (SNP) heritability of 16%. Together, our findings offer insights into the genetic architecture and potential pathobiology of AD, as well as specific targets for future drug development research.Rare k-mers reveal centromere haplogroups underlying human diversity and cancer translocations. bioRxiv (2025).
Abstract: Centromeres are among the most diverse and dynamically evolving regions of the human genome and are commonly affected in various human cancers. However, organized into highly repetitive α-satellite higher-order repeats (HORs), human centromere sequences have long resisted detailed genomic analysis. Although the development of long-read sequencing platforms has enabled the analysis of complete centromere sequences, their application to a large set of samples is still largely limited, preventing our understanding of centromere variation and haplotype structures across large human populations and the structural basis of centromere-involving translocations in cancer. Here we show that rare k-mers present in centromeric regions can serve as effective markers for dissecting the complexity of centromere structure, particularly that of active α-satellite HOR arrays (aHOR arrays), across human populations and for understanding centromere-involving abnormalities in cancer. Based on rare k-mer-based clustering, centromere aHOR arrays are clustered into discrete haplogroups (aHOR-HGs) with distinct structural features. These k-mers were also used to develop a framework that enables the inference of haplogroups in a given sample based on short-read whole genome sequencing (WGS) data (ascairn). By applying ascairn to large-scale human population datasets ( n > 3,300), we revealed the diversity of aHOR-HGs and their geographic histories across populations. The rare k-mer-based approach was also applied to investigate the structure of 1p/19q co-deletion, a highly recurrent centromere-involving translocation in IDH -mutated oligodendrogliomas. Analyzing short-read WGS data from 142 cases with 1p/19q co-deletion using rare k-mers, we showed that breakpoints of 1p/19q co-deletion were mapped to aHOR arrays in chromosomes 1 ( D1Z7 ) and 19 ( D19Z3 ), which was validated by long-read sequencing of two 1p/19q co-deletion-positive cases. Notably, the translocation preferentially involved haplogroups composed of haplotypes containing larger regions susceptible to rearrangement. These results highlight the role of rare k-mers in dissecting the complexity of centromere sequences and their evolutionary history as well as understanding centromere-involving abnormalities associated with human diseases.
Publications
Clonal hematopoiesis related to ionizing radiation in atomic bomb survivors. Leukemia (2026).
Genome-wide association analyses of gestational phenotypes identify context-specific genetic effects. Nature Genetics 58, 1845–1854 (2026).
The approximately 40-week gestational period is central to human reproduction, yet the genetic architecture of diverse gestational phenotypes and their links to maternal late-life health remain unclear. In 111 phenotypes from up to 121,579 Chinese pregnancies (median n = 78,535 per phenotype), we identified 4,688 independent genome-wide significant signals, including 1,703 new associations. Gestation-specific effects were observed for 7.8% of variants across 30 phenotypes; 18.7% of signals for 24 longitudinal hematological traits exhibited genotype-by-gestational-timing interactions across five antenatal and postpartum periods. Dynamic genetic effects were enriched in growth-regulatory and hormone-regulatory pathways, reflecting maternal-fetal interactions. Genetic correlation and Mendelian randomization analyses with 80 diseases and medication traits in BioBank Japan females revealed shared genetic overlaps and potential causal links between gestational phenotypes and maternal mid-life and late-life health. These results establish a dynamic genetic atlas of human gestation, providing a framework for precision maternal health.Polygenic Prediction of Nongoal Response to Statin Therapy. Circulation: Genomic and Precision Medicine (2026).
BACKGROUND:: Genetic differences may contribute to interindividual variability in LDL-C (low-density lipoprotein cholesterol) lowering with statin therapy. Polygenic risk scores may help identify individuals unlikely to achieve guideline-concordant LDL-C targets on statins, enabling earlier therapy intensification. METHODS:: We developed a multiancestry polygenic risk score for statin nongoal response (PRS-NGR) and evaluated its association with failure to achieve an on-statin LDL-C level of ≤70 mg/dL and with percent LDL-C reduction. This longitudinal cohort study used genotyping and electronic health record-linked data from the All of Us Research Program (2018–2025), the UK Biobank (2014–2023), and the Biobank Japan (2003–2008). Participants were statin users with at least 1 prestatin and 1 on-statin LDL-C measurement. Associations were assessed overall and by genetic ancestry, with replication in the UK Biobank and Biobank Japan. RESULTS:: The study included 46 564 participants from All of Us, 37 009 from the UK Biobank, and 3613 from Biobank Japan. In All of Us, higher PRS-NGR was associated with increased odds of nongoal response (odds ratio per SD, 1.43 [95% CI, 1.37–1.49]). Compared with the middle quintile, individuals in the top 1% had a higher risk, whereas those in the bottom 1% had a lower risk of nongoal response. Associations of PRS-NGR were consistent across African, European, and Latin American ancestry groups. Each SD increase in PRS-NGR corresponded to a 1.2-percentage-point smaller LDL-C reduction. Findings were replicated in the UK Biobank (odds ratio per SD, 2.39 [95% CI, 2.23–2.57]) and in Biobank Japan (odds ratio per SD, 1.31 [95% CI, 1.17–1.46]). Integration of PRS-NGR with guideline-based criteria identified individuals who derived a higher LDL-C% change and increased identification of statin-eligible individuals. CONCLUSIONS:: We developed and validated a multiancestry polygenic risk score that estimates the risk of nongoal LDL-C response to statin therapy. Incorporation of polygenic risk into lipid-lowering treatment paradigms may improve risk stratification and support more tailored therapy intensification strategies.Cross-Species Comparison between Mouse Kidney Multi-Omics and Human Genome-Wide Association Studies Highlights Molecular Features Associated with Oxalate Nephropathy. Kidney360 (2026).
Background:: Urolithiasis, a recurrent and increasingly prevalent disease worldwide, has a complex pathophysiology that hinders a comprehensive understanding. Methods:: We performed a cross-species integrative analysis combining multi-omics profiling of kidneys from a hyperoxaluric mouse model and summary statistics from a large-scale human genome-wide association study (GWAS). Results:: In the multi-omics analysis of kidneys from a hyperoxaluric mouse model using RNA-seq, whole-cell proteomics, and phosphoproteomics, we identified 1,173 genes, 342 proteins, and 516 phosphorylated peptides that were differentially expressed compared with control mice. We utilized publicly available large-scale meta-GWAS summary statistics of urolithiasis from BioBank Japan, UK Biobank, and FinnGen (n = 198,769) and prioritized genes using gene enrichment analysis. Through cross-species integration of mouse and human omics, we identified 46 molecules potentially relevant to urolithiasis, hereafter referred to as cross-species urolithiasis-related molecules. We examined the expression and genetic associations of these 46 molecules in human urolithiasis. Among these, CRYAB and SHROOM3 showed differential expression in human renal papilla tissues with and without stones. Colocalization analysis between urolithiasis GWAS and expression quantitative trait loci (eQTL) data of human renal tubules and glomeruli suggested shared signals at the GLUD1 , UMOD , SLC34A1 , TCEA3 , and H1-0 regions. Mendelian randomization analysis using the same renal eQTL datasets further supported these findings, with GLUD1 , UMOD , TCEA3 , and H1-0 showing statistically significant associations, suggesting a strong genetic link with urolithiasis. Conclusions:: The study’s findings highlight a set of cross-species urolithiasis-related molecules, supported by both expression and genetic evidence, offering insights for future functional validation and mechanistic investigations.Meta-analysis across six global biobanks identifies recessive coding associations with complex traits and diseases. The American Journal of Human Genetics 113, 1330–1346 (2026).
Rare bi-allelic variation is a major contributor to human disease risk, yet its effects are difficult to study at scale in population cohorts owing to the limited number of individuals with putatively deleterious bi-allelic genotypes and the challenges of accurately phasing low-frequency variants. Here, we present recessive, gene-based analyses of rare and low-frequency variants in up to 948,690 exome- or whole-genome-sequenced individuals across six biobanks with linked electronic health records. Through statistical phasing, we inferred putatively damaging compound-heterozygous genotypes, increasing the number of bi-allelic damaging genotypes by 19%. Restricting to predicted loss-of-function (pLoF) variants, we identified 5,563 genes harboring bi-allelic genotypes, a 19.8% increase in putative knockouts. We then considered all low-frequency variants (minor allele frequency [MAF] <5%) and performed gene-based recessive association testing using putatively damaging bi-allelic genotypes, identifying 58 significant associations (false discovery rate [FDR] ≤1% or prec≤7.5 × 10-7) after meta-analysis and Cauchy combination of nonsynonymous annotations. Comparing recessive and additive models, we found 17 instances where recessive effects were more pronounced, including several previously unreported associations, such as HBB with heart failure (prec = 2.6 × 10-14; padd = 0.98), LECT2 with height (prec = 3.7 × 10-14; padd = 4.1 × 10-10), and ENSG00000267561 with height (prec = 2.9 × 10-9; padd = 0.37). This study demonstrates the potential of federated approaches to study the effects of rare bi-allelic variation.Integrative GWAS and snRNA-seq Reveal a Mesenchymal-Like Endothelial Signature in Moyamoya Disease. Stroke 57, 1336–1348 (2026).
BACKGROUND:: Moyamoya disease (MMD) has a strong genetic basis, with the rare RNF213 p.Arg4810Lys variant (rs112735431) representing a major risk factor, while the broader genetic architecture and disease-relevant vascular cell types remain incompletely understood. METHODS:: We conducted a genome-wide association study in Japanese individuals (n=47 656; 401 MMD cases and 47 255 controls). Population-level features at MMD risk loci were examined by regional allele frequency and haplotype analyses. We performed single-nucleus RNA-seq of superficial temporal arteries from patients with MMD (n=3). Cell type–specific enrichment of genome-wide association study signals was assessed using the Single-Cell Disease Relevance Score. Endothelial signatures were validated by integration with publicly available single-cell data sets from controls (n=5) and immunohistochemistry for candidate markers (n=1). RESULTS:: Beyond rs112735431, we identified a genome-wide significant signal in the HDAC9-TWIST1 region (rs12530920; P =3.3×10 −14 ; odds ratio, 1.77). Conditional analysis on rs112735431 revealed a protective RNF213 missense variant, p.Asp1331Gly (rs8074015; P =3.7×10 − 9 ; odds ratio, 0.53), whose minor allele was mutually exclusive with rs112735431-A on haplotypes. Population analysis revealed geographic variation and extended haplotype structure of the rs112735431-A allele in Japan. Single-nucleus RNA-seq identified a mesenchymal-like endothelial cell (MEC) population with selective FN1 expression. Genome-wide association study–prioritized disease genes were strongly enriched in MECs. MECs showed mesenchymal pathway activation with a regulatory program distinct from canonical endothelial states. The proportion of MECs was markedly increased in MMD (72% versus 28% in controls), and fibronectin 1 (FN1) expression in endothelial regions was confirmed by immunohistochemistry. CONCLUSIONS:: Our findings identify a protective RNF213 p.Asp1331Gly variant (rs8074015) that is mutually exclusive with the known rs112735431-A allele. Genetic risk converges on an MEC state markedly expanded in MMD.Elucidating genetic backgrounds of myasthenia gravis in Japanese by genome-wide association studies and multi-omics analyses of thymoma. Nature Communications 17, (2026).
Myasthenia gravis (MG) is an autoimmune disorder characterized by impaired neuromuscular transmission and motor symptoms. Its genetic background remains unclear, particularly beyond specific subtypes reported in European populations. Here, we perform a genome-wide association study (GWAS) of 1,434 MG cases covering all disease subtypes and 42,913 controls of Japanese, which newly identify the TERT locus (odds ratio [OR] = 1.31, P = 1.7×10-10). Subtype-stratified GWASs show stronger signals for generalized MG (gMG; OR = 1.38, P = 1.6×10-12), anti AChR antibody-positive gMG (g-AChR-Ab(+)MG; OR= 1.49, P = 2.1×10-15), and thymoma-associated gMG (g-TAMG; OR = 1.92, P = 1.1×10-15). Fine-mapping of the major histocompatibility complex region reveal distinct associations of HLA-DRB1 with late onset gMG (g-LOMG) and HLA-A with early onset gMG (g-EOMG). The MG risk TERT lead variant rs2736099 is associated with poor treatment response, especially in g-AChR-Ab(+)MG and g-EOMG (P < 0.0042). The biobank-based phenome-wide association study identify pleiotropic effects on lung cancer, hematological traits, and telomere length. Single cell transcriptomics and immunohistochemistry identified immature lymphocyte-specific TERT expression in thymoma specimens. Full-length transcriptomics reveal allele-specific decreasing effect of rs2736099-A on TERT expression. Our study unveils genetics of MG distinctly across disease subtypes, and involvement of TERT in its pathogenesis.Genetic regulation across germline and somatic variation on the Y chromosome contributes to type 2 diabetes. Nature Medicine 894–905 (2026).
Our understanding of the biological role of the Y chromosome remains limited. Here, we systematically profile germline Y haplogroups and somatic loss of the Y chromosome (LOY) in 122,683 East Asian males from BioBank Japan and 181,472 European males from the UK Biobank. A phenome-wide scan uncovers male-specific genetic regulation of complex traits, including pleiotropic effects of the Japanese-specific haplogroup D on height and type 2 diabetes (T2D). LOY increases T2D risk in East Asians but is associated with reduced T2D risk in Europeans. In East Asians, LOY contributes to T2D incidence particularly among males with lower polygenic risk scores, providing a compensatory explanation for disease risk beyond germline genetics. Incorporating sex-chromosome variation improves polygenic prediction of T2D risk in both sexes. Single-cell analyses reveal cell type-specific accumulation of LOY across tissues and disease contexts, with LOY in pancreatic β cells potentially impairing glucose metabolism. Our study demonstrates the clinical relevance of Y chromosome variation for diabetes risk prediction and management.Delineating the Genetic Basis of RNF213-related vasculopathies: The association of PKHD1 variants with bilateral cerebral vasculopathy. European Journal of Human Genetics 788–795 (2026).
Moyamoya disease (MMD) is an idiopathic cerebrovascular disorder characterized by progressive stenosis of the internal carotid artery termini and the formation of an abnormal network of fragile perforators. Although the RNF213 gene has been implicated as a susceptibility factor for MMD, the precise genetic basis of the disease remains elusive. This study aimed to investigate other genetic factors contributing to differences in disease manifestation and vascular phenotypes. We conducted whole-exome sequencing of patients with RNF213-related vasculopathy and healthy controls. In total, 122 patients (comprising 69 with bilateral MMD, 13 with unilateral MMD, and 40 with intracranial artery stenosis [ICAS]) and 458 controls were enrolled. Following appropriate quality control measures, case-control analysis and in-case analysis (bilateral MMD vs. unilateral MMD and ICAS) were conducted using single-variant association testing and gene-based aggregation testing. Although no significant locus or gene was found in the case-control analysis, the PKHD1 gene emerged as a top candidate associated with bilateral MMD in the in-case analysis. Two rare damaging variants, p.Ile2364Asn and p.Ser3210Cys, were significantly more prevalent in the bilateral MMD group and were further validated in an independent cohort of 216 individuals. Publicly available single-cell transcriptome data of mouse cerebrovascular and perivascular cells revealed that Pkhd1 expression was significantly higher in specific endothelial-cell clusters. Despite some limitations, our study provides new insights into the potential role of the PKHD1 gene in bilateral MMD, highlighting the need for further investigation of endothelial gene expression in patients with RNF213 and PKHD1 mutations.Global multi-ancestry genome-wide analyses identify genes and biological pathways associated with thyroid cancer and benign thyroid diseases. Nature Genetics 58, 307–316 (2026).
Thyroid diseases are common and highly heritable. We performed a meta-analysis of genome-wide association studies from 19 biobanks for five thyroid diseases: thyroid cancer (ThC), benign nodular goiter, Graves’ disease, lymphocytic thyroiditis and primary hypothyroidism. We analyzed genetic association data from ~2.9 million genomes and identified 313 known and 570 new independent loci linked to thyroid diseases. We discovered genetic correlations between ThC, benign nodular goiter and autoimmune thyroid diseases ( rg = 0.16–0.97). Telomere maintenance genes contributed to benign and malignant thyroid nodular disease risk, whereas cell cycle, DNA repair and damage response genes were associated with ThC. We propose a paradigm that explains genetic predisposition to benign and malignant thyroid nodules. We found polygenic risk score associations with ThC risk of structural disease recurrence, tumor size, multifocality, lymph node metastases and extranodal extension. Polygenic risk scores identified individuals with aggressive ThC in a biobank, creating an opportunity for genetically informed population screening.A cross-population compendium of gene–environment interactions. Nature 688–697 (2026).
Environmental differences in genetic effect sizes, namely, gene-environment interactions, may uncover the genetic encoding of phenotypic plasticity1-3. We provide a cross-population atlas of gene-environment interactions comprising 440,210 individuals from European and Japanese populations, with replication in 539,794 individuals from diverse populations. By decomposing the contributions from age, sex and lifestyles, we delineate the aetiology of these gene-environment interactions, including a reverse-causality from a disease-related dietary change. Genome-wide analyses uncovered missing heritability and trait-trait relationships connected by the synergistic effects of genome and environments, which systematically affected polygenic prediction accuracy and cross-population portability. Single-cell projection revealed aging shift of pathways and cell types responsible for genetic regulation. Omics-level gene-environment analyses identified multiple sex-discordant genetic effects in lipid metabolism, informing clinical trial failures for genetically supported drug development. Our comprehensive gene-environment study decodes the dynamics of genetic associations, offering insights into complex trait biology, personalized medicine and drug development.Proteogenomics in cerebrospinal fluid and plasma reveals new biological fingerprint of cerebral small vessel disease. Nature Aging 2514–2531 (2025).
Cerebral small vessel disease (cSVD) is a leading cause of stroke and dementia with no specific treatment, of which molecular mechanisms remain poorly understood. To identify potential biomarkers and therapeutic targets, we applied Mendelian randomization to examine over 2,500 proteins measured in plasma and, uniquely, cerebrospinal fluid, in relation to magnetic resonance imaging (MRI) markers of cSVD in more than 40,000 individuals. Here we show that 49 proteins are associated with MRI markers of cSVD, most prominently in cerebrospinal fluid. We highlight associations that are consistent across platforms and ancestries, and supported by complementary observational analyses, and we explore differences between fluids. The proteins are enriched in pathways related to the extracellular matrix, immune response and microglial activity. Many also associate with stroke and dementia, and several correspond to existing drug targets. Together, these findings reveal a robust biological fingerprint of cSVD and highlight opportunities for biomarker and drug discovery and repositioning.Deciphering state-dependent immune features from multi-layer omics data at single-cell resolution. Nature Genetics 57, 1905–1921 (2025).
Current molecular quantitative trait locus catalogs are mostly at bulk resolution and centered on Europeans. Here, we constructed an immune cell atlas with single-cell transcriptomics of >1.5 million peripheral blood mononuclear cells, host genetics, plasma proteomics and gut metagenomics from 235 Japanese persons, including patients with coronavirus disease 2019 (COVID-19) and healthy individuals. We mapped germline genetic effects on gene expression within immune cell types and across cell states. We elucidated cell type- and context-specific human leukocyte antigen (HLA) and genome-wide associations with T and B cell receptor repertoires. Colocalization using dynamic genetic regulation provided better understanding of genome-wide association signals. Differential gene and protein expression analyses depicted cell type- and context-specific effects of polygenic risks. Various somatic mutations including mosaic chromosomal alterations, loss of Y chromosome and mitochondrial DNA (mtDNA) heteroplasmy were projected into single-cell resolution. We identified immune features specific to somatically mutated cells. Overall, immune cells are dynamically regulated in a cell state-dependent manner characterized with multiomic profiles.Whole-genome sequencing reveals rare and structural variants contributing to psoriasis and identifies CERCAM as a risk gene. Cell Genomics 5, 100978 (2025).
Psoriasis vulgaris (PsV) is an immune-mediated inflammatory skin disorder with complex genetic architecture. Most genome-wide association studies (GWASs) of PsV have been limited to analyzing common single-nucleotide variants in Europeans, lacking diversity in the variant spectrum and ancestral background. To investigate the contribution of rare variants (RVs) and structural variants (SVs), we perform a whole-genome sequencing study involving 1,415 PsV cases and 3,968 controls in Japanese. A GWAS signal at IFNLR1 is fine-mapped to a 3.3-kb deletion SV disrupting an epithelium-specific putative enhancer, which is validated by PacBio long-read sequencing. Gene-based RV analyses identify two susceptibility genes: IFIH1 (p = 9.8 × 10-6) and CERCAM (p = 4.1 × 10-7). Notably, IL36RN, a causative gene for generalized pustular psoriasis, a rare and lethal multi-systemic inflammatory disorder, is associated with common PsV (p = 1.2 × 10-4). Finally, Cercam knockout (Cercam-/-) in an imiquimod-induced psoriasis mouse model aggravates dermatitis with elevated T cell retention in the subepidermis. Our study elucidates the overlooked genetic basis of PsV.Dissecting cross-population polygenic heterogeneity across respiratory and cardiometabolic diseases. Nature Communications 16, 3765 (2025).
Biological mechanisms underlying multimorbidity remain elusive. To dissect the polygenic heterogeneity of multimorbidity in twelve complex traits across populations, we leveraged biobank resources of genome-wide association studies (GWAS) for 232,987 East Asian individuals (the 1st and 2nd cohorts of BioBank Japan) and 751,051 European individuals (UK Biobank and FinnGen). Cross-trait analyses of respiratory and cardiometabolic diseases, rheumatoid arthritis, and smoking identified negative genetic correlations between respiratory and cardiometabolic diseases in East Asian individuals, opposite from the positive associations in European individuals. Associating genome-wide polygenic risk scores (PRS) with 325 blood metabolome and 2917 proteome biomarkers supported the negative cross-trait genetic correlations in East Asian individuals. Bayesian pathway PRS analysis revealed a negative association between asthma and dyslipidemia in a gene set of peroxisome proliferator-activated receptors. The pathway suggested heterogeneity of cell type specificity in the enrichment analysis of the lung single-cell RNA-sequencing dataset. Our study highlights the heterogeneous pleiotropy of immunometabolic dysfunction in multimorbidity.Contribution of germline and somatic mutations to risk of neuromyelitis optica spectrum disorder. Cell Genomics 5, 100776 (2025).
Neuromyelitis optica spectrum disorder (NMOSD) is a rare autoimmune disease characterized by optic neuritis and transverse myelitis, with an unclear genetic background. A genome-wide meta-analysis of NMOSD in Japanese individuals (240 patients and 50,578 controls) identified significant associations with the major histocompatibility complex region and a common variant close to CCR6 (rs12193698; p = 1.8 × 10-8, odds ratio [OR] = 1.73). In single-cell RNA sequencing (scRNA-seq) analysis (25 patients and 101 controls), the CCR6 risk variant showed disease-specific expression quantitative trait loci effects in CD4+ T (CD4T) cell subsets. Furthermore, we detected somatic mosaic chromosomal alterations (mCAs) in various autoimmune diseases and found that mCAs increase the risk of NMOSD (OR = 3.37 for copy number alteration). In scRNA-seq data, CD4T cells with 21q loss, a recurrently observed somatic event in NMOSD, showed dysregulation of type I interferon-related genes. Our integrated study identified novel germline and somatic mutations associated with NMOSD pathogenesis.Germline variants and mosaic chromosomal alterations affect COVID-19 vaccine immunogenicity. Cell Genomics 5, 100783 (2025).
Vaccine immunogenicity is influenced by the vaccinee’s genetic background. Here, we perform a genome-wide association study of vaccine-induced SARS-CoV-2-specific immunoglobulin G (IgG) antibody titers and T cell immune responses in 1,559 mRNA-1273 and 537 BNT162b2 vaccinees of Japanese ancestry. SARS-CoV-2-specific antibody titers are associated with the immunoglobulin heavy chain (IGH) and major histocompatibility complex (MHC) locus, and T cell responses are associated with MHC. The lead variants at IGH contain a population-specific missense variant (rs1043109-C; p.Leu192Val) in the immunoglobulin heavy constant gamma 1 gene (IGHG1), with a strong decreasing effect (β = -0.54). Antibody-titer-associated variants modulate circulating immune regulatory proteins (e.g., LILRB4 and FCRL6). Age-related hematopoietic expanded mosaic chromosomal alterations (mCAs) affecting MHC and IGH also impair antibody production. MHC-/IGH-affecting mCAs confer infectious and immune disease risk, including sepsis and Graves’ disease. Impacts of expanded mosaic loss of chromosomes X/Y on these phenotypes were examined. Altogether, both germline and somatic mutations contribute to adaptive immunity functions.Blood DNA virome associates with autoimmune diseases and COVID-19. Nature Genetics 57, 65–79 (2025).
Aberrant immune responses to viral pathogens contribute to pathogenesis, but our understanding of pathological immune responses caused by viruses within the human virome, especially at a population scale, remains limited. We analyzed whole-genome sequencing datasets of 6,321 Japanese individuals, including patients with autoimmune diseases (psoriasis vulgaris, rheumatoid arthritis (RA), systemic lupus erythematosus (SLE), pulmonary alveolar proteinosis (PAP) or multiple sclerosis) and coronavirus disease 2019 (COVID-19), or healthy controls. We systematically quantified two constituents of the blood DNA virome, endogenous HHV-6 (eHHV-6) and anellovirus. Participants with eHHV-6B had higher risks of SLE and PAP; the former was validated in All of Us. eHHV-6B-positivity and high SLE disease activity index scores had strong correlations. Genome-wide association study and long-read sequencing mapped the integration of the HHV-6B genome to a locus on chromosome 22q. Epitope mapping and single-cell RNA sequencing revealed distinctive immune induction by eHHV-6B in patients with SLE. In addition, high anellovirus load correlated strongly with SLE, RA and COVID-19 status. Our analyses unveil relationships between the human virome and autoimmune and infectious diseases.Kato, C., Morimoto, S., Takahashi, S., Namba, S., Wang, Q. S., Okada, Y. & Okano, H. Spinal cord motor neuron phenotypes and polygenic risk scores in sporadic amyotrophic lateral sclerosis: deciphering the disease pathology and therapeutic potential of ropinirole hydrochloride. Journal of Neurology, Neurosurgery & Psychiatry 96, 199–201 (2025).
Genetic legacy of ancient hunter-gatherer Jomon in Japanese populations. Nature Communications 15, 9780 (2024).
The tripartite ancestral structure is a recently proposed model for the genetic origin of modern Japanese, comprising indigenous Jomon hunter-gatherers and two additional continental ancestors from Northeast Asia and East Asia. To investigate the impact of the tripartite structure on genetic and phenotypic variation today, we conducted biobank-scale analyses by merging Biobank Japan (BBJ; n = 171,287) with ancient Japanese and Eurasian genomes (n = 22). We demonstrate the applicability of the tripartite model to Japanese populations throughout the archipelago, with an extremely strong correlation between Jomon ancestry and genomic variation among individuals. We also find that the genetic legacy of Jomon ancestry underlies an elevated body mass index (BMI). Genome-wide association analysis with rigorous adjustments for geographical and ancestral substructures identifies 132 variants that are informative for predicting individual Jomon ancestry. This prediction model is validated using independent Japanese cohorts (Nagahama cohort, n = 2993; the second cohort of BBJ, n = 72,695). We further confirm the phenotypic association between Jomon ancestry and BMI using East Asian individuals from UK Biobank (n = 2286). Our extensive analysis of ancient and modern genomes, involving over 250,000 participants, provides valuable insights into the genetic legacy of ancient hunter-gatherers in contemporary populations.Inconsistent embryo selection across polygenic score methods. Nature Human Behaviour 2264–2267 (2024).
Private enterprises offer preimplantation genetic testing with polygenic scores to select embryos with ‘desirable’ potential. In silico simulations using biobank resources show that the selected embryo would rely substantially on the choice of polygenic score method and randomness in score construction, which raises ethical concerns.Statistically and functionally fine-mapped blood eQTLs and pQTLs from 1,405 humans reveal distinct regulation patterns and disease relevance. Nature Genetics 2054–2067 (2024).
Studying the genetic regulation of protein expression (through protein quantitative trait loci (pQTLs)) offers a deeper understanding of regulatory variants uncharacterized by mRNA expression regulation (expression QTLs (eQTLs)) studies. Here we report cis-eQTL and cis-pQTL statistical fine-mapping from 1,405 genotyped samples with blood mRNA and 2,932 plasma samples of protein expression, as part of the Japan COVID-19 Task Force (JCTF). Fine-mapped eQTLs (n = 3,464) were enriched for 932 variants validated with a massively parallel reporter assay. Fine-mapped pQTLs (n = 582) were enriched for missense variations on structured and extracellular domains, although the possibility of epitope-binding artifacts remains. Trans-eQTL and trans-pQTL analysis highlighted associations of class I HLA allele variation with KIR genes. We contrast the multi-tissue origin of plasma protein with blood mRNA, contributing to the limited colocalization level, distinct regulatory mechanisms and trait relevance of eQTLs and pQTLs. We report a negative correlation between ABO mRNA and protein expression because of linkage disequilibrium between distinct nearby eQTLs and pQTLs.Machine learning reveals heterogeneous associations between environmental factors and cardiometabolic diseases across polygenic risk scores. Communications Medicine 4, 181 (2024).
Background: Although polygenic risk scores (PRSs) are expected to be helpful in precision medicine, it remains unclear whether high-PRS groups are more likely to benefit from preventive interventions for diseases. Recent methodological advancements enable us to predict treatment effects at the individual level. Methods: We employed causal forest to explore the relationship between PRSs and individual risk of diseases associated with certain environmental factors. Following simulations illustrating its performance, we applied our approach to investigate the individual risk of cardiometabolic diseases, including coronary artery diseases (CAD) and type 2 diabetes (T2D), associated with obesity and smoking among individuals from UK Biobank (UKB; n = 369,942) and BioBank Japan (BBJ; n = 149,421). Results: Here we find the heterogeneous association of obesity and smoking with diseases across PRS values, complicated by the multi-dimensional combination of individual characteristics such as age and sex. The highest positive correlations of PRSs and the exposure-related disease risks are observed between obesity and T2D in UKB and between smoking and CAD in BBJ (Spearman’s ρ = 0.61 and 0.32, respectively). However, most relationships are weak or negative, suggesting that high-PRS groups will not necessarily benefit most from environmental factor prevention. Conclusions: Our study highlights the importance of individual-level prediction of disease risks associated with target exposure in precision medicine. This study aimed to understand if people with a high genetic risk for certain diseases benefit more from preventive strategies. Using a machine-learning-based method, we analyzed data from large groups of people in the UK and Japan. We examined the risk of heart and metabolic diseases in relation to obesity and smoking. The results showed that the link between genetic risk and disease is complex and varies widely among individuals. Our results suggested that those with a high genetic risk for disease may not always benefit more from the prevention of obesity and smoking. This finding suggests that we need to consider more than risk in decisions on how to prevent diseases in individuals.Quantification of escape from X chromosome inactivation with single-cell omics data reveals heterogeneity across cell types and tissues. Cell Genomics 4, 100625 (2024).
Several X-linked genes escape from X chromosome inactivation (XCI), while differences in escape across cell types and tissues are still poorly characterized. Here, we developed scLinaX for directly quantifying relative gene expression from the inactivated X chromosome with droplet-based single-cell RNA sequencing (scRNA-seq) data. The scLinaX and differentially expressed gene analyses with large-scale blood scRNA-seq datasets consistently identified the stronger escape in lymphocytes than in myeloid cells. An extension of scLinaX to a 10x multiome dataset (scLinaX-multi) suggested a stronger escape in lymphocytes than in myeloid cells at the chromatin-accessibility level. The scLinaX analysis of human multiple-organ scRNA-seq datasets also identified the relatively strong degree of escape from XCI in lymphoid tissues and lymphocytes. Finally, effect size comparisons of genome-wide association studies between sexes suggested the underlying impact of escape on the genotype-phenotype association. Overall, scLinaX and the quantified escape catalog identified the heterogeneity of escape across cell types and tissues.Body mass index stratification optimizes polygenic prediction of type 2 diabetes in cross-biobank analyses. Nature Genetics 56, 1100–1109 (2024).
Type 2 diabetes (T2D) shows heterogeneous body mass index (BMI) sensitivity. Here, we performed stratification based on BMI to optimize predictions for BMI-related diseases. We obtained BMI-stratified datasets using data from more than 195,000 individuals (nT2D = 55,284) from BioBank Japan (BBJ) and UK Biobank. T2D heritability in the low-BMI group was greater than that in the high-BMI group. Polygenic predictions of T2D toward low-BMI targets had pseudo-R2 values that were more than 22% higher than BMI-unstratified targets. Polygenic risk scores (PRSs) from low-BMI discovery outperformed PRSs from high BMI, while PRSs from BMI-unstratified discovery performed best. Pathway-specific PRSs demonstrated the biological contributions of pathogenic pathways. Low-BMI T2D cases showed higher rates of neuropathy and retinopathy. Combining BMI stratification and a method integrating cross-population effects, T2D predictions showed greater than 37% improvements over unstratified-matched-population prediction. We replicated findings in the Tohoku Medical Megabank (n = 26,000) and the second BBJ cohort (n = 33,096). Our findings suggest that target stratification based on existing traits can improve the polygenic prediction of heterogeneous diseases.Poor accuracy and sustainability of the first-step FIB4 EASL pathway for stratifying steatotic liver disease risk in the general population. Alimentary Pharmacology & Therapeutics 59, 1402–1412 (2024).
Summary: Background and Aims: The European Association for the Study of the Liver introduced a clinical pathway (EASL CP) for screening significant/advanced fibrosis in people at risk of steatotic liver disease (SLD). We assessed the performance of the first‐step FIB4 EASL CP in the general population across different SLD risk groups (MASLD, Met‐ALD and ALD) and various age classes.Methods: We analysed a total of 3372 individuals at risk of SLD from the 2017–2018 National Health and Nutrition Examination Survey (NHANES17‐18), projected to 152.3 million U.S. adults, 300,329 from the UK Biobank (UKBB) and 57,644 from the Biobank Japan (BBJ). We assessed liver stiffness measurement (LSM) ≥8 kPa and liver‐related events occurring within 3 and 10 years (3/10 year‐LREs) as outcomes. We defined MASLD, MetALD, and ALD according to recent international recommendations.Results: FIB4 sensitivity for LSM ≥ 8 kPa was low (27.7%), but it ranged approximately 80%‐90% for 3‐year LREs. Using FIB4, 22%–57% of subjects across the three cohorts were identified as candidates for vibration‐controlled transient elastography (VCTE), which was mostly avoidable (positive predictive value of FIB4 ≥ 1.3 for LSM ≥ 8 kPa ranging 9.5%–13% across different SLD categories). Sensitivity for LSM ≥ 8 kPa and LREs increased with increasing alcohol intake (ALD>MetALD>MASLD) and age classes. For individuals aged ≥65 years, using the recommended age‐adjusted FIB4 cut‐off (≥2) substantially reduced sensitivity for LSM ≥ 8 kPa and LREs.Conclusions: The first‐step FIB4 EASL CP is poorly accurate and feasible for individuals at risk of SLD in the general population. It is crucial to enhance the screening strategy with a first‐step approach able to reduce unnecessary VCTEs and optimise their yield.Novel ancestry-specific primary open-angle glaucoma loci and shared biology with vascular mechanisms and cell proliferation. Cell Reports Medicine 5, 101430 (2024).
Primary open-angle glaucoma (POAG), a leading cause of irreversible blindness globally, shows disparity in prevalence and manifestations across ancestries. We perform meta-analysis across 15 biobanks (of the Global Biobank Meta-analysis Initiative) (n = 1,487,441: cases = 26,848) and merge with previous multi-ancestry studies, with the combined dataset representing the largest and most diverse POAG study to date (n = 1,478,037: cases = 46,325) and identify 17 novel significant loci, 5 of which were ancestry specific. Gene-enrichment and transcriptome-wide association analyses implicate vascular and cancer genes, a fifth of which are primary ciliary related. We perform an extensive statistical analysis of SIX6 and CDKN2B-AS1 loci in human GTEx data and across large electronic health records showing interaction between SIX6 gene and causal variants in the chr9p21.3 locus, with expression effect on CDKN2A/B. Our results suggest that some POAG risk variants may be ancestry specific, sex specific, or both, and support the contribution of genes involved in programmed cell death in POAG pathogenesis.Genetic drivers of heterogeneity in type 2 diabetes pathophysiology. Nature 627, 347–357 (2024).
Type 2 diabetes (T2D) is a heterogeneous disease that develops through diverse pathophysiological processes 1,2 and molecular mechanisms that are often specific to cell type 3,4. Here, to characterize the genetic contribution to these processes across ancestry groups, we aggregate genome-wide association study data from 2,535,601 individuals (39.7% not of European ancestry), including 428,452 cases of T2D. We identify 1,289 independent association signals at genome-wide significance ( P < 5 × 10 −8) that map to 611 loci, of which 145 loci are, to our knowledge, previously unreported. We define eight non-overlapping clusters of T2D signals that are characterized by distinct profiles of cardiometabolic trait associations. These clusters are differentially enriched for cell-type-specific regions of open chromatin, including pancreatic islets, adipocytes, endothelial cells and enteroendocrine cells. We build cluster-specific partitioned polygenic scores 5 in a further 279,552 individuals of diverse ancestry, including 30,288 cases of T2D, and test their association with T2D-related vascular outcomes. Cluster-specific partitioned polygenic scores are associated with coronary artery disease, peripheral artery disease and end-stage diabetic nephropathy across ancestry groups, highlighting the importance of obesity-related processes in the development of vascular outcomes. Our findings show the value of integrating multi-ancestry genome-wide association study data with single-cell epigenomics to disentangle the aetiological heterogeneity that drives the development and progression of T2D. This might offer a route to optimize global access to genetically informed diabetes care.Genetic architecture of alcohol consumption identified by a genotype-stratified GWAS and impact on esophageal cancer risk in Japanese people. Science Advances 10, eade2780 (2024).
An East Asian–specific variant on aldehyde dehydrogenase 2 ( ALDH2 rs671, G>A) is the major genetic determinant of alcohol consumption. We performed an rs671 genotype-stratified genome-wide association study meta-analysis of alcohol consumption in 175,672 Japanese individuals to explore gene-gene interactions with rs671 behind drinking behavior. The analysis identified three genome-wide significant loci ( GCKR, KLB, and ADH1B) in wild-type homozygotes and six ( GCKR, ADH1B, ALDH1B1, ALDH1A1, ALDH2, and GOT2) in heterozygotes, with five showing genome-wide significant interaction with rs671. Genetic correlation analyses revealed ancestry-specific genetic architecture in heterozygotes. Of the discovered loci, four ( GCKR, ADH1B, ALDH1A1, and ALDH2) were suggested to interact with rs671 in the risk of esophageal cancer, a representative alcohol-related disease. Our results identify the genotype-specific genetic architecture of alcohol consumption and reveal its potential impact on alcohol-related disease risk.Boosting the power of genome-wide association studies within and across ancestries by using polygenic scores. Nature Genetics 55, 1769–1776 (2023).
Genome-wide association studies (GWASs) have been mostly conducted in populations of European ancestry, which currently limits the transferability of their findings to other populations. Here, we show, through theory, simulations and applications to real data, that adjustment of GWAS analyses for polygenic scores (PGSs) increases the statistical power for discovery across all ancestries. We applied this method to analyze seven traits available in three large biobanks with participants of East Asian ancestry (n = 340,000 in total) and report 139 additional associations across traits. We also present a two-stage meta-analysis strategy whereby, in contributing cohorts, a PGS-adjusted GWAS is rerun using PGSs derived from a first round of a standard meta-analysis. On average, across traits, this approach yields a 1.26-fold increase in the number of detected associations (range 1.07- to 1.76-fold increase). Altogether, our study demonstrates the value of using PGSs to increase the power of GWASs in underrepresented populations and promotes such an analytical strategy for future GWAS meta-analyses.Pan-cancer and cross-population genome-wide association studies dissect shared genetic backgrounds underlying carcinogenesis. Nature Communications 14, 3671 (2023).
Integrating genomic data of multiple cancers allows de novo cancer grouping and elucidating the shared genetic basis across cancers. Here, we conduct the pan-cancer and cross-population genome-wide association study (GWAS) meta-analysis and replication studies on 13 cancers including 250,015 East Asians (Biobank Japan) and 377,441 Europeans (UK Biobank). We identify ten cancer risk variants including five pleiotropic associations (e.g., rs2076295 at DSP on 6p24 associated with lung cancer and rs2525548 at TRIM4 on 7q22 nominally associated with six cancers). Quantifying shared heritability among the cancers detects positive genetic correlations between breast and prostate cancer across populations. Common genetic components increase the statistical power, and the large-scale meta-analysis of 277,896 breast/prostate cancer cases and 901,858 controls identifies 91 newly genome-wide significant loci. Enrichment analysis of pathways and cell types reveals shared genetic backgrounds across said cancers. Focusing on genetically correlated cancers can contribute to enhancing our insights into carcinogenesis.Extracting immunological and clinical heterogeneity across autoimmune rheumatic diseases by cohort-wide immunophenotyping. Annals of the Rheumatic Diseases 83, 242–252 (2023).
Objective: Extracting immunological and clinical heterogeneity across autoimmune rheumatic diseases (AIRDs) is essential towards personalised medicine. Methods: We conducted large-scale and cohort-wide immunophenotyping of 46 peripheral immune cells using Human Immunology Protocol of comprehensive 8-colour flow cytometric analysis. Dataset consisted of >1000 Japanese patients of 11 AIRDs with deep clinical information registered at the FLOW study, including rheumatoid arthritis (RA) and systemic lupus erythematosus (SLE). In-depth clinical and immunological characterisation was conducted for the identified RA patient clusters, including associations of inborn human genetics represented by Polygenic Risk Score (PRS). Results: Multimodal clustering of immunophenotypes deciphered underlying disease-cell type network in immune cell, disease and patient cluster resolutions. This provided immune cell type specificity shared or distinct across AIRDs, such as close immunological network between mixed connective tissue disease and SLE. Individual patient-level clustering dissected patients with AIRD into several clusters with different immunological features. Of these, RA-like or SLE-like clusters were exclusively dominant, showing immunological differentiation between RA and SLE across AIRDs. In-depth clinical analysis of RA revealed that such patient clusters differentially defined clinical heterogeneity in disease activity and treatment responses, such as treatment resistance in patients with RA with SLE-like immunophenotypes. PRS based on RA case-control and within-case stratified genome-wide association studies were associated with clinical and immunological characteristics. This pointed immune cell type implicated in disease biology such as dendritic cells for RA-interstitial lung disease. Conclusion: Cohort-wide and cross-disease immunophenotyping elucidate clinically heterogeneous patient subtypes existing within single disease in immune cell type-specific manner.Multifaceted analysis of cross-tissue transcriptomes reveals phenotype–endotype associations in atopic dermatitis. Nature Communications 14, 6133 (2023).
Atopic dermatitis (AD) is a skin disease that is heterogeneous both in terms of clinical manifestations and molecular profiles. It is increasingly recognized that AD is a systemic rather than a local disease and should be assessed in the context of whole-body pathophysiology. Here we show, via integrated RNA-sequencing of skin tissue and peripheral blood mononuclear cell (PBMC) samples along with clinical data from 115 AD patients and 14 matched healthy controls, that specific clinical presentations associate with matching differential molecular signatures. We establish a regression model based on transcriptome modules identified in weighted gene co-expression network analysis to extract molecular features associated with detailed clinical phenotypes of AD. The two main, qualitatively differential skin manifestations of AD, erythema and papulation are distinguished by differential immunological signatures. We further apply the regression model to a longitudinal dataset of 30 AD patients for personalized monitoring, highlighting patient heterogeneity in disease trajectories. The longitudinal features of blood tests and PBMC transcriptome modules identify three patient clusters which are aligned with clinical severity and reflect treatment history. Our approach thus serves as a framework for effective clinical investigation to gain a holistic view on the pathophysiology of complex human diseases.Single-cell analyses and host genetics highlight the role of innate immune cells in COVID-19 severity. Nature Genetics 55, 753–767 (2023).
Mechanisms underpinning the dysfunctional immune response in severe acute respiratory syndrome coronavirus 2 infection are elusive. We analyzed single-cell transcriptomes and T and B cell receptors (BCR) of >895,000 peripheral blood mononuclear cells from 73 coronavirus disease 2019 (COVID-19) patients and 75 healthy controls of Japanese ancestry with host genetic data. COVID-19 patients showed a low fraction of nonclassical monocytes (ncMono). We report downregulated cell transitions from classical monocytes to ncMono in COVID-19 with reduced CXCL10 expression in ncMono in severe disease. Cell–cell communication analysis inferred decreased cellular interactions involving ncMono in severe COVID-19. Clonal expansions of BCR were evident in the plasmablasts of patients. Putative disease genes identified by COVID-19 genome-wide association study showed cell type-specific expressions in monocytes and dendritic cells. A COVID-19-associated risk variant at the IFNAR2 locus (rs13050728) had context-specific and monocyte-specific expression quantitative trait loci effects. Our study highlights biological and host genetic involvement of innate immune cells in COVID-19 severity.Global Biobank analyses provide lessons for developing polygenic risk scores across diverse cohorts. Cell Genomics 3, 100241 (2023).
Polygenic risk scores (PRSs) have been widely explored in precision medicine. However, few studies have thoroughly investigated their best practices in global populations across different diseases. We here utilized data from Global Biobank Meta-analysis Initiative (GBMI) to explore methodological considerations and PRS performance in 9 different biobanks for 14 disease endpoints. Specifically, we constructed PRSs using pruning and thresholding (P + T) and PRS-continuous shrinkage (CS). For both methods, using a European-based linkage disequilibrium (LD) reference panel resulted in comparable or higher prediction accuracy compared with several other non-European-based panels. PRS-CS overall outperformed the classic P + T method, especially for endpoints with higher SNP-based heritability. Notably, prediction accuracy is heterogeneous across endpoints, biobanks, and ancestries, especially for asthma, which has known variation in disease prevalence across populations. Overall, we provide lessons for PRS construction, evaluation, and interpretation using GBMI resources and highlight the importance of best practices for PRS in the biobank-scale genomics era.Multi-ancestry meta-analysis of asthma identifies novel associations and highlights the value of increased power and diversity. Cell Genomics 2, 100212 (2022).
Asthma is a complex disease that varies widely in prevalence across populations. The extent to which genetic variation contributes to these disparities is unclear, as the genetics underlying asthma have been investigated primarily in populations of European descent. As part of the Global Biobank Meta-analysis Initiative, we conducted a large-scale genome-wide association study of asthma (153,763 cases and 1,647,022 controls) via meta-analysis across 22 biobanks spanning multiple ancestries. We discovered 179 asthma-associated loci, 49 of which were not previously reported. Despite the wide range in asthma prevalence among biobanks, we found largely consistent genetic effects across biobanks and ancestries. The meta-analysis also improved polygenic risk prediction in non-European populations compared with previous studies. Additionally, we found considerable genetic overlap between age-of-onset subtypes and between asthma and comorbid diseases. Our work underscores the multi-factorial nature of asthma development and offers insight into its shared genetic architecture.Common germline risk variants impact somatic alterations and clinical features across cancers. Cancer Research 83, 20–27 (2022).
Aggregation of genome-wide common risk variants, such as polygenic risk score (PRS), can measure genetic susceptibility to cancer. A better understanding of how common germline variants associate with somatic alterations and clinical features could facilitate personalized cancer prevention and early detection. We constructed PRSs from 14 genome-wide association studies (median n = 64,905) for 12 cancer types by multiple methods and calibrated them using the UK Biobank resources (n = 335,048). Meta-analyses across cancer types in The Cancer Genome Atlas (n = 7,965) revealed that higher PRS values were associated with earlier cancer onset and lower burden of somatic alterations, including total mutations, chromosome/arm somatic copy-number alterations (SCNA), and focal SCNAs. This contrasts with rare germline pathogenic variants (e.g., BRCA1/2 variants), showing heterogeneous associations with somatic alterations. Our results suggest that common germline cancer risk variants allow early tumor development before the accumulation of many somatic alterations characteristic of later stages of carcinogenesis. Significance:: Meta-analyses across cancers show that common germline risk variants affect not only cancer predisposition but the age of cancer onset and burden of somatic alterations, including total mutations and copy-number alterations.Global Biobank Meta-analysis Initiative: Powering genetic discovery across human disease. Cell Genomics 2, 100192 (2022).
Biobanks facilitate genome-wide association studies (GWASs), which have mapped genomic loci across a range of human diseases and traits. However, most biobanks are primarily composed of individuals of European ancestry. We introduce the Global Biobank Meta-analysis Initiative (GBMI)-a collaborative network of 23 biobanks from 4 continents representing more than 2.2 million consented individuals with genetic data linked to electronic health records. GBMI meta-analyzes summary statistics from GWASs generated using harmonized genotypes and phenotypes from member biobanks for 14 exemplar diseases and endpoints. This strategy validates that GWASs conducted in diverse biobanks can be integrated despite heterogeneity in case definitions, recruitment strategies, and baseline characteristics. This collaborative effort improves GWAS power for diseases, benefits understudied diseases, and improves risk prediction while also enabling the nomination of disease genes and drug candidates by incorporating gene and protein expression data and providing insight into the underlying biology of human diseases and traits.*Namba, S., *Konuma, T., Wu, K.-H., Zhou, W. & Okada, Y. A practical guideline of genomics-driven drug discovery in the era of global biobank meta-analysis. Cell Genomics 2, 100190 (2022).
Genomics-driven drug discovery is indispensable for accelerating the development of novel therapeutic targets. However, the drug discovery framework based on evidence from genome-wide association studies (GWASs) has not been established, especially for cross-population GWAS meta-analysis. Here, we introduce a practical guideline for genomics-driven drug discovery for cross-population meta-analysis, as lessons from the Global Biobank Meta-analysis Initiative (GBMI). Our drug discovery framework encompassed three methodologies and was applied to the 13 common diseases targeted by GBMI (N mean = 1,329,242). Individual methodologies complementarily prioritized drugs and drug targets, which were systematically validated by referring previously known drug-disease relationships. Integration of the three methodologies provided a comprehensive catalog of candidate drugs for repositioning, nominating promising drug candidates targeting the genes involved in the coagulation process for venous thromboembolism and the interleukin-4 and interleukin-13 signaling pathway for gout. Our study highlighted key factors for successful genomics-driven drug discovery using cross-population meta-analyses.Stroke genetics informs drug discovery and risk prediction across ancestries. Nature 611, 115–123 (2022).
Previous genome-wide association studies (GWASs) of stroke - the second leading cause of death worldwide - were conducted predominantly in populations of European ancestry1,2. Here, in cross-ancestry GWAS meta-analyses of 110,182 patients who have had a stroke (five ancestries, 33% non-European) and 1,503,898 control individuals, we identify association signals for stroke and its subtypes at 89 (61 new) independent loci: 60 in primary inverse-variance-weighted analyses and 29 in secondary meta-regression and multitrait analyses. On the basis of internal cross-ancestry validation and an independent follow-up in 89,084 additional cases of stroke (30% non-European) and 1,013,843 control individuals, 87% of the primary stroke risk loci and 60% of the secondary stroke risk loci were replicated (P < 0.05). Effect sizes were highly correlated across ancestries. Cross-ancestry fine-mapping, in silico mutagenesis analysis3, and transcriptome-wide and proteome-wide association analyses revealed putative causal genes (such as SH3PXD2A and FURIN) and variants (such as at GRK5 and NOS3). Using a three-pronged approach4, we provide genetic evidence for putative drug effects, highlighting F11, KLKB1, PROC, GP1BA, LAMC2 and VCAM1 as possible targets, with drugs already under investigation for stroke for F11 and PROC. A polygenic score integrating cross-ancestry and ancestry-specific stroke GWASs with vascular-risk factor GWASs (integrative polygenic scores) strongly predicted ischaemic stroke in populations of European, East Asian and African ancestry5. Stroke genetic risk scores were predictive of ischaemic stroke independent of clinical risk factors in 52,600 clinical-trial participants with cardiometabolic disease. Our results provide insights to inform biology, reveal potential drug targets and derive genetic risk prediction tools across ancestries.Genetic footprints of assortative mating in the Japanese population. Nature human behaviour 7, 65–73 (2022).
Assortative mating (AM) is a pattern characterized by phenotypic similarities between mating partners. Detecting the evidence of AM has been challenging due to the lack of large-scale datasets that include phenotypic data on both partners, especially in populations of non-European ancestries. Gametic phase disequilibrium between trait-associated alleles is a signature of parental AM on a polygenic trait, which can be detected even without partner data. Here, using polygenic scores for 81 traits in the Japanese population using BioBank Japan Project genome-wide association studies data (n = 172,270), we found evidence of AM on the liability to type 2 diabetes and coronary artery disease, as well as on dietary habits. In cross-population comparison using United Kingdom Biobank data (n = 337,139) we found shared but heterogeneous impacts of AM between populations.The whole blood transcriptional regulation landscape in 465 COVID-19 infected samples from Japan COVID-19 Task Force. Nature communications 13, 4830 (2022).
Coronavirus disease 2019 (COVID-19) is a recently-emerged infectious disease that has caused millions of deaths, where comprehensive understanding of disease mechanisms is still unestablished. In particular, studies of gene expression dynamics and regulation landscape in COVID-19 infected individuals are limited. Here, we report on a thorough analysis of whole blood RNA-seq data from 465 genotyped samples from the Japan COVID-19 Task Force, including 359 severe and 106 non-severe COVID-19 cases. We discover 1169 putative causal expression quantitative trait loci (eQTLs) including 34 possible colocalizations with biobank fine-mapping results of hematopoietic traits in a Japanese population, 1549 putative causal splice QTLs (sQTLs; e.g. two independent sQTLs at TOR1AIP1), as well as biologically interpretable trans-eQTL examples (e.g., REST and STING1), all fine-mapped at single variant resolution. We perform differential gene expression analysis to elucidate 198 genes with increased expression in severe COVID-19 cases and enriched for innate immune-related functions. Finally, we evaluate the limited but non-zero effect of COVID-19 phenotype on eQTL discovery, and highlight the presence of COVID-19 severity-interaction eQTLs (ieQTLs; e.g., CLEC4C and MYBL2). Our study provides a comprehensive catalog of whole blood regulatory variants in Japanese, as well as a reference for transcriptional landscapes in response to COVID-19 infection.DOCK2 is involved in the host genetics and biology of severe COVID-19. Nature 609, 754–760 (2022).
Identifying the host genetic factors underlying severe COVID-19 is an emerging challenge1-5. Here we conducted a genome-wide association study (GWAS) involving 2,393 cases of COVID-19 in a cohort of Japanese individuals collected during the initial waves of the pandemic, with 3,289 unaffected controls. We identified a variant on chromosome 5 at 5q35 (rs60200309-A), close to the dedicator of cytokinesis 2 gene (DOCK2), which was associated with severe COVID-19 in patients less than 65 years of age. This risk allele was prevalent in East Asian individuals but rare in Europeans, highlighting the value of genome-wide association studies in non-European populations. RNA-sequencing analysis of 473 bulk peripheral blood samples identified decreased expression of DOCK2 associated with the risk allele in these younger patients. DOCK2 expression was suppressed in patients with severe cases of COVID-19. Single-cell RNA-sequencing analysis (n = 61 individuals) identified cell-type-specific downregulation of DOCK2 and a COVID-19-specific decreasing effect of the risk allele on DOCK2 expression in non-classical monocytes. Immunohistochemistry of lung specimens from patients with severe COVID-19 pneumonia showed suppressed DOCK2 expression. Moreover, inhibition of DOCK2 function with CPYPP increased the severity of pneumonia in a Syrian hamster model of SARS-CoV-2 infection, characterized by weight loss, lung oedema, enhanced viral loads, impaired macrophage recruitment and dysregulated type I interferon responses. We conclude that DOCK2 has an important role in the host immune response to SARS-CoV-2 infection and the development of severe COVID-19, and could be further explored as a potential biomarker and/or therapeutic target.Multi-trait and cross-population genome-wide association studies across autoimmune and allergic diseases identify shared and distinct genetic component. Annals of the rheumatic diseases 81, 1301–1312 (2022).
OBJECTIVES: Autoimmune and allergic diseases are outcomes of the dysregulation of the immune system. Our study aimed to elucidate differences or shared components in genetic backgrounds between autoimmune and allergic diseases. METHODS: We estimated genetic correlation and performed multi-trait and cross-population genome-wide association study (GWAS) meta-analysis of six immune-related diseases: rheumatoid arthritis, Graves’ disease, type 1 diabetes for autoimmune diseases and asthma, atopic dermatitis and pollinosis for allergic diseases. By integrating large-scale biobank resources (Biobank Japan and UK biobank), our study included 105,721 cases and 433,663 controls. Newly identified variants were evaluated in 21,778 cases and 712,767 controls for two additional autoimmune diseases: psoriasis and systemic lupus erythematosus. We performed enrichment analyses of cell types and biological pathways to highlight shared and distinct perspectives. RESULTS: Autoimmune and allergic diseases were not only mutually classified based on genetic backgrounds but also they had multiple positive genetic correlations beyond the classifications. Multi-trait GWAS meta-analysis newly identified six allergic disease-associated loci. We identified four loci shared between the six autoimmune and allergic diseases (rs10803431 at PRDM2, OR=1.07, p=2.3×10-8, rs2053062 at G3BP1, OR=0.90, p=2.9×10-8, rs2210366 at HBS1L, OR=1.07, p=2.5×10-8 in Japanese and rs4529910 at POU2AF1, OR=0.96, p=1.9×10-10 across ancestries). Associations of rs10803431 and rs4529910 were confirmed at the two additional autoimmune diseases. Enrichment analysis demonstrated link to T cells, natural killer cells and various cytokine signals, including innate immune pathways. CONCLUSION: Our multi-trait and cross-population study should elucidate complex pathogenesis shared components across autoimmune and allergic diseases.Polygenic risk scores for prediction of breast cancer risk in Asian populations. Genetics in medicine 24, 586–600 (2022).
PURPOSE: Non-European populations are under-represented in genetics studies, hindering clinical implementation of breast cancer polygenic risk scores (PRSs). We aimed to develop PRSs using the largest available studies of Asian ancestry and to assess the transferability of PRS across ethnic subgroups. METHODS: The development data set comprised 138,309 women from 17 case-control studies. PRSs were generated using a clumping and thresholding method, lasso penalized regression, an Empirical Bayes approach, a Bayesian polygenic prediction approach, or linear combinations of multiple PRSs. These PRSs were evaluated in 89,898 women from 3 prospective studies (1592 incident cases). RESULTS: The best performing PRS (genome-wide set of single-nucleotide variations [formerly single-nucleotide polymorphism]) had a hazard ratio per unit SD of 1.62 (95% CI = 1.46-1.80) and an area under the receiver operating curve of 0.635 (95% CI = 0.622-0.649). Combined Asian and European PRSs (333 single-nucleotide variations) had a hazard ratio per SD of 1.53 (95% CI = 1.37-1.71) and an area under the receiver operating curve of 0.621 (95% CI = 0.608-0.635). The distribution of the latter PRS was different across ethnic subgroups, confirming the importance of population-specific calibration for valid estimation of breast cancer risk. CONCLUSION: PRSs developed in this study, from association data from multiple ancestries, can enhance risk stratification for women of Asian ancestry.Transcript-targeted analysis reveals isoform alterations and double-hop fusions in breast cancer. Communications biology 4, 1320 (2021).
Although transcriptome alteration is an essential driver of carcinogenesis, the effects of chromosomal structural alterations on the cancer transcriptome are not yet fully understood. Short-read transcript sequencing has prevented researchers from directly exploring full-length transcripts, forcing them to focus on individual splice sites. Here, we develop a pipeline for Multi-Sample long-read Transcriptome Assembly (MuSTA), which enables construction of a transcriptome from long-read sequence data. Using the constructed transcriptome as a reference, we analyze RNA extracted from 22 clinical breast cancer specimens. We identify a comprehensive set of subtype-specific and differentially used isoforms, which extended our knowledge of isoform regulation to unannotated isoforms including a short form TNS3. We also find that the exon-intron structure of fusion transcripts depends on their genomic context, and we identify double-hop fusion transcripts that are transcribed from complex structural rearrangements. For example, a double-hop fusion results in aberrant expression of an endogenous retroviral gene, ERVFRD-1, which is normally expressed exclusively in placenta and is thought to protect fetus from maternal rejection; expression is elevated in several TCGA samples with ERVFRD-1 fusions. Our analyses provide direct evidence that full-length transcript sequencing of clinical samples can add to our understanding of cancer biology and genomics in general.Differential regulation of CpG island methylation within divergent and unidirectional promoters in colorectal cancer. Cancer science 110, 1096–1104 (2019).
The silencing of tumor suppressor genes by promoter CpG island (CGI) methylation is an important cause of oncogenesis. Silencing of MLH1 and BRCA1, two examples of oncogenic events, results from promoter CGI methylation. Interestingly, both MLH1 and BRCA1 have a divergent promoter, from which another gene on the opposite strand is also transcribed. Although studies have shown that divergent transcription is an important factor in transcriptional regulation, little is known about its implication in aberrant promoter methylation in cancer. In this study, we analyzed the methylation status of CGI in divergent promoters using a recently enriched transcriptome database. We measured the extent of CGI methylation in 119 colorectal cancer (CRC) clinical samples (65 microsatellite instability high [MSI-H] CRC with CGI methylator phenotype, 28 MSI-H CRC without CGI methylator phenotype and 26 microsatellite stable CRC) and 21 normal colorectal tissues using Infinium MethylationEPIC BeadChip. We found that CGI within divergent promoters are less frequently methylated than CGI within unidirectional promoters in normal cells. In the genome of CRC cells, CGI within unidirectional promoters are more vulnerable to aberrant methylation than CGI within divergent promoters. In addition, we identified three DNA sequence motifs that correlate with methylated CGI. We also showed that methylated CGI are associated with genes whose expression is low in normal cells. Thus, we here provide fundamental observations regarding the methylation of divergent promoters that are essential for the understanding of carcinogenesis and development of cancer prevention strategies.
Reviews (Japanese)
Namba, S. & Okada, Y. [CURRENT PIPELINES FOR WHOLE-GENOME SEQUENCING ANALYSES]. Arerugi 72, 1110–1112 (2023).
Namba, S. & Okada, Y. [UTILIZING GENOMIC INFORMATION FOR BIOMARKER IDENTIFICATION AND GENOMIC DRUG DISCOVERY]. Arerugi 69, 952–957 (2020).