Article Online

Articles Online (Volume 23, Issue 5)

Preface

Harnessing Large Cohorts and AI to Bridge Genomic Discovery and Clinical Practice

Bitao Zhong, Shaoqi Wang, Xiaoxi Jing, Aniruddh P Patel, Yajie Zhao, Minxian Wang

no abstract

Page qzaf104


Perspective

Toward Responsible and Sustainable Data Sharing in Large-scale Cohort-based Genomic Research

Jie Song, Wenwen Chen, Jin Huang, Huan Song

no abstract

Page qzaf100


Original Research

Whole-genome Sequence Analysis Revealed Novel Subjective Cognitive Decline-associated Genes in 10,763 Chinese

Mengying Wang, Liyang Sun, Xin Xu, Ruoqi Dai , Qilong Tan , Yun Zhu , Andi Xu, Weifang Zheng, Yuanxing Tu, Dan Zhou, Wenyuan Li, Xifeng Wu

Subjective cognitive decline (SCD) is widely regarded as a potential preclinical stage of Alzheimer’s disease (AD), yet its genetic basis remains poorly understood. To address this gap, we investigated genetic biomarkers associated with SCD using whole-genome sequencing (WGS) in 10,763 Chinese participants from the Healthy Zhejiang One Million People Cohort (HOPE Cohort). The discovery stage included 9284 samples, with 1479 samples used for validation. Using a two-stage design, we systematically investigated both common and rare variants associated with SCD. In rare variant analyses, we identified and replicated an association between the upstream region of SEPHS2 and SCD. SEPHS2 is involved in selenophosphate synthesis, and a Mendelian randomization analysis reveals that its expression levels in both blood and brain cerebellum are associated with AD. Additionally, we identified CLVS2, which encodes a protein primarily expressed in neuronal cells, as a potential regulator for SCD based on missense rare variants. Multi-omics evidence suggests that both SEPHS2 and CLVS2 may play roles in neurodegenerative diseases. For common variants, we validated 8 known loci related to cognitive decline, 3 of which originated from the only existing SCD genetic study conducted under a migraine background. Overall, our WGS-based study fills the gap in SCD research by providing vital genetic evidence from an East Asian population and offers insights into the pathogenic mechanisms of SCD.

Page qzaf063


Original Research

Regulatory Genomic Circuitry of Brain Age by Integrative Functional Genomic Analyses

Xingzhong Zhao, Anyi Yang, Jing Ding, Yucheng T Yang, Xing-Ming Zhao

Brain age gap (BAG) is a valuable biomarker for evaluating brain healthy status and detecting age-associated cognitive degeneration. However, the genetic architecture of BAG and the underlying mechanisms are poorly understood. Here, we estimated brain age from magnetic resonance imaging with improved accuracy using our proposed adversarial convolution network (ACN), and applied the ACN model to an elderly cohort from the UK Biobank. The genetic heritability of BAG was significantly enriched in regulatory regions and implicated in glial cells. We prioritized a set of BAG-associated genes, and further characterized their expression patterns across brain cell types and regions. Two BAG-associated genes, RUNX2 and KLF3, were found to be associated with epigenetic clock and diverse aging-related biological pathways. Finally, two BAG-associated hub transcription factor genes, KLF3 and SOX10, were identified as regulators of pleiotropic risk genes for diverse brain disorders. Altogether, we improve the estimation of BAG, and identify BAG-associated genes and regulatory networks implicated in brain disorders.
研究问题: 随着人类寿命延长,保持大脑健康成为延缓衰老的重要研究方向。大脑年龄差(brain age gap,BAG)是指利用大脑磁共振成像预测的大脑年龄与实际年龄之间的差值,它是衡量脑衰老程度的有力生物标志物。然而BAG的遗传结构和调控关系尚未被系统揭示,限制了其在疾病预防和干预中的应用。研究团队希望回答以下问题:脑龄差的遗传基础是什么?哪些基因和调控网络在大脑衰老中发挥关键作用?脑龄差与精神疾病及衰老相关疾病之间有何因果关系? 研究方法: 本研究构建域对抗卷积网络(adversarial convolution network,ACN),将性别与采集站点视为“域”进行对抗训练以提取与年龄相关且去偏的特征;模型在五个独立队列(西南大学成人生命周期数据集SALD、花旗集团生物医学影像中心CBIC、澳大利亚影像生物标志物与生活方式老龄化研究AIBL、图像信息提取IXI、开放获取影像研究系列OASIS)共2011名4–95岁健康个体上训练,并在英国生物银行3.57万人(45–85岁)上预测BAG,平均绝对误差(mean absolute error, MAE)≈2.70岁;随后整合基因型数据进行全基因组关联研究(genome-wide association study, GWAS)分析,并进行多层次的功能基因组整合分析,包括大脑衰老遗传力的组织与细胞类型富集分析,大脑衰老与多种脑疾病的遗传关联,大脑衰老相关基因的鉴定及表达动态特征分析,大脑衰老相关遗传变异/基因与表观遗传时钟的关联,以及大脑衰老相关基因的调控网络分析。 主要结果: 1. ACN模型在跨数据集上表现优异。与LASSO、弹性网络、GPR、SVR、3D ResNet等模型比较,ACN在多个数据集上的平均绝对误差最低。该模型学习到的特征可在支持向量机(support vector machine, SVM)中以约79%的准确率区分阿尔茨海默病患者和健康人群。 2. GWAS鉴定大量新遗传变异。在3.6万名个体的GWAS中,共发现3868个达到全基因组显著性的单核苷酸多态性(single nucleotide polymorphism, SNP),并筛选出10个独立的lead SNP,其中9个为首次报道;BAG的遗传力约为0.21。 3. BAG的遗传力主要富集在调控区和胶质细胞中。遗传力富集分析显示BAG的遗传贡献主要富集在超级增强子(富集倍数为2.61,假发现率FDR=2.97E−08)和H3K27ac标记的增强子区域;在细胞层面,在胶质细胞的富集度明显高于神经元。 4. BAG与多种脑病和代谢疾病的遗传风险相关。多基因风险评分分析表明BAG与阿尔茨海默病、帕金森病、2型糖尿病、重度抑郁症、精神分裂症等的风险显著正相关。Mendelian随机化分析进一步指出精神分裂症和双相情感障碍可能促进大脑加速衰老。 5. 识别关键基因及其调控网络。基于基因的分析和功能映射,共确认近200个BAG相关基因,其富集于WNT信号通路、突触传递、细菌脂蛋白代谢等功能;这些基因在成人大脑和胶质细胞中的表达显著高于幼年期。RUNX2、KLF3、SOX10等枢纽基因不仅与DNA甲基化年龄差相关,还调控多种神经疾病风险基因。

Page qzaf064


Original Research

Mitochondrial Genome Variants and Nuclear Mitochondrial DNA Segments in 7331 Individuals from NyuWa and 1KGP

Yuanxin Wang, Jiajia Wang, Yanyan Li, Peng Zhang, Zhonglong Wang, Shuai Liu, Yiwei Niu, Yirong Shi, Sijia Zhang, Tingrui Song, Tao Xu, Shunmin He

Dysfunctional mitochondria are implicated in various diseases, but comprehensive characterization of mitochondrial DNA (mtDNA) in the Chinese population remains limited. Here, we conducted a systematic analysis of mtDNA from 7331 samples, comprising 4129 Chinese samples from the NyuWa cohort and 3202 samples from the 1000 Genomes Project (1KGP). We identified 7216 high-quality mtDNA variants, which classified 7266 samples into 22 macro-haplogroups, and detected 1466 nuclear mitochondrial DNA segments (NUMTs). Among these, 88 mtDNA variants and 642 NUMTs were specific to NyuWa. Genome-wide association analyses revealed significant correlations between 12 mtDNA variants and 199 nuclear DNA (nDNA) variants. Our findings demonstrated that all individuals in both NyuWa and 1KGP harbored common NUMTs, while one-fifth possessed ultra-rare NUMTs that tended to insert into nuclear gene regions. Notably, rare NUMTs in the NyuWa cohort showed significant enrichment of nuclear breakpoints in long interspersed nuclear elements (LINEs) compared to 1KGP. Overall, this study provides the first comprehensive profile of NUMTs in the Chinese population and establishes the most extensive resource of Chinese mtDNA variants and NUMTs to date based on high-depth whole-genome sequencing, providing valuable reference resources for genetic research on mtDNA-related diseases.
研究问题: 以往的线粒体DNA(mitochondrial DNA, mtDNA)研究主要基于欧洲裔个体,针对东亚人群,尤其是中国人群的mtDNA变异和核内线粒体DNA片段(Nuclear Mitochondrial DNA Segments, NUMTs)深度测序研究仍相对缺乏。 研究方法: 全基因组测序(Whole Genome Sequencing, WGS)数据比对至GRCh38参考基因组(包含线粒体DNA,NC_012920.1),并使用BWA-mem生成BAM(Binary Alignment/Map)文件。使用线粒体变异检测流程(https://github.com/gatk-workflows/gatk4-mitochondria-pipeline)及其默认参数对每个样本进行线粒体DNA变异检测并通过NUMTs-detection识别了 7331 个 WGS 样本中的NUMTs。 主要结果: 研究基于大规模人群测序,共鉴定出 7216 个高质量 mtDNA 变异,可将样本划分为 22 个宏单倍群;同时识别出 1466 个独立 NUMTs,其中 88 个 mtDNA 变异及 642 个 NUMTs 为 NyuWa 特有。在进一步的全基因组关联分析中,研究者发现 12 个 mtDNA 变异与 199 个核基因组变异显著相关。本工作首次系统描绘了中国人群 NUMTs 的全景图谱,显示 NyuWa 与1000 Genomes Project(1KGP)个体普遍携带常见NUMTs,且约五分之一的个体携带极罕见 NUMTs,这些罕见片段更倾向于插入核基因区域。通过整合上述信息,研究构建了迄今最全面的 中国人群 mtDNA 变异与 NUMTs 资源,为解析线粒体相关疾病的遗传基础提供了关键参考。 数据链接或代码连接或其他: https://github.com/gatk-workflows/gatk4-mitochondria-pipeline https://github.com/WeiWei060512/NUMTs-detection.git http://bigdata.ibp.ac.cn/NMVR/

Page qzaf098


Original Research

Genome-wide Association Studies of over 30,000 Samples with Bone Mineral Density at Multiple Skeletal Sites and Its Clinical Relevance

Yu Qian, Jiangwei Xia, Pingyu Wang, Chao Xie, Hong-Li Lin, Gloria Hoi-Yee Li, Cheng-Da Yuan, Mo-Chang Qiu, Yi-Hu Fang, Chun-Fu Yu, Xiang-Chun Cai, Saber Khederzadeh, Pian-Pian Zhao, Meng-Yuan Yang, Jia-Dong Zhong, Xin Li, Peng-Lin Guan, Jia-Xuan Gu, Si-Rui Gai, Xiang-Jiao Yi, Jian-Guo Tao, Xiang Chen, Mao-Mao Miao, Guo-Bo Chen, Lin Xu, Shu-Yang Xie, Geng Tian, Hua Yue, Guangfei Li, Wenjin Xiao, David Karasik, Youjia Xu, Liu Yang, Ching-Lung Cheung, Fei Huang, Zhenlin Zhang, Hou-Feng Zheng

The ultimate goal of a genome-wide association study (GWAS) is to translate its discoveries into clinical practice. To explore the clinical use of GWAS findings in the bone field, we conducted a GWAS of dual-energy X-ray absorptiometry (DXA)-derived bone mineral density (BMD) traits at 11 skeletal sites, within over 30,000 European individuals from the UK Biobank. A total of 91 unique and independent loci were identified for 11 DXA-derived BMD traits and fractures, including 5 novel loci (harboring the genes ABCA1, CHSY1, CYP24A1, SWAP70, and PAX1) for 6 BMD traits. These loci exhibited evidence of association in both males and females, which could serve as independent replication. We demonstrated that each polygenic risk score (PRS) was independently associated with fracture risk. Although incorporating multiple PRSs (i.e., metaPRS) with clinical risk factors from the Fracture Risk Assessment Tool (FRAX) yielded the highest predictive performance, the improvement was modest in fracture prediction. Additionally, we uncovered genetic correlation and shared polygenicity between head BMD and intracranial aneurysm (IA). Finally, by integrating gene expression and GWAS datasets, we prioritized genes (e.g., ESR1 and SREBF1) encoding druggable human proteins along with their respective inhibitors/antagonists. In conclusion, this comprehensive investigation reveals a new genetic basis for BMD and its clinical relevance to fracture prediction. More importantly, it suggests that head BMD is genetically correlated with IA. The prioritization of genetically supported targets implies the potential repurposing of drugs [e.g., omega-3 polyunsaturated fatty acid (PUFA) supplements] for the prevention of osteoporosis.
研究问题: 骨矿物密度是诊断骨质疏松和预测骨折风险的关键指标。尽管全基因组关联研究(Genome wide association study, GWAS)已鉴定出数百个骨矿物密度相关位点,但其结果多源于不同研究的汇总数据,存在固有局限。本研究利用英国生物银行(UK Biobank)中超过3万人的个体基因型数据和11个骨骼部位的双能X射线吸收测量法(Dual-energy X-ray absorptiometry, DXA)骨密度数据开展个体水平的大规模GWAS分析,旨在实现四大目标:一、发现新的骨矿物密度遗传易感位点;二、系统评估多部位骨密度多基因风险评分对骨折的预测效能;三、揭示骨密度与神经退行性疾病、心血管疾病等复杂疾病的共享遗传基础;四、整合多组学数据,筛选具有遗传学证据的潜在药物靶点。 研究方法: 对UK Biobank中约3万名欧洲个体的11个部位骨密度和骨折数据进行GWAS分析。利用多基因风险评分评估骨折风险预测能力。采用连锁不平衡分数回归、MiXeR、共定位分析等方法探究骨密度与13种常见慢性病的遗传相关性。结合ChEMBL可药基因组数据库、表达数量性状基因座(expression quantitative trait locus, eQTL)数据和孟德尔随机化方法,系统性筛选并验证骨质疏松的潜在药物靶点。 研究结果: 本研究共鉴定出91个独立的与骨密度相关的遗传位点,包括5个新位点(涉及ABCA1、CHSY1、CYP24A1、SWAP70和PAX1基因),这些位点在男性和女性群体中均得到一致验证。进一步分析显示,基于多部位骨密度构建的多基因风险评分能够显著预测骨折风险,然而将其纳入临床风险因素模型后,对骨折预测准确性的提升非常有限,表明在风险分层中临床风险因素仍占据主导地位。值得关注的是,本研究首次揭示了头骨骨密度与颅内动脉瘤之间存在显著的负向遗传相关性,并鉴定出包括PLCE1在内的4个共享基因位点,从而提示骨骼系统与脑血管健康之间可能存在遗传联系。此外,通过遗传优先策略,筛选出ESR1、SREBF1、CCR1等一系列潜在治疗靶点,提示Omega-3脂肪酸(靶向SREBF1)等现有药物在骨质疏松预防方面具有“老药新用”的潜在价值。

Page qzaf097


Original Research

Integrative Genome-wide Association Meta-analysis of Aortic Aneurysm and Dissection Identifies Five Novel Genes

Yifan Du, Yunlong Guan, Zhonghe Shao, Minghui Jiang, Minghan Qu, Yifan Kong, Hongji Wu, Da Luo, Shu Peng, Si Li, Xi Cao, Jing Chen, Ping Ye, Jiahong Xia, Xingjie Hao

Aortic aneurysm and dissection (AAD) is a multifaceted condition characterized by significant genetic predisposition and a considerable contribution to cardiovascular-related mortality. Previous studies have suggested that AAD subtypes share similar genetic mechanisms; however, these studies investigated the subtypes separately. Here, we performed a large genome-wide association study (GWAS) meta-analysis for AAD by combining its subtypes, including 11,148 cases and 708,468 controls of European ancestry. We identified 24 susceptibility loci, including four novel loci at 1p21.2 (PALMD), 2p22.2 (CRIM1), 6q22.1 (FRK), and 12q14.3 (HMGA2), which were partially validated in both internal and external populations. Cell type-specific analysis highlighted the artery as the most relevant tissue where the susceptibility variants may exert their effects in a tissue-specific manner. By using four approaches, we prioritized 53 genes, reinforcing the importance of elastic fiber formation and transforming growth factor-beta (TGF-β) signaling in the formation of AAD, and suggested potential target drugs for the treatment. Additionally, various cardiovascular diseases were genetically correlated with AAD, and several cardiovascular risk factors [e.g., body mass index (BMI), lipid levels, and pulse pressure] showed causal associations with AAD, underscoring their shared genetic structures and mechanisms underlying the comorbidity. Moreover, five prioritized genes (PALMD, CRIM1, FRK, HMGA2, and NT5DC1) at the novel loci were supported as regulators of smooth muscle and endothelial cell functions through ex vivo and in vitro experiments. Together, these findings enhance our understanding of the genetic architecture of AAD and provide novel insights into future biological mechanism studies and therapeutic strategies.

Page qzaf039


Original Research

An Integrative Polygenic and Epigenetic Risk Score for Overweight-related Hypertension in Chinese Population

Yaning Zhang, Qiwen Zheng, Qili Qian, Na Yuan, Tianzi Liu, Xingjian Gao, Xiu Fan, Youkun Bi, Guangju Ji, Peilin Jia, Sijia Wang, Fan Liu, Changqing Zeng

Overweight-related hypertension (OrH), defined by the coexistence of excess body weight and hypertension (HTN), is an increasing health concern elevating cardiovascular disease risks. In this study, we evaluated the prediction performance of polygenic risk scores (PRSs) and methylation risk scores (MRSs) for OrH in 7605 Chinese participants from two cohorts: the Chinese Academy of Sciences (CAS) and the National Survey of Physical Traits (NSPT). In the CAS cohort, which predominantly consists of academics, males showed significantly higher prevalence of obesity, HTN, and OrH, along with worse metabolic syndrome indicators, compared to females. This disparity was less pronounced in the NSPT cohort and in broader Chinese epidemiological studies. Among ten PRS methods, PRS-CSx was the most effective, enhancing prediction accuracy for obesity [area under the curve (AUC) = 0.75], HTN (AUC = 0.74), and OrH (AUC = 0.75), compared to baseline models using only age and sex (AUC = 0.55–0.71). Similarly, least absolute shrinkage and selection operator (LASSO)-based MRS models improved prediction accuracy for obesity (AUC = 0.70), HTN (AUC = 0.73), and OrH (AUC = 0.78). Combining PRS and MRS further boosted prediction accuracy, achieving AUC values of 0.77, 0.76, and 0.80 for obesity, HTN, and OrH, respectively. These models stratified individuals into high (> 0.6) or low (< 0.1) risk categories, covering 59.95% for obesity, 31.75% for HTN, and 43.89% for OrH. Our findings highlight a higher OrH risk among male academics, emphasize the influence of metabolic and lifestyle factors on MRS predictions, and highlight the value of multi-omics approaches in enhancing risk stratification.

Page qzaf048


Original Research

Association of Multiple-trait Polygenic Risk Score with Obesity and Cardiometabolic Diseases in Korean Population

Jinyeon Jo, Nayoung Ha, Yunmi Ji, Ahra Do, Je Hyun Seo, Bumjo Oh, Sungkyoung Choi, Eun Kyung Choe, Woojoo Lee, Jang Won Son, Sungho Won

We conducted a comprehensive genetic investigation of obesity in a cohort of 93,673 Korean individuals, categorized by body mass index and waist circumference using Korean-specific and international criteria. To explore the genetic architecture of obesity and its related comorbidities, we performed genome-wide association studies and constructed polygenic risk scores (PRSs) using both conventional single-trait and advanced multiple-trait models, including the PRSsum approach. Our analyses identified genome-wide significant loci and demonstrated their higher heritability for general obesity than for abdominal obesity, and for moderate obesity than for severe obesity. East Asian populations showed stronger genetic correlations between abdominal obesity and obesity-related diseases. Both single-trait and multiple-trait PRSs stratified individuals by risk, with low-PRS individuals exhibiting reduced risk for obesity, hypertension, and type 2 diabetes, while high-PRS individuals displayed elevated risk, particularly under the multiple-trait model. Interaction and mediation analyses revealed distinct genetic pathways through which obesity contributes to disease development. Collectively, our findings revealed key loci and shared genetic mechanisms linking obesity and its comorbidities in the Korean population. These insights highlight the value of multiple-trait PRS models and underscore the importance of ancestry-specific genetic research for addressing the obesity epidemic.

Page qzaf102


Original Research

Benefits of Better Cardiovascular Health for Calcific Aortic Valve Stenosis Stratified by Polygenic Risk Score

Yuexin Zhu (朱岳鑫) , Qiuli Chen (陈秋丽) , Bokang Qiao (乔博康) , Lixin Jia (贾立昕) , Haichu Wen (温海初) , Wei Pan (潘威) , Yifan Wang (王一帆) , Shuangli Mi (米双利) , Minxian Wang (汪敏先) , Jie Du (杜杰)

With global aging, the prevalence of calcific aortic valve stenosis (CAVS) has significantly increased, and even mild CAVS elevates mortality risk, highlighting an urgent need for preventive strategies. In this study, we analyzed 153,312 participants from the UK Biobank. We found that a higher Life’s Essential 8 (LE8) score was independently associated with a decreased CAVS risk [hazard ratio (HR)/standard deviation (SD): 0.72, 95% confidence interval (CI): 0.68–0.76, P < 2E−16], whereas a higher genetic risk score was independently associated with an increased CAVS risk (HR/SD: 1.63, 95% CI: 1.54–1.71, P < 2E−16). Restricted cubic spline analysis revealed approximately linear inverse associations between LE8 score and CAVS risk across all genetic risk groups. Although no multiplicative interaction was found between cardiovascular health (CVH) and genetic risk for CAVS, a significant additive interaction was identified (relative excess risk due to interaction: 3.77, 95% CI: 1.30–8.50). Among participants with high genetic risk, those with ideal CVH had a lower 10-year cumulative CAVS incidence rate than those with poor CVH (0.33% vs. 1.80%, P < 0.001). Besides, a significant multiplicative interaction was observed between LE8 score and age (P = 0.007). Similar trends were also observed for early-onset and late-onset CAVS. In conclusion, regardless of genetic risk groups, a higher LE8 score was associated with a lower CAVS risk in an approximately linear pattern. In particular, high genetic risk and poor CVH had a synergistic effect on CAVS risk, meaning that participants with high genetic risk and poor CVH would amplify their CAVS risk. Therefore, early and sustained optimization of CVH, particularly among those with high genetic risk, is essential to mitigate the risk of CAVS and its subtypes.

Page qzaf099


Original Research

Cross-ethnic Molecular Signatures Underpin the Adverse Impact of Statin Use on Type 2 Diabetes

Fengzhe Xu, Min Yang, Wei Hu, Shuai Yuan, Xue Cai, Wanglong Gou, Zelei Miao, Bang-yan Li, Liang Yue, Zhangzhi Xue, Menglei Shuai, Luqi Shen, Yuanqing Fu, Tiannan Guo, Yu-ming Chen, Ju-Sheng Zheng

The use of statins as the primary therapy for reducing low-density lipoprotein cholesterol has raised concerns regarding their potential side effects in increasing the risk of type 2 diabetes (T2D). However, the underlying mechanism remains largely unknown. In this study, we utilized multi-omics molecular signatures to unravel the etiology of statin-induced T2D. Through systematic screening of 102 gut microbial features, 40 blood metabolites, and 131 circulating proteins in East Asians and Europeans, we identified a set of blood metabolites and proteins potentially influenced by genetically proxied statin use. Notably, Mendelian randomization analyses provided evidence that elevated circulating levels of gastric inhibitory polypeptide (GIP) were associated with an increased risk of T2D. This association between genetically proxied statin use and GIP was consistently observed across East Asian and European populations, highlighting the pivotal role of GIP in modulating the risks of statin-induced T2D. Furthermore, this study established an extensive atlas of multi-omics molecular signatures associated with statin-induced T2D, offering valuable insights for prioritizing intervention targets.
研究问题 他汀类药物是临床上常见的降胆固醇药物,能够降低血液中低密度脂蛋白胆固醇,但临床观察发现该药物的使用可能会增加2型糖尿病的风险,而这种关联背后的潜在机制仍然未知。本研究旨在通过结合遗传数据和多组学分子图谱,在东亚和欧洲人群中解析他汀类药物诱导2型糖尿病风险升高背后的潜在机制。 研究方法 本研究结合了东亚人和欧洲人的遗传和多组学数据,并使用孟德尔随机化方法,以编码HMG-CoA还原酶(HMGCR)基因变异作为他汀类药物使用的遗传学工具,系统筛选了受他汀类药物影响的分子图谱,并评估了相关分子特征与2型糖尿病风险的潜在因果关系。此外,研究还利用机器学习模型验证了以上探索发现的分子特征对2型糖尿病的预测能力。 主要结果: 1. 整合遗传和蛋白质组研究发现,血液抑胃肽(Gastric Inhibitory Polypeptide,GIP)水平升高能够显著增加2型糖尿病发病风险。这一关联在东亚和欧洲人群中具有一致性。 2.肾上腺酸(adrenic acid)是遗传预测的他汀类药物使用与2型糖尿病风险之间的潜在中介分子。 3.他汀类药物使用有关的多组学分子特征与传统风险因素整合,能够显著提升对2型糖尿病的预测性能。 数据链接或代码连接或其他: https://omics.lab.westlake.edu.cn/resource/statint2d.html

Page qzaf101


Original Research

Distinct Co-methylation Patterns in African and European Populations and Their Genetic Associations

Zheng Dong, Nicole Gladish, Maggie P Y Fu, Samantha L Schaffner, Keegan Korthauer, Michael S Kobor

Human populations have substantial genetic diversity, but the extent of epigenetic diversity remains unclear, as population-specific DNA methylation (DNAm) has only been studied for ∼ 3.0% of CpGs. In this study, we quantified DNAm using whole-genome bisulfite sequencing (WGBS) and analyzed it alongside whole-genome genotype data to provide a more comprehensive view of population-specific DNAm. Using a co-methylated region (CMR) approach, 36,657 CMRs were identified in WGBS data from 62 lymphoblastoid B-cell line (LCL) samples, with subsequent validation in a combined array dataset of 326 LCL samples. Between individuals of European and African ancestry, 101 CMRs exhibited population-specific DNAm patterns (Pop-CMRs), including 91 Pop-CMRs not reported in previous investigations. These regions spanned genes (e.g., CCDC42, GYPE, MAP3K20, and OBI1) related to diseases (e.g., malaria infection and diabetes) with differing prevalence and incidence between populations. Over half of the Pop-CMRs were associated with genetic variants, displaying population-specific allele frequencies and primarily mapped to genes involved in metabolic and infectious processes. Additionally, subsets of Pop-CMRs were applicable in East Asian populations and peripheral blood-based tissues. This study highlights genome-wide DNAm differences between populations and examines their associations with genetic varation and biological relevance, advancing our understanding of epigenetic contributions to population specificity.

Page qzaf096


Original Research

Boosting the Power of Rare Variant Association Studies by Imputation Using Large-scale Sequencing Population

Jinglan Dai, Yixin Zhang, Yuan Gao, Hongru Li, Sha Du, Hao Hong, Dongfang You, Zaiming Li, Ruyang Zhang, Yang Zhao, Zhonghua Liu, David C Christiani, Feng Chen, Sipeng Shen

With the emergence of population-scale whole-genome sequencing (WGS), rare variants can be captured precisely. Studying rare variants explains part of the heritability of complex traits that is overlooked by conventional genome-wide association studies (GWASs). However, the extent to which imputed data can approximate or improve upon the power of WGS data in rare variant association studies remains unclear. Using the UK Biobank WGS data (n = 150,119) as the ground truth, we first evaluated the consistency of rare variants in the single-nucleotide polymorphism (SNP) array data imputed using TOPMed or HRC+UK10K reference panel. Imputation quality (average R2) of the TOPMed-imputed data reached 0.6 even for extremely rare variants with minor allele count 5. TOPMed-imputed data were closer to WGS data across three ethnic groups, with average Cramer’s V 0.75. Furthermore, association tests were performed on 45 traits. At the same sample size (n = 150,119), neither imputed dataset outperformed WGS data, but the results of the TOPMed-imputed data were more consistent with those of WGS data. When the sample size was increased to 488,377, the number of significant rare variants identified from the TOPMed-imputed data increased by 27.71% for quantitative traits and by approximately 10-fold for binary traits. Finally, we meta-analyzed the association results of SNP array and WGS for lung cancer and epithelial ovarian cancer, respectively. Compared to WGS-based results, more significant variants and genes were identified. Our findings highlight that incorporating rare variants imputed using large-scale sequencing populations can boost the power of rare variant association studies when WGS has limited sample sizes.
研究问题: 全基因组测序(Whole-genome sequencing,WGS)技术的出现使得精准捕获罕见变异成为可能,这些变异在解释人类复杂性状和疾病遗传性方面具有重要意义。然而,由于单核苷酸多态性(Single-nucleotide polymorphisms,SNP)芯片技术的局限性,罕见变异在传统全基因组关联研究(Genome-wide association study,GWAS)中往往被忽视。基因型填补技术可通过高质量外部参考面板(如TOPMed和HRC+UK10K)对缺失的基因型进行填补,提供了一种弥补上述不足的途径。尽管已有研究表明应用填补数据开展罕见变异关联性分析能够发现未报道的新信号,但其在不同样本量和表型类型下是否能够接近或超越WGS数据的效能,仍需进一步系统性评估。 研究方法: 本研究将英国生物样本库(UK Biobank,UKB)中150,119名个体的WGS数据作为金标准,评估了分别基于TOPMed和HRC+UK10K参考面板进行基因型填补后的罕见变异覆盖率和准确性。首先,对白人、亚洲人、黑人三个种族分别进行WGS数据与填补数据之间的基因型一致性评估,以观察填补数据在不同族群中的表现。随后,利用这两种填补数据与WGS数据,在不同样本量(n = 150,119 和 n = 488,377)下,对WGS数据和填补数据中30个生化指标及15种复杂疾病进行关联分析,比较两者在罕见变异关联中的效能差异。最后,对肺癌和上皮性卵巢癌基于不同数据集的关联性分析结果进行了meta分析,以验证结合WGS和填补数据确能识别更多罕见变异关联信号与相关基因。 主要结果: 1. 填补数据与WGS数据的罕见变异覆盖情况显示,样本量150,119的TOPMed填补数据覆盖了WGS数据中22.2%的单核苷酸变异(Single-nucleotide variants,SNVs),远高于HRC+UK10K填补数据的10.0%。TOPMed填补数据共包含4000万个次要等位基因计数(Minor allele count,MAC)为1和2的超罕见变异,该数量是HRC+UK10K填补数据的4倍,不过仍少于WGS数据中同类的3.32亿个变异。 2. 对不同种族(白人、亚洲人和黑人)进行的基因型一致性分析显示,样本量150,119的TOPMed填补数据与WGS数据的相似性显著高于HRC+UK10K填补数据,TOPMed数据的平均Cramer's V值超过0.75,表明在所有MAC区间内,TOPMed填补数据与WGS数据在罕见变异上的一致性较强。 3. 针对30个生化指标和15种复杂疾病,分别使用总样本量150,119的WGS数据、TOPMed填补数据与HRC+UK10K填补数据开展罕见变异关联性分析。结果显示使用TOPMed填补数据发现的显著关联罕见变异数为WGS数据发现数的41.88%,而HRC+UK10K则为32.69%。 4. 当参与关联性分析的总样本量增加至488,377时,30个生化指标的分析结果显示TOPMed填补数据发现的显著关联罕见变异数相比总样本量为150,119的WGS数据发现数增加了27.71%,而HRC+UK10K填补数据发现数相比之下仅增加4.7%。对于15种复杂疾病,大样本填补数据发现的显著罕见变异数较WGS数据发现数能提高近10倍。 5. 对肺癌和上皮性卵巢癌分别基于WGS和SNP芯片数据开展关联分析,并对两组结果进行meta分析。单变异检验中,肺癌meta结果与上皮性卵巢癌meta结果分别发现12个和22个显著关联罕见变异,而单独使用WGS数据进行关联分析则未能保留任何显著信号。即便使用更严格的阈值(如单变异P < 10⁻⁷、基因水平P < 10⁻⁵),WGS+SNP芯片数据的meta策略仍能发现多数显著关联变异。 分析代码链接: https://ngdc.cncb.ac.cn/biocode/tool/BT007792

Page qzaf084


Method

SRPS: Survival Reinforced Transfer Learning for Multicentric Proteomic Subtyping and Biomarker Discovery

Linhai Xie, Pei Jiang, Cheng Chang

Omics-based molecular subtyping in large-scale and multicentric cohort studies is a prerequisite for proteomics-driven precision medicine (PDPM). However, maintaining subtypes with robust molecular features and significant prognostic associations across different cohorts remains challenging due to biological heterogeneity and technical inconsistency. Herein, we propose a subtyping algorithm, named Survival Reinforced Patient Stratification (SRPS), to adapt known subtypes from a discovery cohort to another by simultaneously preserving the distinct prognosis and molecular characteristics of each subtype. SRPS was benchmarked on simulated and real-world datasets, demonstrating a 12% increase in classification accuracy and best prognostic discrimination. Moreover, based on the calculated subtype significance score, an “unpopular” protein, peptidyl-prolyl cis-trans isomerase C (PPIC), was identified as the top 1 remarkable protein for subtyping hepatocellular carcinoma (HCC) patients with the worst prognosis. Eventually, PPIC was experimentally validated as a pro-cancer protein in HCC, confirming our work as a demonstration of interpretable machine learning-guided biological discovery in PDPM research. SRPS is publicly available at https://github.com/PHOENIXcenter/SRPS and https://ngdc.cncb.ac.cn/biocode/tool/BT007770.

Page qzaf052


Database

GDBIG: A Pioneering Birth Cohort Genomic Platform Facilitating Intergenerational Genetic Research

Shujia Huang, Chengrui Wang, Mingxi Huang, Jinhua Lu, Jian-Rong He, Shanshan Lin, Siyang Liu, Huimin Xia, Xiu Qiu

High-quality genome databases derived from large-scale, family-based birth cohorts are vital resources for investigating the genetic determinants of early-life traits and the impact of early-life environments on the health of both parents and offspring. Here, we established a genomic platform for the Born in Guangzhou Cohort Study (BIGCS), the Genome Database of BIGCS (GDBIG), which represents the first birth cohort-based genomic database in China and is designed to facilitate intergenerational genetic research. Based on the phase I results of the BIGCS, GDBIG includes low-coverage (∼ 6.63×) whole-genome sequencing (WGS) data and extensive pregnancy phenotypes from 4053 Chinese participants. These participants are from 30 of China’s 34 provincial-level administrative divisions, encompassing Han and 12 minority ethnic groups. Currently, GDBIG provides a range of services, including allele frequency queries for 56.23 million variants across two generations, a genotype imputation server featuring a high-quality family-based reference panel, and a genome-wide association study (GWAS) meta-analysis interface for various maternal and infant phenotypes. The GDBIG database addresses the lack of Asian birth cohort-based genomic resources and provides a valuable platform for conducting genetic analysis, accessible online or via application programming interfaces at http://gdbig.bigcs.com.cn/.

Page qzaf045


Database

HiLand Resource: A Comprehensive Database of Highland Human Populations

Weijie Zhang, Xiaoning Chen, Yanling Sun, Yibo Wang, Bixia Tang, Yu Zhang, Kai Liu, Wenming Zhao, Bing Su, Yaoxi He

Over 80 million people worldwide live at high altitudes (> 2500 m), where numerous studies have documented the remarkable biological adaptations of highland populations to these extreme environments. However, current resources for accessing and analyzing highlander-specific data remain limited. To address this gap, we present the HiLand Resource (HLR), a comprehensive database that integrates phenomic, genomic, and genetic association data from 23,336 highlanders across three major high-altitude regions: the Qinghai-Tibet Plateau, the Andean Plateau, and the Ethiopian Plateau. HLR offers six key functions: (1) visualization of phenotypic patterns among highlanders from the Qinghai-Tibet Plateau across different altitudes, as well as comparison between highlanders and lowlanders, and between sexes; (2) an interactive interface to explore genomic diversity, population structure, ancestral composition, and signatures of natural selection of high-altitude populations; (3) access to a comprehensive catalog of genome-wide variants and genes identified in highlanders; (4) a genome browser built on a high-quality Tibetan genome assembly; (5) a curated collection of genotype–phenotype associations derived from genome-wide association studies (GWASs) in highland populations; and (6) an online, user-friendly tool for genotype imputation using a highland-specific reference panel. Collectively, HLR provides a novel and in-depth resource for understanding the biological features of high-altitude human populations. It holds significant potential for advancing research on human adaptation to hypoxic environments and improving medical studies focused on highland communities. The HLR database is freely available at https://ngdc.cncb.ac.cn/hiland/.
研究问题: 高海拔地区低氧、低压、强紫外线等极端环境对人类生存构成严峻挑战。目前,全球超过8000万人长期居住在海拔2500米以上的高原区域。高原世居人群在长期自然选择过程中,已形成一系列生物学适应机制。近年来,多项大规模基因组研究陆续揭示了这些人群特有的遗传适应特征。然而,现有数据仍分散于不同研究中,缺乏一个集成多组学数据、支持在线分析及可视化查询的综合人群数据资源平台。 研究方法: 本研究系统收集并整合了来自青藏高原、安第斯高原和埃塞俄比亚高原三大高海拔区域的共29,977个体的表型、基因组及遗传关联数据。通过对数据开展系统质控、标准化处理与统一注释,进行了表型—海拔关联分析及群体适应性表型识别;综合运用多种群体遗传学方法与自然选择检测手段,解析高原人群的基因组多样性、群体结构、祖先成分及自然选择信号;整合高原人群全基因组关联分析(GWAS)结果,构建基于高原人群参考基因组的基因组浏览器,并依托藏族人群基因组参考面板(1KTGP)提供高精度基因型填充的分析工具。 主要结果: 我们成功构建了国际首个面向全球高原人群的综合数据库——HiLand Resource(HLR)。该数据库整合了来自青藏高原、安第斯高原和埃塞俄比亚高原的共29,977个体的表型、基因组与遗传关联数据,提供高效的数据浏览与检索功能。HLR作为全面、开放的数据平台,致力于系统解析高原适应的生物学机制,并为进化生物学和高海拔医学研究提供关键数据资源。 数据库链接: https://ngdc.cncb.ac.cn/hiland/

Page qzaf083