Article Online

Articles Online (Volume 24, Issue 2)

Review Article

Single-cell Omics Assessment of Mitochondrial Function: Current Status and Future Perspectives

Xu Zhang, Peng An, Zhengyang Zhang, Yanling Hao, Xiaoxian Guo, Yinhua Zhu, Zhenglong Gu, Yongting Luo, Junjie Luo

In recent years, significant advancements in single-cell omics technologies have offered new insights into the study of mitochondria. These technologies are particularly suitable for investigating mitochondria due to their capacity to address both intracellular and intercellular heterogeneity. In this review, we categorize the mitochondrial dysfunction and variability identified in both pathological and physiological contexts through single-cell omics assessments. We examine the cutting-edge single-cell omics technologies and trace the evolution of studies on mitochondria, highlighting the transition from low-throughput to high-throughput capabilities and from single data types to the integration of multiple genomic and phenomic profiles. Furthermore, we emphasize the applications of single-cell mitochondrial assessment methods in exploring disease mechanisms, screening, and prevention, as well as their potential impacts on lineage tracing, drug discovery, and genetic counseling. Insights gained from single-cell technologies may lead to the development of novel therapeutic strategies, offering promising avenues for addressing diseases associated with mitochondrial dysfunction. Lastly, we identify the limitations of current methodologies and propose areas of focus for future research.

Page qzaf081


Review Article

Approaches to Studying Viral Pangenome Variation Graphs

Tim Downing

Pangenome variation graphs (PVGs) allow for the representation of genetic diversity in a more nuanced way than traditional reference-based approaches. Here, I focus on how PVGs are a powerful tool for studying genetic variation in viruses, offering insights into the complexities of viral quasispecies, mutation rates, and population dynamics. PVGs originated in human genomics and hold great promise for viral genomics. Previous work has been constrained by small sample sizes and gene-centric methods, whereas PVGs enable a more comprehensive approach to studying viral diversity. Large viral genome collections should be used to make PVGs, which offer significant advantages. Here, I outline accessible tools to achieve their construction. These tools span PVG construction, PVG file formats, PVG manipulation and analysis, PVG visualisation, PVG openness measurement, and read mapping to PVGs. Additionally, the development of PVG-specific formats for mutation representation and personalised PVGs that reflect specific research questions will further enhance PVG applications. Challenges remain, particularly in managing nested variants, optimising error detection, optimising k-mer/minimizer-based approaches for AT-rich genomes, incorporating long-read sequencing data, and developing scalable visualisation approaches. Nevertheless, PVGs offer a new opportunity for viral population genomics, and a testing ground for tool development prior to application to larger eukaryotic genomes. These advances will enable more accurate and comprehensive detection of viral mutations, contributing to a deeper understanding of viral evolution and genotype–phenotype associations.

Page qzag003


Original Research

Benchmarking Generative AI Protein Models Reveals Differences Between Structural and Sequence-based Approaches

Alexander J Barnett, Rajendra KC, Pratikshya Pandey, Pamodha Somasiri, Kirsten A Fairfax, Sandy Hung, Alex W Hewitt

Recent advances in artificial intelligence have led to the development of generative models for de novo protein design. In this study, we compared 13 state-of-the-art generative protein models, assessing their ability to produce feasible, diverse, and novel protein monomers. Structural diffusion models generally create designs with higher confidence in predicted structures and more biologically plausible energy distributions, but exhibit limited diversity and strong sequence biases. Conversely, protein language models generate more diverse and novel designs but with lower structural confidence. We also evaluated the ability of these models to generate unique proteins, conditionally based on the tobacco etch virus (TEV) protease. Generative models are successful in producing functional enzymes, albeit with diminished activity compared to the wild-type TEV. Our systematic benchmarking provides a foundation for evaluating and selecting generative protein models, while highlighting the complementary strengths of different generative paradigms. This framework will facilitate informed application of these tools for biomedical engineering and design.

Page qzag014


Method

scATAnno: Automated Cell Type Annotation for Single-cell ATAC-seq Data

Yijia Jiang, Zhirui Hu, Feng Lu, Allen W Lynch, Junchen Jiang, Alexander Zhu, Ziqi Zeng, Yi Zhang, Gongwei Wu, Yingtian Xie, Rong Li, Ningxuan Zhou, Cliff A Meyer, Paloma Cejas, Myles Brown, Henry W Long, Xintao Qiu

Recent advances in single-cell epigenomic techniques have increased the demand for single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) analysis. One key analytical task is to determine cell type identity based on epigenetic data. Here, we introduce scATAnno, a Python package designed to automatically annotate scATAC-seq data using large-scale scATAC-seq reference atlases. This workflow generates reference atlases from publicly available datasets, enabling accurate cell type annotation by integrating query data with reference atlases without the use of single-cell RNA sequencing (scRNA-seq) data. To enhance annotation accuracy, we incorporated k-nearest neighbors (KNN)-based and weighted distance-based uncertainty scores to effectively detect cell populations within the query data that are distinct from all cell types in the reference data. We compared and benchmarked scATAnno against five other published cell annotation approaches, demonstrating its superior performance across multiple datasets and metrics. We further showcased the utility of scATAnno across multiple datasets, including peripheral blood mononuclear cells (PBMCs), triple-negative breast cancer (TNBC), and basal cell carcinoma (BCC), and demonstrated that scATAnno accurately annotates cell types across diverse biological conditions. Overall, scATAnno is a useful tool for scATAC-seq reference atlas construction and cell type annotation and can facilitate the interpretation of new scATAC-seq datasets in complex biological systems. scATAnno is publicly available at https://scatanno-main.readthedocs.io/.

Page qzaf108


Database

PASSpedia: A Polyadenylation Site Database Across Different Species at Single‐cell Resolution

Pei-Hong Zhang, Hua Feng, Xu-Kai Ma, Fang Nan, Li Yang

Polyadenylation site (PAS) selection plays important roles in gene expression regulation and function. RNA sequencing (RNA-seq) data derived from 3′ tag sequencing contain intrinsic information about PAS usage and have been analyzed for alternative polyadenylation (APA) isoform expression in both bulk and single‐cell samples. Here, we upgraded our previously developed deep learning-based PAS analysis pipeline SCAPTURE v2 to profile PASs from 1330 published 3′ tag-based single-cell RNA-seq (scRNA-seq) datasets across seven species, resulting in a comprehensive PAS landscape across species. Validation with long-read sequencing data from matched human tissues showed high accuracy of single-cell PAS profiling by SCAPTURE, including previously unannotated ones. Further comparisons revealed distinct PAS usage preferences in different species, such as human versus mouse, independent of conservation of gene expression. Finally, we present PASSpedia, a comprehensive database for PAS analysis and comparison across seven species at single‐cell resolution, which is freely accessible online at https://bits.fudan.edu.cn/PASSpedia/.

Page qzaf089


Database

circASbase: A Comprehensive Database of Alternative Splicing Events in circRNAs

Lingxiao Zou, Jian Zhao, Haojie Li, Chen Xu, Yulan Wang, Xuejiang Guo, Xiaofeng Song

Although extensive evidence has underscored the critical role of alternative splicing (AS) in generating mature circular RNA (circRNA) isoforms and augmenting their functional diversity, a significant gap remains in the availability of specialized databases housing circRNA AS events. To bridge this gap, we develop circASbase, a pioneering and comprehensive database that catalogs 452,129 AS events in 884,047 full-length circRNAs from 581 samples across 13 species, and provides rich annotations to facilitate understanding the splicing regulation of circRNA. Our findings reveal substantial differences between circRNAs and linear transcripts regarding the distribution and occurrence of AS events, highlighting the unique regulatory landscape of circRNAs. These special splicing events result in functional differences of circRNAs by affecting internal ribosome entry sites, N6-methyladenosine sites, open reading frames, protein features, microRNA targets, and more. In summary, circASbase not only meets the urgent need of the research community for data repositories, but also represents a significant advancement in our understanding of circRNA biology. With its user-friendly interfaces and web-based visualization tools, circASbase is poised to become an indispensable resource for researchers exploring the regulatory mechanisms and functional roles of AS events in circRNAs. This database will continuously drive new insights and discoveries in the field, setting the stage for further advancements in circRNA research. circASbase is freely available at http://reprod.njmu.edu.cn/cgi-bin/circASbase/.

Page qzaf121


Method

Ψ-Atlas: An Integrated Atlas for Pseudouridine Epitranscriptome

Xiaochen Wang, Jinjing Luo, Xiaoqiang Lang, Yongqing Ling, Yiming Zhou, Guoxian Liu, Xiangye Chen, Yibo Chen, Yingshun Zhou, Yi Cao, Zhonghui Zhang, Changjun Ding, Demeng Chen, Qi Liu

Pseudouridine (Ψ) is a C5 glycosidic isomer of uridine, formed by breaking the N1 glycosyl bond and undergoing a 180° base rotation. This modification is one of the most widespread post-transcriptional alterations in RNA and is universally distributed among diverse RNA species. The pervasiveness of this modification enhances RNA structural integrity, confers unique structural and functional attributes upon the RNA molecules it adorns, and facilitates additional hydrogen bonding. However, a convenient, integrated, and intuitive visualization database that includes all currently reported species and RNA types is lacking. Here, we present Ψ-Atlas, an extensive database meticulously curated for the comprehensive collection and annotation of RNA pseudouridine. This database encompasses 554,895 Ψ modification sites across various RNA categories, including mRNA, non-coding RNA (ncRNA), tRNA, and rRNA, in 77 distinct species reported in the current literature. The sequencing methodologies employed comprise next-generation sequencing techniques such as Ψ-Seq, Pseudo-Seq, CeU-Seq, PSI-Seq, RBS-Seq, HydraPsiSeq, BID-Seq, and PRAISE-Seq, as well as third-generation sequencing methods like direct RNA sequencing. Ψ-Atlas is the most comprehensive and integrated resource for RNA Ψ modifications to date. Ψ-Atlas offers an intuitive interface for information display and a myriad of analytical tools, including PsiVar and PsiFinder. Overall, this platform serves as a robust search and visualization tool for the study of pseudouridylation in epitranscriptomics. Ψ-Atlas is available at https://rnainformatics.org.cn/PsiAtlas.
研究问题:目前虽已发现大量假尿苷修饰位点,但缺乏一个整合多物种、多RNA类型、多检测技术的专用数据库,阻碍了该领域的系统研究。因此,亟需构建一个集中、友好、功能齐全的假尿苷研究平台。本研究构建了一个目前最全面的RNA假尿苷修饰数据库,整合了78个物种、55万多个修饰位点,并提供可视化界面与分析工具,以支持表观转录组学研究。 研究方法:本研究从132项已发表研究中收集假尿苷数据,涵盖化学测序、二代测序和纳米孔直接RNA测序等技术。利用MySQL构建数据库,采用HTML/PHP/JavaScript开发网页界面,并集成JBrowse基因组浏览器、深度学习预测工具PsiFinder和变异影响分析工具PsiVar。 主要结果: 1. Ψ-Atlas共整合了来自78个物种的554,895个假尿苷修饰位点。 2. 数据库包含了16种人类疾病相关的假尿苷修饰信息。 3. 开发了PsiFinder工具,可基于深度学习模型高精度预测假尿苷位点。 4. 开发了PsiVar工具,能有效评估基因变异对假尿苷修饰的潜在影响。 数据链接或代码连接或其他: 数据库访问地址:https://rnainformatics.org.cn/PsiAtlas 代码地址:代码存放于国家基因组科学数据中心BioCode,编号为BT007923, BT007924, BT007926。

Page qzag004


Database

MoRNiNG: A Database of RNA Modification Sites Associated with RNA Secondary Structure Dynamics

Yicen Zhou, Shanxin Lyu, Shiau Wei Liew, Xi Mou, Ian Hoffecker, Jian Yan, Yu Li, Chun Kit Kwok, Jilin Zhang

RNA structures are essential building blocks of functional RNA molecules. Profiling secondary structures in vivo and in real time remains challenging because RNAs exhibit dynamic structures and complex conformations. In addition to the canonical stem-loop secondary structure, the non-canonical RNA G-quadruplex (rG4) structure has attracted interest for its potential as a drug target. Early studies have demonstrated that RNAs can form distinct secondary structures. However, how distinct RNA structures formed from the same RNA sequence function within the transcriptome is poorly understood, and the factors that drive and regulate structural transitions remain to be investigated. Inspired by the ability of a HOXB9 segment to form multiple structures, we found that many RNA segments across the transcriptome exhibit multi-faceted structure-forming potential. In the case of HOXB9, we demonstrated that N6-methyladenosine (m6A) modification influences RNA structure and binding to RNA-binding proteins (RBPs). Therefore, we collected RNA modification sites naturally occurring within the putative G-quadruplex-forming sequences (PQSs) of transcripts and developed MoRNiNG, a database for RNA modifications in natural rG4 structures. MoRNiNG is organized into reliability tiers determined by the resolution of RNA modification sites and is designed to accommodate various large datasets. We experimentally validated the influence of m6A, 5-methylcytosine (m5C), and adenosine-to-inosine (A-to-I) editing on rG4-forming sequences, providing evidence to support the modification switch concept. The diversity and transition of secondary structures from the same RNA segment offer valuable insights into the regulation of RNA structural dynamics. MoRNiNG is freely accessible at https://www.cityu.edu.hk/bms/morning.
研究问题: RNA分子中普遍存在的G-四链体(RNA G-quadruplex, rG4)与茎环(stem-loop)等二级结构,其结构动态对RNA行使生物学功能至关重要。虽然多种RNA修饰存在于这些结构区域内,但这些修饰作为潜在的“开关”,驱动RNA结构变化并调控RNA-蛋白质互作的机制尚不明确。我们构建了rG4座位内RNA修饰位点的数据库——MoRNiNG,用于填补RNA表观遗传学与RNA结构动力学研究之间的知识空白。该数据库为解析RNA结构动态调控提供了核心数据资源。 研究方法: 本研究采用多层次技术路线系统性研究rG4与RNA修饰的关系:首先,整合了多个体内实验数据和七种预测工具的分层预测结果获得高可靠性rG4座位;其次,收集了7种脊椎动物和2个病毒中9类常见RNA修饰位点,筛选得到与rG4相关的修饰位点,并对rG4和RNA修饰位点分别建立了三级可信度评估体系;最后,运用ThT/SYBR Gold凝胶染色、配体增强荧光光谱、圆二色谱和紫外熔解等实验手段验证rG4形成能力和热稳定性,并通过EMSA评估RNA修饰在LIN28A蛋白与rG4相互作用中的调控角色。 主要结果: 1. 成功构建了名为MoRNiNG的数据库,收录九个物种中rG4内的九种RNA修饰位点。 2. 为m6A修饰调控rG4热稳定性,进而影响RNA与m6A anti-reader LIN28A的结合提供了实验证据。 3. 多个物种的转录组中众多rG4形成序列既具有形成茎环结构的潜力,又在其区域内同时存在RNA修饰。 4. 体外实验验证了m6A、m5C及A-to-I 编辑等多种RNA修饰均可影响rG4的属性,如热稳定性、多聚化行为及rG4结构的形成能力。 数据链接或代码连接或其他: https://www.cityu.edu.hk/bms/morning

Page qzaf106


Database

cfMethDB: A Comprehensive cfDNA Methylation Data Resource for Cancer Biomarkers

Yuanhui Sun, Zhixian Zhu, Qiangwei Zhou, Zhe Wang, Yuying Hou, Xionghui Zhou, Guoliang Li

Cancer is a major global health threat, and early detection is crucial for improving patient outcomes. DNA methylation in circulating cell-free DNA (cfDNA) has emerged as a promising biomarker for non-invasive cancer diagnosis. However, the integration and utilization of existing cfDNA methylation data have been limited, hindering comprehensive research efforts, particularly in the discovery of cfDNA methylation biomarkers. To address this challenge, we introduced cfMethDB, a comprehensive database dedicated to cfDNA methylation in cancer that encompasses 4828 publicly available datasets. Through standardized analysis, we identified 1,048,770 differentially methylated cytosines (DMCs) as candidate biomarkers across seven cancer types. With cfMethDB, we not only identified known cfDNA methylation biomarkers, but also discovered several genes, such as ZIC4, that could be novel biomarkers. Moreover, cfMethDB offers a suite of user-friendly tools, including biomarker evaluation, pan-cancer search, and end motif analysis. We hope that cfMethDB will serve as a valuable platform for the discovery of novel cancer cfDNA methylation biomarkers and facilitate cancer research and clinical applications. cfMethDB is publicly available at https://cfmethdb.hzau.edu.cn/home.
研究问题: 癌症是全球范围内的重大健康威胁,而早期检测对改善患者预后具有决定性作用。细胞游离DNA(Cell-free DNA,cfDNA)的DNA甲基化已成为极具前景的无创癌症诊断生物标志物。然而,目前公开的cfDNA甲基化数据在整合与应用方面仍相对有限,从而阻碍了新型甲基化标志物的系统发现。如何对这些数据进行系统整合、统一处理和高效利用,是亟待解决的重要问题。 研究方法: 为了构建综合性癌症cfDNA甲基化数据库cfMethDB,本研究从SRA和GSA数据库收集截至2024年6月的cfDNA甲基化数据,采用统一的质控与预处理流程(fastp、Trim Galore、BatMeth2等)对原始测序数据进行标准化处理。在过滤比对率低、质量差和覆盖度不足的数据后,最终保留4828个高质量数据集,并计算每个胞嘧啶位点的甲基化水平,同时依据CHG甲基化率评估亚硫酸氢盐转化效率。在差异甲基化分析部分,使用SMART2软件鉴定全基因组范围内的差异甲基化胞嘧啶位点(Differentially methylated cytosine,DMC),并通过ChIPseeker进行功能注释。cfMethDB数据库由MySQL、Django、Nginx与JBrowse工具驱动,可为用户提供高效的数据检索、可视化展示与分析支持。 主要结果: 1.构建了包含7种癌症类型的cfDNA甲基化数据库,包含Whole Genome Bisulfite Sequencing(WGBS)、Reduced Representation Bisulfite Sequencing(RRBS)和Padlock等多种DNA甲基化测序技术数据。 2.在全基因组水平鉴定了1,048,770个DMC位点,并发现大量泛癌DMC位点,可作为潜在的生物标志物。 3.cfMethDB提供包括基因搜索(Gene search)、区间搜索(Region search)、泛癌DMC分析、区间标志物诊断性能评估模型、区间末端特征序列模式分析和基因组浏览器等在线检索和分析工具。 4.cfMethDB提供批量数据下载和应用程序编程接口(Application programming interface,API),方便与其他数据库集成与二次分析。 数据库链接: https://cfmethdb.hzau.edu.cn/home

Page qzaf092


Application Note

ClusterGVis: An Advanced Visualization and Clustering Tool for Gene Expression Analysis

Jun Zhang, Hongyuan Li, Wenjun Tao, Jun Zhou

Both single-cell RNA sequencing and bulk RNA sequencing data provide valuable insights into physiological and pathological processes. The effective interpretation of such data relies on the availability of sophisticated analytical and visualization tools. Here, we introduce ClusterGVis, an advanced bioinformatics software specifically designed to simplify the analysis and visualization of gene expression data. ClusterGVis provides a user-friendly interface that allows researchers to perform fuzzy c-means and k-means clustering on transcriptomic data. It enables researchers to effectively uncover patterns and relationships within complex gene expression profiles. The integrated heatmap visualization features support intuitive exploration of co-expression networks and identification of differentially expressed genes across diverse experimental conditions. ClusterGVis serves a dual purpose: aiding in the identification of potential biomarkers and enriching the understanding of gene function and regulatory mechanisms. The tutorials, manual, source code, and demo data of ClusterGVis are publicly available at https://github.com/junjunlab/ClusterGVis and https://bioconductor.org/packages/ClusterGVis. The ClusterGVis Shiny app has been deployed on shinyapps.io and is accessible at https://laojunjun.shinyapps.io/clustergvis_app_v0/. The Shiny app source code is hosted on GitHub at https://github.com/junjunlab/ClusterGvis-app.
研究问题: 高通量测序产生的基因表达数据分析对阐明生物学机制至关重要。然而在基因组学领域,现有的基因聚类、通路富集分析及其可视化工具往往呈现碎片化且相互独立的状态。虽然已有多种功能富集R包,但仍显著缺乏能够根据不同表达模式对基因进行聚类、并进行批量富集分析和可视化的工具。因此,如何构建一个集成化且用户友好的生物信息学工具,以实现基因表达数据的系统聚类、批量富集与综合性可视化,已成为提升非生物信息学背景研究人员组学数据分析效率、加速生物学机制发现过程中亟待解决的关键问题。 研究方法: 基于R语言开发环境,使用K均值(K-means)及模糊C均值聚类(Fuzzy c-means clustering,FCM)算法对基因表达数据进行聚类获得不同表达趋势的基因亚群,使用clusterProfiler R包对不同亚群进行功能富集分析,采用ComplexHeatmap R包对分析结果进行整合可视化展示,最后基于R Shiny构建交互式网页应用端分析工具。 主要结果: 1.开发ClusterGVis R包,简化对基因表达数据的聚类、功能富集分析及可视化步骤。 2.支持对接加权基因共表达网络分析(Weighted gene co-correlation network analysis,WGCNA)、Seurat和Monocle等经典分析软件输出结果。 3.基于Shiny开发用户端交互分析平台,支持在线和本地运行分析。 数据链接或代码连接或其他: R包网址: Bioconductor:https://bioconductor.org/packages/ClusterGVis Github:https://github.com/junjunlab/ClusterGVis Shiny网址: Shiny应用端:https://laojunjun.shinyapps.io/clustergvis_app_v0/ Shiny代码:https://github.com/junjunlab/ClusterGvis-app

Page qzag005