Chattopadhyay, Amit K., Abdollahinejad, Yeganeh and Flower, Darren R. (2026) PyScale: A Curated Nucleotide Propensity-Scale Resource With Reproducible Profile Analysis and a Robust Profile-Level Divergence Index. IEEE Access, 14. pp. 144001-144014. ISSN 2169-3536
Preview |
PDF
Download (1MB) | Preview |
Abstract
Nucleotide sequences encode functional and structural information that is not always easy to detect using alignment algorithms alone, particularly when sequences are highly diverged, rearranged, or contain complex repeat structures. Propensity-scale profiling offers an alignment-free alternative by converting sequences into quantitative profiles reflecting physicochemical and structural properties. This approach is well established for proteins but less standardized for nucleic acids. We present PyScale, an algorithmic toolkit that curates nucleotide propensity scales and provides reproducible software to transform DNA or RNA sequences into time-series-like profiles for statistical analysis. PyScale bundles 306 nucleotide scales, including 243 mononucleotide, 46 dinucleotide, and 17 trinucleotide scales, and supports four paradigms: window averaging with variable step sizes, periodicity assumption, reverse reading, and complementary strand conversion. To characterize sequence irregularity, PyScale incorporates a Rosenstein-style profile divergence exponent, a comparative profile-level index that is distinct from a dynamical-system Lyapunov exponent, validated using the chaotic logistic map. To improve reproducibility, the estimator adds K-nearest-neighbor averaging and automatic fit-region selection, which reduce the variance of estimates across trials by roughly 50% and remove dependence on analyst-chosen fit intervals, at the cost of a negative bias in the point estimate relative to the single-neighbor baseline (7.9% for automatic fit). Incorporating both a Python command-line pipeline for batch processing and a graphical interface for interactive visualization and multi-sequence comparison, PyScale provides an accessible framework for alignment-free nucleotide sequence characterization using interpretable property profiles and reproducible profile-divergence estimation. PyScale can be seamlessly adapted for protein sequence analysis.
| Item Type: | Article |
|---|---|
| Additional Information: | This work has a Creative Commons CC BY 4.0 International License: https://creativecommons.org/licenses/by/4.0/ |
| Uncontrolled Keywords: | Alignment-free sequence analysis; bioinformatics software; DNA and RNA visualization; Lyapunov exponent; nucleotide propensity scales; periodicity; Rosenstein algorithm; sequence profiling; window averaging |
| Subjects: | H Social Sciences > HA Statistics Q Science > QR Microbiology Q Science > QA Mathematics > Algebra > Algorithms Q Science > QH Natural history > QH301 Biology > Methods of research. Technique. Experimental biology > Data processing. Bioinformatics |
| Divisions: | School of Business and Social Sciences > Staff Research and Publications |
| Depositing User: | Tamara Malone |
| Date Deposited: | 25 Sep 2026 14:18 |
| Last Modified: | 25 Sep 2026 14:18 |
| URI: | https://norma.ncirl.ie/id/eprint/9943 |
Actions (login required)
![]() |
View Item |
Tools
Tools