Shannon entropy is a standard proxy for DNA sequence complexity, but under fixed-length sampling it often exhibits saturation and limited dynamic range, reducing its utility for discriminating sequences that share similar base composition. We introduce the entropy–rank ratio, a distribution-aware complexity rank defined as the cumulative proportion of all length- blocks (under non-overlapping -tuples) whose Shannon entropy is less than or equal to that of the target block. By calibrating entropy against its attainability distribution at fixed, provides a comparable and expanded complexity scale under a constrained protocol. We operationalize as a crop-selection signal for fixed-length data augmentation: among candidate crops, we select segments whose value best matches the full-sequence complexity signature while penalizing excessive displacement from the center. Across two independent classification problems (viral gene classes and human genes with polynucleotide expansions), ratio-guided cropping yields consistent improvements over random, entropy-based, and compression-based crop heuristics when training lightweight CNNs in small labeled-data regimes. Because the method is model-agnostic and the reference distribution can be cached per, it scales to the sequence lengths used in our benchmarks while preserving a strict fixed-length input protocol.
Entropy–rank ratio: a novel entropy–based perspective for DNA complexity and classification
Pastore, Emmanuel Pio
;Passarino, Giuseppe;Sapia, Peppino;De Rango, Francesco
2026-01-01
Abstract
Shannon entropy is a standard proxy for DNA sequence complexity, but under fixed-length sampling it often exhibits saturation and limited dynamic range, reducing its utility for discriminating sequences that share similar base composition. We introduce the entropy–rank ratio, a distribution-aware complexity rank defined as the cumulative proportion of all length- blocks (under non-overlapping -tuples) whose Shannon entropy is less than or equal to that of the target block. By calibrating entropy against its attainability distribution at fixed, provides a comparable and expanded complexity scale under a constrained protocol. We operationalize as a crop-selection signal for fixed-length data augmentation: among candidate crops, we select segments whose value best matches the full-sequence complexity signature while penalizing excessive displacement from the center. Across two independent classification problems (viral gene classes and human genes with polynucleotide expansions), ratio-guided cropping yields consistent improvements over random, entropy-based, and compression-based crop heuristics when training lightweight CNNs in small labeled-data regimes. Because the method is model-agnostic and the reference distribution can be cached per, it scales to the sequence lengths used in our benchmarks while preserving a strict fixed-length input protocol.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


