Shannon entropy is a standard proxy for DNA sequence complexity, but under fixed-length sampling it often exhibits saturation and limited dynamic range, reducing its utility for discriminating sequences that share similar base composition. We introduce the entropy–rank ratio, a distribution-aware complexity rank defined as the cumulative proportion of all length- blocks (under non-overlapping -tuples) whose Shannon entropy is less than or equal to that of the target block. By calibrating entropy against its attainability distribution at fixed, provides a comparable and expanded complexity scale under a constrained protocol. We operationalize as a crop-selection signal for fixed-length data augmentation: among candidate crops, we select segments whose value best matches the full-sequence complexity signature while penalizing excessive displacement from the center. Across two independent classification problems (viral gene classes and human genes with polynucleotide expansions), ratio-guided cropping yields consistent improvements over random, entropy-based, and compression-based crop heuristics when training lightweight CNNs in small labeled-data regimes. Because the method is model-agnostic and the reference distribution can be cached per, it scales to the sequence lengths used in our benchmarks while preserving a strict fixed-length input protocol.

Entropy–rank ratio: a novel entropy–based perspective for DNA complexity and classification

Pastore, Emmanuel Pio
;
Passarino, Giuseppe;Sapia, Peppino;De Rango, Francesco
2026-01-01

Abstract

Shannon entropy is a standard proxy for DNA sequence complexity, but under fixed-length sampling it often exhibits saturation and limited dynamic range, reducing its utility for discriminating sequences that share similar base composition. We introduce the entropy–rank ratio, a distribution-aware complexity rank defined as the cumulative proportion of all length- blocks (under non-overlapping -tuples) whose Shannon entropy is less than or equal to that of the target block. By calibrating entropy against its attainability distribution at fixed, provides a comparable and expanded complexity scale under a constrained protocol. We operationalize as a crop-selection signal for fixed-length data augmentation: among candidate crops, we select segments whose value best matches the full-sequence complexity signature while penalizing excessive displacement from the center. Across two independent classification problems (viral gene classes and human genes with polynucleotide expansions), ratio-guided cropping yields consistent improvements over random, entropy-based, and compression-based crop heuristics when training lightweight CNNs in small labeled-data regimes. Because the method is model-agnostic and the reference distribution can be cached per, it scales to the sequence lengths used in our benchmarks while preserving a strict fixed-length input protocol.
2026
Data augmentation for convolutional neural networks
DNA sequence complexity
Entropy rank ratio (R)
Shannon entropy
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.11770/411677
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact