The rapid growth of global data generation has far exceeded the physical and energetic limits of traditional silicon-based storage technologies. DNA has been explored as a molecular data storage medium because of its remarkable density and durability,...
The rapid growth of global data generation has far exceeded the physical and energetic limits of traditional silicon-based storage technologies. DNA has been explored as a molecular data storage medium because of its remarkable density and durability, but its high synthesis cost, slow read/write speed, and limited scalability restrict its practical use. To overcome these challenges, sequence-defined polymers (SDPs) have emerged as fully synthetic alternatives offering greater chemical diversity, environmental stability, and higher information density. SDPs can be precisely constructed via stepwise iterative or cross-convergent strategies, enabling molecular-level sequence control and scalable synthesis. Recent advances in decoding methods, including tandem mass spectrometry, spectroscopic techniques, and nanopore sequencing, further establish the foundation for practical, high-density molecular data storage systems based on SDPs.
In Chapter 2, nondestructive decoding of sequence-defined oligoesters composed of enantiopure α-hydroxy acids was demonstrated using 13C nuclear magnetic resonance (NMR) spectroscopy. Libraries of oligo(L-mandelic-co-D-phenyl lactic acid)s (oMPs) and oligo(L-lactic-co-glycolic acid)s (oLGs) were synthesized through a semiautomated flow system, which reduced reaction time to less than 1% of conventional batch processes. Decoding of these oligoesters was enabled by sequence-specific 13C NMR peaks arising from the distinct electronic shielding and deshielding effects of neighboring monomer units, allowing the precise identification of monomer sequences without polymer degradation. Using this approach, a 192-bit bitmap image was nondestructively decoded from 12 equimolar mixtures of 8-bit-storing oligomers, demonstrating the feasibility of permanent digital information storage in synthetic macromolecules.
In Chapter 3, a 512-mer sequence-defined polyester with a molecular weight of 57.3 kDa was synthesized to overcome the analytical size limit of mass spectrometry-based decoding. A four-bit fragmentation code was strategically embedded at aperiodic positions within the polymer sequence, enabling controlled cleavage into 18 readable oligomers without reducing storage capacity. Each fragment was individually sequenced by tandem mass spectrometry, and the complete 512-bit information was computationally reconstructed through an error-detection algorithm. This shotgun sequencing approach eliminates the storage limit of a single polymer chain and enables random access to specific data segments without full-chain sequencing.
In Chapter 4, the concept of SDPs was extended from polyesters to nucleobase-containing polymers capable of molecular recognition. Thymine- and adenine-functionalized isocyanides were used to synthesize uniform poly(hydroxybutyrate) derivatives via Passerini-iterative exponential growth (P-IEG), overcoming the sequence and length limitations of conventional methods. The resulting polymers exhibited specific adenine–thymine hydrogen bonding confirmed by 1H NMR, NMR titration, and variable-temperature studies. Furthermore, adenine and thymine were introduced in a defined order within a single polymer backbone, and the resulting SDPs were successfully sequenced by tandem mass spectrometry. In addition, the association constant of a mismatched sequence was also determined, allowing direct comparison between complementary and non-complementary arrangements. This chapter demonstrates that the P-IEG platform enables the construction of biomimetic macromolecules capable of encoding and reading molecular recognition patterns.
Unlike conventional data storage media such as silicon and DNA, which allow direct random access to specific information, data stored in synthetic polymers has traditionally required full-chain sequencing for retrieval. The studies presented in this dissertation demonstrate that selective decoding is achievable by molecular design. In Chapter 2, sequence-defined oligoesters were decoded nondestructively by 13C NMR spectroscopy, where only the NMR tube containing the desired oligomer needed to be analyzed, enabling targeted information access without degrading the entire library. In Chapter 3, random access was achieved by partial fragmentation of the polymer through chemical activation of the embedded fragmentation code, enabling selective identification of fragments containing the target word without full-chain sequencing. Finally, the nucleobase-containing sequence-defined polymers developed in Chapter 4 introduce the potential for selective molecular recognition through complementary base pairing, suggesting that future polymer-based information systems could achieve selective and programmable decoding analogous to biological processes.