Archiving Debt Securities on a Synthetic DNA Molecule, in the Age of Scriptural Money
This text is the study I submitted on 19 September 2024 for the Special Prize of the Public Treasury of Benin, on the theme "The Treasury's banking function and the optimization of State cash management", supervised by Bruno Monkoun, a biologist.
Abstract
The shift toward a digital economy brings with it a host of constraints, many of which can be anticipated through the implementation of an effective data management policy.
This study fits into that context, proposing a new approach to optimizing State cash management, with a particular focus on the banking function of the Public Treasury and the archiving of debt securities.
It examines the limitations of current data storage methods, whether traditional systems or modern databases, in a context of exponentially growing data volumes and environmental constraints.
To address this, an innovative solution has been identified : the use of the synthetic DNA molecule as a long-term archiving medium.
The study explores the advantages of this technology, including its capacity to store large amounts of data in a small space, as well as its benefits in terms of durability and energy efficiency, and concludes with recommendations for implementing this technology and with its strategic implications for the future.
Keywords of the study : Treasury, scriptural money, database, archiving, deoxyribonucleic acid (DNA).
General Introduction
In a world where data management has become a crucial issue for governments and financial institutions, the Public Treasury, as the guarantor of the sound management of public funds, faces a twofold challenge : on the one hand, the need to modernize its management tools to keep pace with financial technology, and on the other, the need to reduce its ecological footprint while optimizing the efficiency of its processes.
At the same time, with globalization, States have no doubt had to keep reinventing their economic model, notably by giving the Treasury the ability to issue debt securities to meet the State's financing needs.
As a result, the archiving of debt securities, financial transactions and accounting documents, historically carried out on physical and digital media, now shows its limits in terms of cost, durability and security.
Moreover, the growing volume of data requires massive, energy-hungry infrastructures, namely data centers, whose electricity consumption and cooling needs pose major environmental challenges.
Faced with these challenges, this study explores an innovative and promising solution : archiving data on a synthetic DNA molecule. This technology, although still in its infancy, offers extremely dense and durable storage capacity while consuming very few resources.
Taking as a case study the banking function of the Treasury and the optimization of State cash management, this study aims to show how the synthetic DNA molecule could not only meet current storage needs, but also make it possible to design data management strategies based on artificial intelligence technologies that are sustainable and efficient in the future.
Literature Review
Money and the State
The Evolution of Money
Originally, the value of a coin was determined by the intrinsic value of the materials used to make it, such as gold or copper. For example, a dollar was once worth about 1.5 grams of gold (the ounce of gold, 31.1 g, was set at 20.67 dollars). This is the concept of the gold standard.
This meant that anyone holding dollars anywhere in the world could, at any time, exchange them for the corresponding amount of gold. This system rested notably on the idea that each unit of currency was backed by a real gold reserve, which guaranteed its value. Gold coins are commodity money, and convertible banknotes representative money : their value lay in the promise that they could be exchanged for gold. The term fiduciary money (from the Latin fides, trust) applies above all to what came afterward : money that is worth only the trust placed in it, also called fiat money.
However, not every country could follow this model, because of the difficulty of gathering enough gold reserves. The solution found was for other currencies to be pegged to the US dollar, and for the dollar in turn to be pegged to gold : this is the basis of the Bretton Woods agreements (1944).
Under this system, the dollar remained convertible into gold, at 35 dollars an ounce since 1934, but only for foreign central banks ; the United States therefore had to hold enough gold to honor these requests. But as global demand for dollars grew, the United States was forced to issue more dollars than it had gold to back them. This no doubt led to the end of the Bretton Woods agreements, in 1971.
After the end of the Bretton Woods agreements, the dollar and the other currencies ceased to be directly convertible into gold. This lifted the last physical constraint on the creation of money ex nihilo, that is, money resting on nothing. Money therefore no longer needed to be backed by a physical reserve. Without this bound, money can be issued without limit, and lose value if it is issued in excess : this is inflation.
Today, the fiduciary nature of money rests on the trust people place in central banks. That is, if the central bank states that a banknote is worth 20 dollars, that banknote can be used to buy goods or services worth 20 dollars. And it is up to the central bank to guarantee the stability of the currency, by acting as a regulator.
With the computer revolution, the whole financial sector went digital, moving from a paper-based system to a digital economy. As a result, scriptural money (bank money), based on accounting entries, that is, a series of figures written in a computer, became commonplace. This money rests on the promise that your bank will pay you back the equivalent in physical money whenever you need it.
Thus, when you deposit money with your usual bank (a retail bank), you are in fact lending it that money. In return, the bank promises to give it back to you at any time.
In an interbank transfer, the two retail banks rely on the central bank, which acts as a trusted third party. The central bank is entitled to lend to retail banks, just as they lend to individuals. In other words, the sending bank arranges to hold the equivalent of the amount to be sent as a promise from the central bank, which it then sends to the other bank.
In reality, however, retail banks all hold an account at the central bank, and interbank transfers are carried out through a clearing mechanism, which consists of adjusting the balances of these accounts.
To sum up :
- individuals trust retail banks to keep their accounts ;
- retail banks trust the central bank to keep theirs ;
- but who keeps the State's accounts ?
The Treasury
The Treasury is the body responsible for managing public finances, in that it records the State's revenue and expenditure. It plays a crucial role in the economy, by enabling the State to raise funds through the issuance of debt securities, such as Treasury bills. These securities, issued on behalf of the State, can then be bought by companies or individuals with a view to making a future profit.
For example :
- the State needs 100 million francs to finance a project ;
- the Treasury can issue 100 debt securities worth 1 million each ;
- these securities are then bought by companies ;
- the purchase is made under a contract, with the benefits that come with it.
The debt securities issued by the Treasury include :
- Treasury bills : short-term maturities (up to 1 year, and up to 2 years on the market of the West African Economic and Monetary Union, WAEMU) ;
- government bonds : medium- and long-term maturities (from 2 years to more than 10 years) ;
- inflation-indexed bonds : protected against inflation, with medium- or long-term maturities.
Physical documents, such as receipts or contracts, serve as proof during a transaction related to the issuance of a Treasury product.
However, financial management is mainly digital, with records in databases that follow the life cycle of the transaction.
Databases
Using Files to Store Data
From the appearance of the first computers until the 1960s, company data was traditionally organized as files.
Example :
But this way of doing things soon reached its limits.
Even for a single user, it became difficult to keep the lines consistent, or to search for complex records.
The situation grew worse as the number of users increased, for the simple reason that each of them used a different version of the file, leading to conflicts and errors, for lack of synchronized information.
Moreover, the very organization of the files posed another challenge : each file was often tailored to a particular processing task.
All this created a great deal of data redundancy, needlessly weighing down the system, delaying data processing and thus increasing the risk of inaccuracies.
Around 1965, a new idea emerged : separating data from its processing, with the advent of databases. Databases not only fix the problem of data integrity, but also allow a more centralized organization and a more efficient management of data.
Transition to Databases
A database is designed to record facts and events within an organization, either to return them on demand or to draw conclusions by cross-referencing several pieces of information.
In the case of the Public Treasury, it is essential to store all information relating to financial transactions, by centralizing in a database the :
- revenue : amounts collected by the government through taxes, duties, fines, etc. ;
- expenditure : details of payments made for salaries, equipment purchases, subsidies, etc. ;
- public accounts : tracking of public account balances and financial movements.
Beyond centralizing this data, the database makes it possible to establish links between the pieces of data, to facilitate complex analyses. For example :
- an expense is linked to a specific public account, which shows which account the money was taken from ;
- a revenue item is associated with a type of tax, making it possible to see each tax type's contribution to overall revenue.
Seen from the perspective of older design methodologies, each part of this example would require a new application, with its own files and its own programs. Creating a database goes against this way of doing things, by making it possible to centralize, coordinate, integrate and disseminate information.
More generally, database management systems (DBMS) offer several significant advantages over simply managing data as files. These include :
- Organization and structuring of data : DBMSs make it possible to define schemas that structure data as tables, with columns defined by specific data types. This guarantees the consistency and integrity of the data.
- Simple and efficient access : DBMSs offer a query language (such as SQL) for retrieving, inserting, updating or deleting data efficiently and quickly, even for large datasets.
- Scalability : DBMSs are designed to scale and handle massive amounts of data, in terms of both storage and query performance.
- Multi-user support : they efficiently manage simultaneous access by multiple users, ensuring that operations run in a coordinated way, without interference.
Does the Database Replace Files ?
To answer this question, we need to understand the very structure of a database management system (DBMS). A DBMS consists of three layers of functions, stacked from secondary storage up to the users :
- File management system (FMS) : stores data on secondary storage (such as hard drives).
- Internal DBMS : manages the links between the data stored in the files, which become invisible to application programs.
- External DBMS : handles the analysis and interpretation of user queries, as well as the formatting of the data exchanged with the outside world.
Thus, the database does not replace files, but offers a layer of abstraction on top of them. This means that files still exist in the background, but that the database makes it possible to manage them in a more structured and efficient way. In other words, within the database itself, it is files that are being managed, but transparently for the user.
The Data Life Cycle
From a macro view of the Treasury's information system, data management goes through several important stages, those of information life cycle management, of which electronic document management (EDM) is the counterpart for documents and contracts.
Data Collection
This is the first stage, in which data is gathered and stored in an active database (hot data). In other words, it is a systematic approach that consists of gathering and measuring different pieces of information coming from various sources. This approach helps to get an overall view of the business, as explained in section 1.2.
Snapshot and Backup
At regular intervals, the data is offloaded to other locations, generally as files, independently of the original data. A snapshot of the database is created, read-only by default, to preserve the state of the data at that moment t. Capturing the state of a system at a precise moment is called a snapshot, while a backup is the act of copying the data, here the snapshot, to another location.
Data backup is a data protection function, essential for reducing the risk of total or partial data loss in the event of unforeseen events. It gives organizations the ability to restore their systems and applications to a desired earlier state. To reduce the storage space taken by the data to be backed up, it is often recommended to compress the files beforehand, in order to reduce their size.
A backup removes nothing from the active database : it keeps a restorable copy of it. Removing old data from the database is a separate operation, decided on its own. This lighter load mainly benefits future operations that scan an entire table (reports, backups…).
Data Backup Methods
Data can be backed up in various ways. Some methods back up a full copy of the data each time, while others only copy the changes made to the data. Each method has its advantages and drawbacks.
- Full backup : full backups take a complete copy of all the data each time, stored as is, or compressed and encrypted. Synthetic full backups create full backups from a full backup and one or more incremental backups.
- Incremental backup : incremental backups copy all the data that has changed since the last backup, whatever that backup's method. Reverse incremental backups add all the changed data to the last full backup.
- Differential backup : differential backups copy all the data changed since the last full backup, whether or not another backup was made in between with another method.
- Mirror backup : a mirror backup is stored in an uncompressed format that mirrors all the files and configurations of the source data. It can be accessed just like the original data.
Versioning
Keeping a history of the backed-up versions proves useful when the chosen backup method is not a full backup, and even more so if the data is purged from the active database at each backup.
Thus, when certain changes to the system lead to undesirable results (since the data structure is likely to change over time), the Treasury can restore a snapshot of the system at a given point in time.
The versioning frequency follows that of the backups (daily, weekly, etc.). This concerns frequently accessed data.
Archiving
Version history, although essential, cannot be kept indefinitely. Over time, the stored versions may go back several months, or even years. It then becomes necessary to archive the oldest data, which requires an archiving plan.
An archiving plan is a set of practices and rules for managing and durably preserving an organization's documents and data, whether physical or digital.
Historically, archiving mainly concerned physical documents, such as contracts, but with the advent of digital tools, many companies now favor electronic archiving.
This transition is driven by cost reduction, not by risk : while physical archiving exposes documents to possible deterioration, due to the weather or to the conditions in which the archives are kept, digital archiving also exposes data to cyberattacks, or to hardware failures.
This second option can also be applied to physical documents. They just need to be digitized first : this is the dematerialization of data, or document digitization.
However, unlike a simple data backup, archiving meets regulatory requirements. For example, in the OHADA area, to which Benin belongs, accounting books and supporting documents must be kept for ten years (Uniform Act on Accounting Law and Financial Reporting, art. 24), as in France (Commercial Code, art. L123-22) ; for public accounting, retention periods fall under national archives regulations.
The Big Data and Cloud Computing Approaches
The archiving plan requires a centralized space with a large storage capacity to keep the data : a data center. A data center houses both hardware (hard drives…) and software (a database management system…).
In short, the data center will host the data in all its states throughout its life cycle : from its generation to its archiving.
Big Data
Despite the efficiency of data centers for the traditional archiving plan, the rapid development of technologies and the explosion of data volumes pushed companies beyond the capabilities of conventional database management systems (DBMS).
Google, in particular, pioneered pushing back these limits of conventional DBMSs and helped reinvent the field, with new tools for big data.
New types of DBMS therefore emerged in the software layer, while, in the hardware layer, the processing of these vast volumes of data was spread across clusters of ordinary machines, rather than entrusted to ever more powerful processors.
Cloud Computing
Setting up a big data system requires a robust infrastructure. This can mean high costs for its acquisition and maintenance…
On the other hand, the first players to face these challenges (Google, Amazon, Microsoft) quickly had to find solutions, and even build infrastructures beyond comprehension. A little later came the idea of cloud computing.
In essence, cloud computing makes it possible to use other companies' data center infrastructure. It offers the same data storage and processing advantages, without having to manage the infrastructure yourself. It is like having access to a complete data center without owning one, in exchange for a subscription : this is Infrastructure as a Service (IaaS).
The main cloud services today are :
- GCP : Google Cloud Platform, from Google ;
- AWS : Amazon Web Services, from Amazon ;
- Azure : owned by Microsoft.
By choosing IaaS, the Treasury is provided with the data center's infrastructure only. Managing the archiving plan, on the other hand, remains its responsibility.
However, some cloud services go even further, offering software built into the cloud platform that takes care of the archiving process in a few clicks : this is Software as a Service (SaaS).
Methodology
In the previous chapter (section 1.2.3), we looked at the internal and external layers of DBMSs. In this one, we will go a little deeper, down to the bottom layer : the file management system.
To do so, it is essential to first understand how the machine writes information to the disk.
Indeed, every file, before being used by the machine, must be converted into binary, made up only of 0s and 1s, the only language computers understand.
But first, it is important to point out that each letter of the keyboard is associated with an ASCII code, and it is rather the list of these correspondences that will be translated into binary. Converting ASCII to binary therefore amounts to going from the decimal system to the binary system.
ASCII only encodes 128 characters, without accented letters ; a French text today goes through Unicode, whose UTF-8 encoding keeps ASCII codes on one byte and uses two to four bytes for the other characters. The sample file below only uses ASCII characters : one byte each.
The sample file, "Cotonou, le 05 juin 2024" (French for "Cotonou, 5 June 2024"), translated into binary :
| Character | ASCII code | ASCII code in binary (8 bits) |
|---|---|---|
| C | 67 | 01000011 |
| o | 111 | 01101111 |
| t | 116 | 01110100 |
| o | 111 | 01101111 |
| n | 110 | 01101110 |
| o | 111 | 01101111 |
| u | 117 | 01110101 |
| , | 44 | 00101100 |
| space | 32 | 00100000 |
| l | 108 | 01101100 |
| e | 101 | 01100101 |
| space | 32 | 00100000 |
| 0 | 48 | 00110000 |
| 5 | 53 | 00110101 |
| space | 32 | 00100000 |
| j | 106 | 01101010 |
| u | 117 | 01110101 |
| i | 105 | 01101001 |
| n | 110 | 01101110 |
| space | 32 | 00100000 |
| 2 | 50 | 00110010 |
| 0 | 48 | 00110000 |
| 2 | 50 | 00110010 |
| 4 | 52 | 00110100 |
Binary written on 8 bits, that is, with 8 digits (0 or 1), represents a byte.
The sequence that will finally be written to the machine is the following (all these bytes concatenated, one after the other) :
0100001101101111011101000110111101101110011011110111010100101100
0010000001101100011001010010000000110000001101010010000001101010
0111010101101001011011100010000000110010001100000011001000110100
Once this binary sequence corresponding to the input file has been obtained, it can be recorded on a long-term storage medium, such as an external device.
Writing the Binary File to an Optical Medium
Traditionally, writing a binary file to an optical disc, such as a DVD or a Blu-Ray, consisted of burning the binary onto its surface with a burner.
The burner uses a laser beam to record the information as small marks, which can then be read by a compatible player.
To produce the 1s and 0s, the DVD burner switches the beam on and off.
Indeed, when the laser is on, it heats certain parts of the recordable layer to create pits. The unburned areas, called lands, remain intact.
These differences between the pits and the flat areas are what the laser reader interprets as 1s and 0s when reading the data.
A clarification : on a pressed disc, it is actually the transitions between pits and lands that encode the 1s, the length of each area encoding a run of 0s ; and on a recordable disc, the laser does not dig into the surface, it changes the reflectivity of a dye layer. The principle remains the same : two states, readable by the beam.
Taking the binary of "Cotonou, le 05 juin 2024", we burn it as follows :
DNA as an Information Medium
Our body is mostly made up of cells, which themselves contain organelles, including the nucleus, where our genetic material is found.
This genetic material consists of natural DNA (deoxyribonucleic acid) molecules coiled on themselves, forming chromosomes in the nucleus.
Each DNA molecule carries the genetic information essential to the growth, development and functioning of all living organisms.
A Closer Look at the Natural DNA Molecule
DNA is a molecule made of two strands wound around each other, forming a double helix. The strands are made of four different nucleotides, with their nitrogenous bases facing the inside of the molecule : adenine (A), thymine (T), cytosine (C) and guanine (G).
The A bases of one strand always face the T bases of the other strand, while the G bases always face the C bases : this is base complementarity.
Genetic information is determined by the order of the bases in the DNA sequence ; that is, the specific composition of this sequence defines the genes, which are nothing other than the instructions that determine the characteristics of an individual, or of any living being in general.
DNA Seen through Numeral Systems
DNA can be seen as a long text written with a four-letter alphabet (A, T, G, C), representing the four nitrogenous bases.
Mathematically speaking, our default system for writing numbers is the decimal system, made up of 10 digits (0, 1, 2, 3, 4, 5, 6, 7, 8, 9).
But besides this system, we also have :
- the binary system, with an alphabet of 2 digits (0, 1) ;
- the quaternary system, with an alphabet of 4 digits (0, 1, 2, 3) ;
- …
A number expressed in a given numeral system can always be converted into another system, and vice versa : decimal → binary, binary → decimal.
Writing DNA, which uses four symbols, therefore follows a quaternary system.
But to go from standard quaternary (0, 1, 2, 3) to DNA quaternary (A, T, C, G), a correspondence must be established between these two sets of symbols :
0→A 1→T 2→C 3→G
Based on this convention, "Cotonou, le 05 juin 2024" becomes :
| Character | ASCII code | ASCII in standard quaternary | ASCII in DNA quaternary |
|---|---|---|---|
| C | 67 | 1003 | TAAG |
| o | 111 | 1233 | TCGG |
| t | 116 | 1310 | TGTA |
| o | 111 | 1233 | TCGG |
| n | 110 | 1232 | TCGC |
| o | 111 | 1233 | TCGG |
| u | 117 | 1311 | TGTT |
| , | 44 | 0230 | ACGA |
| space | 32 | 0200 | ACAA |
| l | 108 | 1230 | TCGA |
| e | 101 | 1211 | TCTT |
| space | 32 | 0200 | ACAA |
| 0 | 48 | 0300 | AGAA |
| 5 | 53 | 0311 | AGTT |
| space | 32 | 0200 | ACAA |
| j | 106 | 1222 | TCCC |
| u | 117 | 1311 | TGTT |
| i | 105 | 1221 | TCCT |
| n | 110 | 1232 | TCGC |
| space | 32 | 0200 | ACAA |
| 2 | 50 | 0302 | AGAC |
| 0 | 48 | 0300 | AGAA |
| 2 | 50 | 0302 | AGAC |
| 4 | 52 | 0310 | AGTA |
By concatenating all these quaternary codes, one after the other, we get the first strand :
TAAGTCGGTGTATCGGTCGCTCGGTGTTACGAACAATCGATCTTACAA
AGAAAGTTACAATCCCTGTTTCCTTCGCACAAAGACAGAAAGACAGTA
Next, we take the complement of this sequence (A facing T, C facing G) to obtain the second strand :
ATTCAGCCACATAGCCAGCGAGCCACAATGCTTGTTAGCTAGAATGTT
TCTTTCAATGTTAGGGACAAAGGAAGCGTGTTTCTGTCTTTCTGTCAT
A clarification : this direct encoding, one base per pair of bits, is simplified for teaching purposes ; our own sequence bears its mark, with the repeat "ACAAAGAAAG". Real systems avoid such homopolymers, a source of sequencing errors, balance the proportion of G and C, and add error-correcting codes as well as addresses to locate a given fragment (Goldman et al., 2013 ; Erlich and Zielinski, 2017 ; Organick et al., 2018).
We have thus encoded our file on a DNA molecule. To retrieve the original text, we have to go the other way.
Results
Writing in quaternary rather than binary is not in itself an advantage : each DNA base carries 2 bits, no more and no less than the same information spread over twice as many binary units. Encoding the phrase "Cotonou, le 05 juin 2024" thus required 192 bits in binary (section 2.1), exactly 96 bases in DNA quaternary : the same amount of information, in a form twice as compact in number of units, since each base carries two. The complementary strand does not change this count : it carries no additional information ; it doubles the material, not the capacity. In real systems, data is in fact written on short single strands, of 150 to 200 bases, not on a complete double helix.
What really sets DNA apart is the physical size of its memory unit, not the number of symbols in its alphabet. In DNA storage, the memory unit is the nucleotide, which carries two bits ; in flash memory, it is the cell, which today carries three or four. What makes the difference is the size of the unit : a flash cell measures a few tens of nanometers and is stacked over hundreds of layers, whereas a DNA base takes up 0.34 nanometers along a molecule 2 nanometers wide. Relative to volume or mass, the theoretical density of DNA is on the order of several hundred exabytes per gram (Church et al., 2012), and 215 petabytes per gram have been achieved in the laboratory (Erlich and Zielinski, 2017) : several orders of magnitude above current media.
What makes this technology even more promising is that data centers today consume not only enormous amounts of electricity (about 460 TWh in 2022, nearly 2% of the world's electricity according to the International Energy Agency, which would place them around tenth among the world's consuming countries if they were one), but also hundreds of billions of liters of water for their cooling.
Thus, the digital sector alone produces as much greenhouse gas as air traffic, which raises questions of corporate social responsibility (CSR), with an environmental impact that is both physical (water consumption) and digital (electricity consumption).
So many issues that DNA storage can mitigate : once written, the molecule consumes nothing, whereas a disk or a tape must be powered, cooled and replaced. Writing (synthesis) and reading (sequencing), on the other hand, remain costly in energy and reagents, and preservation requires cold storage or encapsulation (section 6). This would therefore contribute to a significant reduction in the carbon footprint of the digital sector as a whole. And the Treasury, as a public body, has a responsibility to lead by example and to promote best practices.
Implications
Following on from section 1.3.5, let us ask ourselves : what happens to the data once the regulatory archiving period is over ?
Indeed, this "cold" data may then go through a destruction process, to free up storage space. But the synthetic DNA molecule stands out as an innovative solution for long-term storage, making it possible to keep information for centuries, and for millennia if the molecule is encapsulated in silica and kept cold (Grass et al., 2015).
A small clarification : the synthetic DNA molecule will not replace traditional data centers, which remain very efficient for the operational management of data. But it could fit into the overall architecture, as the storage medium for "cold" data : the archives.
But What For ?
With the rise of artificial intelligence (AI), keeping data over the long term is essential for building a sound AI strategy. In other words, there is no AI without data. By keeping historical data, the Treasury can accumulate rich and diverse datasets, which can help it improve the accuracy of its algorithms, thereby giving it a solid basis for in-depth analyses and better-informed decisions.
Moreover, unlike standard archiving, where data is often integrated into a single image or a single system, this new storage approach involves encoding the information in DNA capsules, each capsule being independent of the others. This means that the data is decentralized and protected against the loss or corruption of information, as well as against cyberattacks that could affect a centralized system.
In addition, this storage method improves data sovereignty, offering an alternative to cloud computing solutions, all the more so since this is sensitive State data. By using DNA as a storage medium, the Treasury commits to a dynamic of increased security and full control over its data.
Implementation Actions
Just as we presented the burner, in section 2.1, as the tool that turns digital data into a physical form on a medium, we need a tool to synthesize exactly the molecule corresponding to the sequence obtained in DNA quaternary : the DNA synthesizer. And, like the player for the disc, we need a second tool to read it back : the sequencer.
Thus :
- the synthesizer performs the encoding, by building the molecule base by base ;
- the sequencer performs the decoding, by reading the information encoded on the molecule.
This frees humans from this technical task.
The main challenge lies in the speed of synthesizing and reading the data. Currently, synthesizing data on DNA is slow and very expensive.
However, the creation in 2020 of the "DNA Data Storage Alliance" by four American companies (Western Digital, Microsoft, Twist Bioscience and Illumina), which aims to promote this technology and to standardize data encoding, shows the interest in the question. It would therefore be wise to see more States invest in this innovative technology, in order to speed up its development and support its evolution.
Recommendations
Over time, storage technologies, from the floppy disk to the DVD, have emerged and then disappeared, forcing us to constantly migrate our data to new media.
However, DNA, as a storage medium, stands out as a potentially eternal alternative. Unlike digital technologies, DNA has existed for billions of years, and it will always be possible to interpret it as long as there are forms of life.
That said, care should be taken to keep the synthesized DNA molecules cold, dry and away from light, isolated from water, oxygen and the enzymes that degrade them : in practice, encapsulated in silica beads, which has been shown to be the most durable protection.
Finally, although this technology is still in its infancy, its potential for ultra-durable data storage is immense.
General Conclusion
The transition to archiving on synthetic DNA will be a major technological leap for the Public Treasury. By enabling more efficient and more sustainable data management, this innovation meets the challenges posed by the exponential growth of data, and by the need for State bodies to have a reduced ecological footprint.
By adopting these new technologies, the State positions itself not only as a modern, forward-looking player in the management of its finances, but also as a model of environmental responsibility for public institutions.
Bibliography
- Cours bases de données
- Cours complet pour apprendre les systèmes de gestion de bases de données
- What Is Data Quality and Why Is It Important?
- Qu'est-ce que la sauvegarde de données ?
- Plan d'archivage : comment le mettre en place ?
- Les avantages et inconvénients de l'archivage physique et électronique
- Quels sont les délais de conservation des documents pour les entreprises ?
- Politique d'archivage de l'Inspection générale des finances, juin 2021
- Archives nationales du Bénin
- Politique nationale de développement des archives 2021-2030
- L'ADN
- L'ADN, support de l'information génétique
- Les chromosomes, structures universelles des cellules eucaryotes
- Le stockage de données sur ADN synthétique : une révolution nécessaire ?
- International Energy Agency, Electricity 2024, January 2024. iea.org
- DNA Data Storage Alliance
- G. M. Church, Y. Gao, S. Kosuri, "Next-generation digital information storage in DNA", Science, vol. 337, no. 6102, p. 1628, 2012.
- N. Goldman et al., "Towards practical, high-capacity, low-maintenance information storage in synthesized DNA", Nature, vol. 494, pp. 77–80, 2013.
- R. N. Grass, R. Heckel, M. Puddu, D. Paunescu, W. J. Stark, "Robust chemical preservation of digital information on DNA in silica with error-correcting codes", Angewandte Chemie International Edition, vol. 54, no. 8, pp. 2552–2555, 2015.
- Y. Erlich, D. Zielinski, "DNA Fountain enables a robust and efficient storage architecture", Science, vol. 355, no. 6328, pp. 950–954, 2017.
- L. Organick et al., "Random access in large-scale DNA data storage", Nature Biotechnology, vol. 36, pp. 242–248, 2018.
- L. Ceze, J. Nivala, K. Strauss, "Molecular digital data storage using DNA", Nature Reviews Genetics, vol. 20, pp. 456–466, 2019.
- OHADA, Acte uniforme relatif au droit comptable et à l'information financière [Uniform Act on Accounting Law and Financial Reporting], 2017, art. 24.
Illustrations : all the figures are redrawn as SVG from those in the study (a database course for the layers of a DBMS, biology textbooks for the cell and the DNA molecule, the diagram of an optical reader for the burner, the diagram of Oxford Nanopore's MinION for the sequencer). The figures that show the sample file, the two tables and the sequences are computed from the text "Cotonou, le 05 juin 2024" ; the sequencer's current trace is simulated for the first twenty bases of the strand. References 15 to 23 were added during the revision.
Want to talk about it?
Write to me at merlix@monkoun.com
or find me on LinkedIn.