Reading the book of life: AI tool Helixer

Genome research has gained a new reading aid. The AI tool Helixer identifies where individual genes begin and end within a genome – and does so at astonishing speed.

Ein offenes Buch mit einem großen Baum, mehreren Pflanzen, einem Schmetterling, zwei Personen, einem Schwein, einer Schildkröte und einem Fisch darauf.

Imagine you’ve been given a book in a completely unknown language. No translation, no dictionary – just pages containing an endless sequence of only four letters. Now try to figure out where individual sentences begin and where they end. Researchers at the Institute of Bio- and Geosciences – Bioinformatics (IBG-4) and Heinrich Heine University Düsseldorf have trained an artificial intelligence system to solve this very task – for a special kind of “book”: the genome, the book of life. A living organism’s genome contains its blueprint – in its DNA.

Ein aufgeschlagenes Buch.

The AI tool Helixer can identify each individual sentence in the sequence of letters – in other words, every gene in an organism’s DNA sequence – and does so at remarkable speed. Genes are specific segments of DNA that contain the instructions for building various proteins. A key component of genes are the four bases adenine, guanine, cytosine, and thymine, abbreviated by their initial letters A, G, C, and T. However, both the number and the order of these four bases vary. A key difficulty lies in determining at which base a gene begins and where it ends. What complicates matters further is that there are also bases between the genes that do not belong to any gene. The human genome, for example, contains three billion bases, yet only around one third of these make up its roughly 20,000 genes. Helixer was not thrown off by this complexity, decoding the human genome in just a few hours. It required no prior knowledge to do so.

Eine Person mit Brille und grauer Jacke über einem dunklen Oberteil.

The AI tool thus solves a problem that has long held back genome research. DNA sequencing – in other words, determining the sequence of the letters – has been automated for many years. But actually “understanding” the sequence – i.e. which bases together form a gene – has previously required extensive and time-consuming laboratory work. “For almost two decades, there were no fundamentally new approaches in this field of gene annotation,” says Björn Usadel, director at IBG-4 and professor at the university in Düsseldorf. “Helixer shows that modern AI methods can help overcome this bottleneck.”

An AI learns to read

Helixer works in a similar way to large language models. The system was trained using several hundred previously analysed genomes from a wide range of organisms. The tool thus learnt to recognize typical patterns: Where are genes located? Which sections belong together?

Eine Person mit blauer Jacke vor einem unscharfen Hintergrund.

“What makes it stand out,” says Marie Bolger, a bioinformatician at IBG-4, “is that Helixer does not require any additional experimental data. It works solely on the basis of the DNA sequence – and still achieves an accuracy close to that of labour-intensive reference analyses carried out by humans. In theory, anyone with a sufficiently powerful computer can use the AI tool. We have also created a website where anyone can upload their genome for analysis and our cluster handles the evaluation. The results are then available for download.”

The power of this approach became evident in a test using the model plant Arabidopsis thaliana – the thale cress is the “laboratory mouse” of botany, as its genome is particularly well-researched. Surprisingly, Helixer identified over 100 genes that were missing from existing references or had been incorrectly assigned. “This demonstrates the power of Helixer to improve even the best reference annotations available to date,” says Felicitas Kindel, a doctoral researcher at IBG-4.

A circular structure with internal elements, a pair of connected strands, a twisted double helix, and a series of horizontal lines.
Specific sequence of letters: DNA contains our genetic code, the blueprint for our bodies and all their functions. It is made up of four chemical bases: adenine, guanine, cytosine, and thymine – abbreviated as A, G, C, and T. The order of the bases can be determined by sequencing. Annotation reveals which sequences of bases together form a gene – technically referred to as encoding a gene – and where this gene begins and ends. Non-coding regions can often be found between genes, and their functions are still being investigated.

The next stage – the hybrid model

Helixer still has room for improvement – partly due to human factors: “The system learns from existing gene analyses created by humans or other annotation tools – and in doing so, it also inherits their errors. It’s like learning Italian from a flawed textbook or dealing with an AI that’s hallucinating,” explains Kindel.

Eine Person mit hellem Haar und Brille trägt eine dunkle Jacke über einem hellen Oberteil.

The model also reaches its limits when it comes to finer details. For example, it is tricky to determine the start and end points of a gene if different variants of the gene exist or if it contains blueprints for several proteins. This is where the next stage of development comes in. Kindel is working on combining Helixer with a foundation model – an even more powerful class of AI systems that are not trained for just a single task.

A foundation model of this kind could first learn the overall “language” of DNA and thus develop a fundamental understanding of typical patterns and relationships. “Helixer would have access to a sort of lexicon that would enable it to identify the location of genes more quickly. This would also allow for even more accurate results,” explains Kindel. This would help, for example, to provide rapid initial annotations for the thousands of new genomes decoded each year – particularly as many of these genomes come from species that have hardly been studied to date. Another advantage is that Helixer currently includes models specifically tailored to different groups, such as vertebrates, plants, or fungi. With a foundation model that has learnt universal principles, there could be a single model for all of them.

And yet gene annotation is only the first step. Further potential applications range from breeding more resilient crops to researching genetic diseases. “Anyone looking to develop a drought-resistant plant, for example, needs to understand which genes are responsible for resistance, which proteins they produce, and how these interact within more complex networks. And the more precise the annotation, the more effectively genes can, for example, be targeted using the CRISPR gene-editing tool,” emphasizes Kindel.

This text is taken from the 1/26 issue of effzett. Text: Brigitte Stahl-Busse, Pictures from top to bottom: Björn Usadel, Marie Bolger, Felicitas Kindel, Copyright: FZJ/Ralf-Uwe Limbach

Last Modified: 24.08.2026