Proteins and Natural Language: AI Enables the Design of Novel Proteins

Artificial intelligence (AI) has created new possibilities for designing tailor-made proteins to solve everything from medical to ecological problems. A research team at the University of Bayreuth led by Prof. Dr. Birte Höcker has now successfully applied a computer-based natural language processing model to protein research. Completely independently, the ProtGPT2 model designs new proteins that are capable of stable folding and could take over defined functions in larger molecular contexts. The model and its potential are detailed scientifically in Nature Communications.

Natural languages and proteins are actually similar in structure. Amino acids arrange themselves in a multitude of combinations to form structures that have specific functions in the living organism - similar to the way words form sentences in different combinations that express certain facts. In recent years, numerous approaches have therefore been developed to use principles and processes that control the computer-assisted processing of natural language in protein research. "Natural language processing has made extraordinary progress thanks to new AI technologies. Today, models of language processing enable machines not only to understand meaningful sentences but also to generate them themselves. Such a model was the starting point of our research. With detailed information concerning about 50 million sequences of natural proteins, my colleague Noelia Ferruz trained the model and enabled it to generate protein sequences on its own. It now understands the language of proteins and can use it creatively. We have found that these creative designs follow the basic principles of natural proteins," says Prof. Dr. Birte Höcker, Head of the Protein Design Group at the University of Bayreuth.

The language processing model transferred to protein evolution is called "ProtGPT2". It can now be used to design proteins that adopt stable structures through folding and are permanently functional in this state. In addition, the Bayreuth biochemists have found out, through complex investigations, that the model can even create proteins that do not occur in nature and have possibly never existed in the history of evolution. These findings shed light on the immeasurable world of possible proteins and open a door to designing them in novel and unexplored ways. There is a further advantage: Most proteins that have been designed de novo so far have idealised structures. Before such structures can have a potential application, they usually must pass through an elaborate functionalization process - for example by inserting extensions and cavities - so that they can interact with their environment and take on precisely defined functions in larger system contexts. ProtGPT2, on the other hand, generates proteins that have such differentiated structures innately, and are thus already operational in their respective environments.

"Our new model is another impressive demonstration of the systemic affinity of protein design and natural language processing. Artificial intelligence opens up highly interesting and promising possibilities to use methods of language processing for the production of customised proteins. At the University of Bayreuth, we hope to contribute in this way to developing innovative solutions for biomedical, pharmaceutical, and ecological problems," says Prof. Dr. Birte Höcker.

Ferruz N, Schmidt S, Höcker B.
ProtGPT2 is a deep unsupervised language model for protein design.
Nat Commun 13, 4348, 2022. doi: 10.1038/s41467-022-32007-7

Most Popular Now

ChatGPT Shows 'Impressive' Acc…

A new study led by investigators from Mass General Brigham has found that ChatGPT was about 72 percent accurate in overall clinical decision making, from coming up with possible diagnoses...

WiFi SPARK's Healthcare Business Re…

Leading WiFi provider WiFi SPARK is rebranding its healthcare arm as SPARK Technology Services Limited. The new identity marks the completion of the integration of the former Hospedia bedside unit...

Online AI-Based Test for Parkinson'…

An artificial intelligence (AI) tool developed by researchers at the University of Rochester can help people with Parkinson's disease remotely assess the severity of their symptoms within minutes. A study...

ChatGPT is Debunking Myths on Social Med…

ChatGPT could help to increase vaccine uptake by debunking myths around jab safety, say the authors of a study published in the peer-reviewed journal Human Vaccines and Immunotherapeutics. The researchers asked...

AI Performs Comparably to Human Readers …

Using a standardized assessment, researchers in the UK compared the performance of a commercially available artificial intelligence (AI) algorithm with human readers of screening mammograms. Results of their findings were...

Siemens Healthineers Expands Production …

Siemens Healthineers is expanding its site in Rudolstadt, Germany. By mid 2024, a new manufacturing building will be built on the site. The new manufacturing plant will produce electron accelerators...

More Cases of Breast Cancer Detected wit…

One radiologist supported by AI detected more cases of breast cancer in screening mammography than two radiologists working together, reports the ScreenTrustCAD study from Karolinska Institutet in The Lancet Digital...

ChatGPT Performs as Well as Doctors for …

The artificial intelligence chatbot ChatGPT performed as well as a trained doctor in suggesting likely diagnoses for patients being assessed in emergency medicine departments, in a pilot study to be...

Smartphone Technology Expected to Advanc…

Since the 1980s, we have known that neurological soft signs (NSS) can distinguish people with schizophrenia from psychiatrically healthy individuals. NSS are subtle neurological impairments that principally manifest as decreased...

AI may Outperform Most Humans at Creativ…

Large language model (LLM) artificial intelligence (AI) chatbots may be able to outperform the average human at a creative thinking task where the participant devises alternative uses for everyday objects...

MEDICA 2023 + COMPAMED 2023: "Where…

13 - 16 November 2023, Düsseldorf, Germany. The medical technology market is in worldwide motion and the signs ahead of MEDICA 2023 and COMPAMED 2023 in Düsseldorf as the internationally leading...

AI and Machine Learning can Successfully…

Artificial intelligence (AI) and machine learning (ML) can effectively detect and diagnose Polycystic Ovary Syndrome (PCOS), which is the most common hormone disorder among women, typically between ages 15 and...