Friday, September 11, 2026
HomeBusiness and IdeasHow to Prevent AI from 'Spouting Nonsense' in a Serious Manner? Research...

How to Prevent AI from ‘Spouting Nonsense’ in a Serious Manner? Research Still Needs This Key Solution

In the Digital Age, Researchers Face an Unprecedented Flood of Information

In the face of a dazzling array of resources, researchers are challenged not only to discern which content is relevant to their work but also to extract valuable information and form unique insights. AI assistants like ChatGPT and DeepSeek excel at summarizing information and responding quickly to queries, but how can we ensure that the information they provide is accurate, high-quality, and not just “nonsense” presented in a serious manner?

Many researchers may have heard of TDM — text and data mining, which uses computational tools and techniques to analyze large text datasets. It extracts valuable information from vast amounts of scientific data in academic papers, journals, and other scientific publications and identifies patterns, correlations, and trends that are difficult or impossible to uncover through traditional manual analysis. From academic research and medical diagnostics to market trend analysis, an increasing number of fields are using TDM to derive actionable insights from complex data sets.

With the Development of Large-Scale Foundation Models and Other Machine Learning and Deep Learning Models

Today, data scientists use corpora of published research to train their own models. These models are not limited to traditional descriptive analysis; they can also provide predictive and prescriptive analysis. For example, Google’s AlphaFold can predict the folding of proteins, fully demonstrating the powerful capabilities of such tools. TDM can help researchers acquire and process information more efficiently. What sparks could be generated if TDM were combined with AI tools? Could the combined value of both be more fully and responsibly realized?

Last year, during a webinar focused on text and data mining (TDM), Dr. Prathik Roy, Head of Data Solutions and Strategy at Springer Nature, discussed advanced technologies combining TDM and AI, showcasing four case studies from different fields. Two of these cases were from the biomedical field, while the other two came from materials science and fintech. These case studies highlight the enormous potential of combining AI tools with TDM, using Springer Nature’s vast publication resources. We hope to inspire researchers, data scientists, and R&D professionals, and provide recommendations on how to integrate TDM into enterprise R&D structures.

How BenevolentAI, an AI Healthcare Unicorn, Uses TDM to Support Drug Discovery

By combining data from multiple sources such as clinical trial reports, patents, journals, books, and medical records, you can obtain over 1 billion associations between genes, symptoms, diseases, proteins, tissues, species, and potential drug candidates.

BenevolentAI is a leading AI drug discovery company. As early as 2018, BenevolentAI partnered with Springer Nature to leverage our TDM tools to access high-quality, rich resources. The company used these datasets to train and build models to identify genes related to specific medical conditions, and based on this, discover effective candidate compounds. During the pandemic, the company also identified potential drug candidates for Covid-19.

How TDM Helps This Large Antibody Search Engine Company

CiteAb is a company specializing in providing a reagent search engine, dedicated to building and training models to extract reagent information from literature.

CiteAb used Springer Nature’s TDM tools to search through the full texts of 60,000 scientific publications, identifying reagent products and how they were used. This information was then converted into structured data to support its search engine. The company then used this data to train AI, continuously enhancing and refining the model in the process, ultimately enabling the rapid identification and extraction of reagent and antibody information from literature. This process is highly automated, utilizing various text mining methods, from simple pattern matching to AI classifiers, and also includes manual review to check for edge cases that algorithms may not handle.

How TDM Supports Semiconductor Design

In materials science, AI-based material data mining has expanded the boundaries of researchers’ capabilities. Initially, researchers used TDM to search for crystal structure data and material properties from comprehensive materials science databases.

The next phase involves using material compositions to predict their structure and properties. Today, AI models can generate statistically-driven material designs through predictive analysis. This means that researchers can design materials with ideal compositions, structures, and properties using physicochemical data, even conducting virtual experiments before actual synthesis and evaluation.

So far, this approach has had a significant impact on semiconductor design, which has driven the development of integrated circuit (IC) and chip design. It has reduced the delay rate in IC design to below 10%, cutting construction times by as much as 10%.

How TDM Helps Fintech

Even financial institutions are interested in literature-based TDM. By using models and TDM technologies to extract information from research corpora, they can better understand and analyze supply chains—especially the chemical industry supply chain. It also helps to understand how the research patterns of R&D companies affect their stock market performance.

Why Do They Choose Springer Nature?

These use cases are all based on Springer Nature’s databases, and Springer Nature’s TDM provides access to the data used by these models.

Springer Nature has developed a formal TDM process and several API tools. Additionally, our extensive publication resources and databases contain a vast amount of solid, validated research.

Our databases adhere to the FAIR principles, designed to ensure users can easily access data:

  • Findable: Well-organized metadata and other input elements make data discoverable in platforms and/or applications.
  • Accessible: Data must be readable and operable by both humans and machines, with the goal of making it publicly accessible.
  • Interoperable: Data is organized using specialized metadata vocabularies to meet various digital lab application scenarios.
  • Reusable: Verified data is directly linked to relevant research elements.

We have also developed several TDM API tools to facilitate text and data mining of our rich publication resources for researchers.

TDM for Open Access Content

Springer Nature Open Access Content API: Provides Springer Nature’s open-access XML-format metadata and full-text content (where available), covering over 649,000 online documents across various disciplines, including journals from BioMed Central and SpringerOpen. We support multiple data output formats, including XML and JSON.

TDM for Subscription Users

For subscription users, Springer Nature offers a variety of TDM data combinations, such as metadata or full-text APIs, applicable to both open access and subscription content.

In addition to well-known journals from the Nature series and Springer Nature link journals and books, Springer Nature also offers specialized databases like SpringerMaterials, AdisInsight, and SpringerProtocols.

The TDM databases can be customized for subscription users, combining different data modules to facilitate searching and usage.

Owogram
Owogram
Welcome to Owogram.com, Your Ultimate Business & Finance Blog. (Business ideas, personal finance, loans, insurance, CRMs, marketing, and more.)
RELATED ARTICLES

Leave A Reply

Please enter your comment!
Please enter your name here

This site uses Akismet to reduce spam. Learn how your comment data is processed.

More