Skip to main content

Write a PREreview

A Literature Knowledge Graph of Thermoelectric Materials Data Extracted from 10,158 Papers

Posted
Server
Preprints.org
DOI
10.20944/preprints202609.1858.v1

Decades of thermoelectric-materials research are recorded in tens of thousands of papers, but the quantitative results — Seebeck coefficients, conductivities, figures of merit, doping strategies, synthesis conditions — remain locked in unstructured prose, inaccessible to systematic analysis or machine learning. We present a literature knowledge graph that extracts and structures this information at scale. Using large-language-model information extraction over the parsed full text of 10,158 thermoelectric papers, the graph captures 166,178 records across 13 entity types: 22,317 material mentions, 24,121 property measurements, 7,443 dopant/modification records, 32,306 scientific claims, 14,304 experimental conditions, and derived research-coverage, novelty, and failure signals, each attributed to its source paper by DOI and aligned to an OWL ontology. An independent audit on a stratified 300-paper sample gives an extraction precision of 0.97 and a recall of ~0.85 on content-bearing papers. The graph is released in five interoperable formats (CSV, JSON, JSON-LD, RDF/Turtle with 1.11 million triples, and GraphML with 166,178 nodes and 156,020 edges), enabling literature-scale trend analysis, research-gap discovery, and the fusion of literature evidence with computational materials features.

You can write a PREreview of A Literature Knowledge Graph of Thermoelectric Materials Data Extracted from 10,158 Papers. A PREreview is a review of a preprint and can vary from a few sentences to a lengthy report, similar to a journal-organized peer-review report.

Before you start

We will ask you to log in with your ORCID iD. If you don’t have an iD, you can create one.

What is an ORCID iD?

An ORCID iD is a unique identifier that distinguishes you from everyone with the same or similar name.

Start now