Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
versions View all 2 versions
addClaim

SDS Toolbox: End-to-End SDS Retrieval and Structured Data Extraction Using LLMs

Authors: Ali, Asmaa; Edelweiss Connect (Switzerland);

SDS Toolbox: End-to-End SDS Retrieval and Structured Data Extraction Using LLMs

Abstract

SDS Toolbox is an automated system designed to streamline the process of retrieving and extracting structured data from Safety Data Sheets (SDS) using chemical identifiers such as CAS numbers or IUPAC names. It combines intelligent search capabilities with language model–powered data extraction in a modular architecture. Modules The toolbox consists of three core modules: 1. SDS-FIND Function: Searches for SDS files online using a CAS number or IUPAC name. Sources: PDF files are retrieved from trusted sources indexed by search engines. Search Engines Supported: Currently uses SerpAPI (Google) or Brave Search. 2. SDS-STRUCT Function: Parses and extracts structured data from SDS PDF files. Technology: Utilizes LLMs to convert unstructured PDF content into structured formats (e.g., CSV, JSON). Data Output: Extracted data includes chemical identifiers, hazard statements, manufacturer info, and safety classifications. 3. SDS-FLOW Function: An end-to-end pipeline that combines search (SDS-FIND) and extraction (SDS-STRUCT) into a single seamless process. Input: CAS number or IUPAC name. Output: Archived SDS files + extracted structured data. Features: Automated retries Logging Zip output with downloadable results Email delivery (optional) Key Features Supports both single requests and batch processing Modular and extensible architecture Works with both local PDFs and online searches Fully integrated with Streamlit UI Optional email delivery of results

Related Organizations
Keywords

Automation, Data Collection, safety assessment, safety data sheets, large language models, Material Safety Data Sheets

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average