Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Dataset . null
Data sources: ZENODO
addClaim

This Research product is the result of merged Research products in OpenAIRE.

You have already added 0 works in your ORCID record related to the merged Research product.

SCDF: A Speaker Characteristics DeepFake Speech Dataset for Bias Analysis

Authors: Staněk, Vojtěch; Srna, Karel; Firc, Anton; Malinka, Kamil;

SCDF: A Speaker Characteristics DeepFake Speech Dataset for Bias Analysis

Abstract

SCDF - A Speaker Characteristics DeepFake Speech Dataset for Bias Analysis is a dataset for fairness assessment and bias analysis of deepfake speech detectors. It contains over 500 hours of (mostly synthetic) speech in over 237,000 utterances. It is based on the VoxPopuli corpus and provides a richly annotated resource with metadata for systematic analysis of how speaker characteristics influence deepfake speech detectors, which allows for fine-grained analysis of the impact of specific attributes. Importantly, this aligns with the ethical AI principles and emerging regulatory requirements, such as the European Union’s AI Act. The dataset contains mostly synthetic speech; 25 bona-fide utterances per speaker are included, and the rest of the utterances are synthesized by one of the tools mentioned below. The SCDF dataset is balanced across: Speakers: balanced representation of speakers (25 male and 25 female) Languages: 5 languages (Czech, French, English, German, Spanish) Synthesizers: 4 state-of-the-art synthesizers (XTTSv2, F5-TTS, Open Voice v2, DDDM-VC) Age: Wide range of speaker ages All audio files are provided in 16-bit PCM WAV format at 16 kHz, consisting of synthetic (deepfake) and genuine speech samples. License: CC0 1.0Audio type: Synthetic (deepfake) and genuine (bonafide)Language: Czech, French, German, English, SpanishPaper: https://arxiv.org/abs/2508.07944

Powered by OpenAIRE graph
Found an issue? Give us feedback