Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ Estudo Geralarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
Estudo Geral
Master thesis . 2024
Data sources: Estudo Geral
addClaim

Bias Correction in Datasets

Authors: Artur, João Filipe Guiomar;

Bias Correction in Datasets

Abstract

O aumento do recurso a técnicas de aprendizagem computacional para apoiar cenários de tomada de decisões com elevado impacto social, nomeadamente em contextos financeiros e criminais, tem suscitado preocupações relacionadas com a equidade das decisões. Em virtude destas preocupações, verificou-se a emergência do domínio da Aprendizagem Computacional Justa que visa dar resposta a estas preocupações e garantir que o emprego de técnicas de aprendizagem computacional não promove decisões discriminatórias com base em atributos protegidos, como a raça ou o género, sobretudo devido à existência de conjuntos de dados de treino enviesados. Esta tese analisa a literatura deste domínio e, em particular, para a mitigação de enviesamentos em conjuntos de dados tabulares, identificando a predominância de abordagens para problemas de classificação binária com apenas um atributo sensível binário e outras limitações. Em função dos desafios identificados, é proposta uma nova ferramenta, concebida para automatizar a redução de enviesamentos em dados tabulares. O núcleo desta estrutura é o FairGenes, um algoritmo evolutivo que aproveita as técnicas de redução de enviesamento existentes para resolver problemas com atributos sensíveis não binários. O FairGenes evolui, a partir de um conjunto de algoritmos, uma sequência de transformações a aplicar aos dados. Para encontrar a melhor solução para um determinado classificador alvo, o FairGenes utiliza um conjunto de classificadores substitutos para avaliar e melhorar a equidade das previsões. Os resultados experimentais mostram ser possível atingir o objetivo desejado de equidade nasprevisõesdeumclassificadoralvoatravésdautilizaçãodeclassificadoressubstitutos. As comparações com os métodos existentes indicam que a ferramenta proposta tem um desempenho semelhante ao de outras abordagens atuais para problemas que envolvem caraterísticas binárias protegidas, sugerindo direções com potencial para futuras melhorias. Futuro trabalho centrar-se-á na expansão do conjunto de algoritmos, no refinamento da escolha de métricas e classif icadores e na avaliação do desempenho da framework para outros conjuntos de dados e tipos de problemas.

The increasing reliance on machine learning to support decision-making scenarios with high social impact, particularly in financial and in criminal contexts, has raised concerns about the fairness of such decisions. The field of Fair Machine Learning has emerged to provide a response to these concerns and to ensure that the employment of machine learning techniques does not promote discriminatory decisions with respect to sensitive attributes, such as race or gender, in particular due to the existence of biased training data. This thesis examines the stateof-art for Fair Machine Learning in the form of bias mitigation in tabular datasets, identifying the predominance of approaches for binary classification problems with only one binary sensitive attribute and other limitations. To address these challenges, a novel framework is proposed, designed to automate bias reduction from tabular data. The core of this framework is FairGenes, an evolutionary algorithm that leverages existing bias reduction techniques to solve problems with non-binary sensitive attributes. FairGenes evolves a pipeline of algorithms that encodes a sequence of transformations to are to be applied to the data. To find the best solution for a given target classifier, FairGenes uses a pool of surrogate classifiers to evaluate and improve the fairness of the predictions. The experimental results show that it is possible to achieve the desired goal of fairness in the predictions of a target classifier through the use of surrogate classifiers. Comparisons with existing methods indicate that the proposed framework performs similarly to other current approaches for problems involving binary protected features, suggesting promising directions for further improvement. Future work will focus onexpandingthepoolofalgorithms, refiningthechoiceoffairnessmetrics and classifiers, and evaluating the performance of the framework for other datasets and types of problems.

Outro - CRAI: This workwassupportedbythePortuguese Recovery and Resilience Plan (PRR) through project C645008882-00000055, Center for Responsible AI, by the FCT, I.P./MCTES through national funds (PIDDAC).

Dissertação de Mestrado em Engenharia Informática apresentada à Faculdade de Ciências e Tecnologia

Country
Portugal
Related Organizations
Keywords

fair machine learning, bias mitigation, ferramenta justa, pre-processing, fairness framework, Fair Genes, aprendizagem computacional justa, mitigação de enviesamentos, pré-processamento

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average
Green