
Overview EquiLens CorpusGen is a reproducible, balanced corpus generator for bias auditing in Small Language Models (SLMs). This release includes all code, configuration, and documentation needed to generate and validate audit corpora for gender bias and other demographic comparisons. What's Included File Name Description generate_corpus.py Main corpus generator script test_config.py Strict configuration validator word_lists.json Corpus configuration (edit for your audit) word_lists_schema.json JSON schema for config validation corpus/audit_corpus_.csv Generated corpus (after running the script) README.md Full usage, setup, and documentation CITATION.cff Citation file LICENSE.md Apache 2.0 license package_for_zenodo.ps1 Packaging script for Zenodo or other archives How to Use See README.md for setup and usage instructions (Windows PowerShell and Linux/macOS Bash supported). Validate your configuration with test_config.py. Generate your corpus with generate_corpus.py. Package your release with package_for_zenodo.ps1 (Windows) or manually (Linux/macOS). Citation See CITATION.cff or use the following BibTeX: @misc{equilens2025, author = {Krishna GSVV}, title = {EquiLens Corpus Generator: A Framework for Reproducible Bias Auditing in Small Language Models}, year = {2025}, publisher = {GitHub}, journal = {GitHub repository}, howpublished = {\\url{https://github.com/Life-Experimentalists/EquiLens}}, license = {Apache-2.0} } License Apache License 2.0 (see LICENSE.md) 5. Additional Metadata (optional) DOI: https://doi.org/10.5281/zenodo.17014104
The EquiLens Corpus Generator is an open-source tool for creating balanced, reproducible datasets for bias auditing in Small Language Models (SLMs). This release (v1.0.0-corpus) includes the first stable version with controlled generation of names × professions × traits × templates, designed for auditing gender bias and extendable to other demographic categories.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
