doi: 10.5281/zenodo.21809656
Research article: Batch Inference Scheduling: Maximizing GPU Utilization for Cost-Effective Enterprise AI