
We present a comprehensive empirical study examining the scalability characteristics of Latent Posterior Factor (LPF) models across varying evidence pool sizes. Through systematic experiments spanning pool sizes from 10 to 500 evidence pieces per entity, we demonstrate that both LPF-SPN and LPF-Learned architectures maintain robust performance while exhibiting distinct scaling behaviours. Our findings reveal that LPF-SPN achieves superior calibration (ECE = 0.050–0.163) with computational efficiency (14–15 ms inference), while LPF-Learned attains near-perfect accuracy (98.5–100%) at the cost of increased latency (35–42 ms). Notably, performance remains stable across a 50× increase in evidence volume, validating the architectural design for real-world deployment scenarios where knowledge bases accumulate evidence over time. Keywords:Scalability analysis, Latent Posterior Factors (LPF), evidence aggregation, multi-evidence reasoning, large-scale AI systems, probabilistic reasoning, neural-symbolic AI, sum-product networks (SPN), learned aggregation, uncertainty calibration, expected calibration error (ECE), high-throughput inference, model efficiency, performance scaling, knowledge base systems, machine learning scalability, real-world deployment AI
large-scale AI systems, neural-symbolic AI, real-world deployment AI, knowledge base systems, learned aggregation, Scalability analysis, sum-product networks (SPN), evidence aggregation, model efficiency, Latent Posterior Factors (LPF), expected calibration error (ECE), machine learning scalability, uncertainty calibration, high-throughput inference, performance scaling, probabilistic reasoning, multi-evidence reasoning
large-scale AI systems, neural-symbolic AI, real-world deployment AI, knowledge base systems, learned aggregation, Scalability analysis, sum-product networks (SPN), evidence aggregation, model efficiency, Latent Posterior Factors (LPF), expected calibration error (ECE), machine learning scalability, uncertainty calibration, high-throughput inference, performance scaling, probabilistic reasoning, multi-evidence reasoning
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
