Accessibility notice:
If you need help accessing this archived item, Ask a Librarian.
Theory-Constrained Inductive Bias for High-Stakes Tabular Classification: Enabling Ante-Hoc Governance in Small-Data Decision
| dc.contributor.author | Zubia Mughal | |
| dc.date.accessioned | 2026-07-29T20:44:08Z | |
| dc.date.issued | 2026-03 | |
| dc.description.abstract | Small and medium enterprises (SMEs) face a persistent and underappreciated paradox: they require high-stakes predictive analytics for workforce decisions and yet lack the data volumes ([Equation]) that conventional machine learning approaches demand in order to generalise reliably. While black-box models such as Random Forest and XGBoost dominate performance benchmarks in big-data contexts, their weak inductive biases cause overfitting and produce uninterpretable predictions in low-data organisational settings. Drawing on Design Science Research (Hevner et al., 2004; Peffers et al., 2007), this study instantiates a theory-constrained decision intelligence agent that encodes the Learning Transfer System Inventory (LTSI; Holton et al., 2000) as a strong inductive bias, thereby constraining the hypothesis space to theoretically plausible relationships between LTSI constructs and workforce outcomes. The instantiation compares theory-guided logistic regression against black-box alternatives using five-fold stratified cross-validation on synthetic data ([Equation] records, representing 60 analysts across 5 areas of practice) generated from LTSI-specified theoretical relationships. This constitutes a proof-of-concept evaluation in which all three model architectures achieve comparable discriminative performance (AUC approximately 0.62), yet the theory-constrained model achieves substantially lower cross-validation variance (CV SD = 0.022 versus 0.034 to 0.061 for tree-based alternatives), representing a 1.5 to 2.8 times reduction in prediction instability across evaluation folds. Governance stability, rather than marginal accuracy gains, is the appropriate evaluation criterion for high-stakes SME workforce decisions where managerial accountability is non-negotiable. Model coefficients serve as governance artifacts that are reviewable by managers prior to deployment, which is the defining property of ante-hoc interpretability as articulated by Rudin (2019). A human-in-the-loop Decision Queue with immutable audit logging ensures that managers validate all AI-generated intervention plans before any action reaches an employee. The live implementation is publicly accessible and fully reproducible from the published source code. The broader contribution is a replicable DSR instantiation demonstrating that organisational theory can function as statistical inductive bias, enabling responsible and auditable AI governance in SME contexts where data are scarce and accountability is non-negotiable. | |
| dc.identifier.uri | https://digital.library.wisc.edu/1793/97749 | |
| dc.title | Theory-Constrained Inductive Bias for High-Stakes Tabular Classification: Enabling Ante-Hoc Governance in Small-Data Decision | |
| dc.type | Working Paper |
Files
Original bundle
showing 1 - 1 results
Loading...
- Name:
- Mughal_2026_Theory_Constrained_Inductive_Bias_Zubia Mughal.pdf
- Size:
- 729.34 KB
- Format:
- Adobe Portable Document Format
License bundle
showing 1 - 1 results
Loading...
- Name:
- license.txt
- Size:
- 2.59 KB
- Format:
- Item-specific license agreed upon to submission
- Description: