Accessibility notice:

If you need help accessing this archived item, Ask a Librarian.

Theory-Constrained Inductive Bias for High-Stakes Tabular Classification: Enabling Ante-Hoc Governance in Small-Data Decision

Loading...
Thumbnail Image

Advisors

License

DOI

Type

Working Paper

Journal Title

Journal ISSN

Volume Title

Publisher

Grantor

Abstract

Small and medium enterprises (SMEs) face a persistent and underappreciated paradox: they require high-stakes predictive analytics for workforce decisions and yet lack the data volumes ([Equation]) that conventional machine learning approaches demand in order to generalise reliably. While black-box models such as Random Forest and XGBoost dominate performance benchmarks in big-data contexts, their weak inductive biases cause overfitting and produce uninterpretable predictions in low-data organisational settings. Drawing on Design Science Research (Hevner et al., 2004; Peffers et al., 2007), this study instantiates a theory-constrained decision intelligence agent that encodes the Learning Transfer System Inventory (LTSI; Holton et al., 2000) as a strong inductive bias, thereby constraining the hypothesis space to theoretically plausible relationships between LTSI constructs and workforce outcomes. The instantiation compares theory-guided logistic regression against black-box alternatives using five-fold stratified cross-validation on synthetic data ([Equation] records, representing 60 analysts across 5 areas of practice) generated from LTSI-specified theoretical relationships. This constitutes a proof-of-concept evaluation in which all three model architectures achieve comparable discriminative performance (AUC approximately 0.62), yet the theory-constrained model achieves substantially lower cross-validation variance (CV SD = 0.022 versus 0.034 to 0.061 for tree-based alternatives), representing a 1.5 to 2.8 times reduction in prediction instability across evaluation folds. Governance stability, rather than marginal accuracy gains, is the appropriate evaluation criterion for high-stakes SME workforce decisions where managerial accountability is non-negotiable. Model coefficients serve as governance artifacts that are reviewable by managers prior to deployment, which is the defining property of ante-hoc interpretability as articulated by Rudin (2019). A human-in-the-loop Decision Queue with immutable audit logging ensures that managers validate all AI-generated intervention plans before any action reaches an employee. The live implementation is publicly accessible and fully reproducible from the published source code. The broader contribution is a replicable DSR instantiation demonstrating that organisational theory can function as statistical inductive bias, enabling responsible and auditable AI governance in SME contexts where data are scarce and accountability is non-negotiable.

Description

Keywords

Related Material and Data

Citation

Sponsorship

Endorsement

Review

Supplemented By

Referenced By