References
level tested testing records above reference level; hypothesis- generating, not confirmed in cattle Each variable in Table 3 requires, at minimum, a recorded source, timestamp, permissible range, and missing-value code before it enters the analysis dataset, and repeated variables require a defined measurement window rather than a single value carried forward. Breed composition and genomic estimated breeding value are the two variables least likely to change within a heifer's breeding season and are therefore recorded once and applied as time-invariant effect modifiers; body condition score, growth rate, forage dry matter intake, and thermal exposure are recorded repeatedly because their trajectories, not only their values at a single timepoint, carry predictive information. Where genomic data are unavailable for a given heifer, breed composition from pedigree is used as a lower-resolution substitute rather than as a missing-data placeholder, and the model architecture comparison in Section 4 is repeated separately for the genotyped and non- genotyped subpopulations to test whether genomic resolution changes which architecture performs best. 5.1 Reliability and quality control A predictor is only as useful as its reliability allows. Body condition score assigned by visual appraisal is subject to inter-observer disagreement that widens with the number of scorers involved across a multiherd study; the protocol therefore requires either a single trained scorer per herd-visit, periodic inter-observer calibration sessions with a reported agreement statistic, or adoption of the image-based automated scoring methods documented in Section 2, which remove between-observer variability at the cost of requiring calibration against visual scores at least once per herd. Temperature-humidity index derived from a single fixed weather station may not represent the microclimate a given heifer actually experiences if shade, airflow, or stocking density vary within a pen; on-animal or in-pen sensors, where feasible, should be prioritized over station-level estimates for the subset of herds able to support them, with the resulting discrepancy between station and in-pen measurement reported as a methods finding in its own right rather than discarded. Genomic estimated breeding values carry their own reliability statistic, reported by the genomic evaluation provider, that should be retained alongside the point estimate and used to downweight low-reliability genomic values in the mixed-effects model rather than treating all genomic estimated breeding values as equally precise regardless of the reference population size behind them. Laboratory or diagnostic assays used for pregnancy confirmation should report their own sensitivity and specificity at the gestational age of testing, since early transrectal palpation and early ultrasonography carry materially different false- negative rates, and a difference in diagnostic method across herds that is not recorded as a covariate could otherwise masquerade as a difference in true pregnancy rate. Reliability reporting also extends to the contaminant-burden indicator in the environmental exposure domain. Because water-source testing is neither universal nor performed on a fixed schedule, the study should record the test date, analyte panel, and detection limit for every herd that contributes this variable, and should treat a herd with an old or partial test differently from one with a recent, full-panel test rather than collapsing both into a single binary tested-or-not indicator. A stale or partial test can understate current exposure in a way that biases hypothesis H6 toward the null, which is a further reason H6 is designated exploratory rather than confirmatory in the sequencing set out below. Laboratory batch and technician effects, well documented as sources of variance in bovine in vitro production work, apply equally to on-farm diagnostic and scoring procedures and should be captured as random effects in the primary model wherever more than one technician or diagnostic team contributes records within a herd. Where a single technician effect is confounded with a single herd, as is common on smaller operations, this should be reported as a limitation on the interpretability of any residual herd-level variance component rather than silently absorbed into the herd random intercept. 6 Testable hypotheses Six hypotheses follow directly from the physiological foundations in Section 2 and the framework in Section 3. Each is stated with a direction, the comparison that would support or refute it, and the minimum evidence a prospective study should require before treating the hypothesis as supported. H1 (condition by heat interaction). The association between integrated temperature-humidity index and pregnancy success is more strongly negative among heifers in the lowest body- condition-score tertile than among heifers in the highest tertile. Support requires a statistically and practically meaningful interaction coefficient in the mixed-effects model that survives herd- held-out validation, not merely a difference in raw pregnancy rate across unadjusted subgroups. H2 (genomic merit by nutrition interaction). Genomic estimated breeding value for fertility predicts pregnancy success more strongly among heifers above herd-median forage dry matter intake than among heifers below it. Support requires the interaction term, not only the genomic EBV main effect, to add predictive value in held-out herds. H3 (breed composition by thermal load). Proportion Bos indicus ancestry is associated with higher pregnancy success specifically in herd-year-season strata with integrated temperature- humidity index above the study median, with no meaningful association, or a reversed association favoring conventional genetics, below that median. This hypothesis treats breed composition as conditionally rather than universally advantageous. H4 (measurement resolution). A model incorporating genomic estimated breeding value and a genomic relationship matrix achieves better calibration in an unseen herd than an otherwise identical model using pedigree-based breed composition alone, among genotyped heifers. H5 (architecture transfer gap). The performance gap between development-herd and held-out- herd evaluation is larger for gradient-boosted trees than for the mixed-effects baseline, reflecting greater susceptibility of flexible learners to herd-specific management signal that does not transfer. H6 (contaminant burden). Among herds with water-source testing on record, categorical contaminant-burden status is associated with pregnancy success after adjustment for temperature-humidity index and body condition score. This hypothesis is explicitly exploratory given the absence of bovine-specific dose-response data and is retained to generate, not confirm, an effect estimate. For each hypothesis, a null or reversed result is treated as informative. A failed interaction term indicates either that the proposed physiological mechanism does not operate at the measured resolution, that the measurement itself is too coarse to detect it, or that a confound absorbed the effect; distinguishing among these explanations requires the process- and implementation-level evidence specified in Sections 4 and 5, not a redefinition of pregnancy success after the fact. 6.1 Sequencing of hypothesis tests The six hypotheses are not of equal priority and should not be tested as an undifferentiated set. H1 carries the strongest a priori physiological support from the neuroendocrine convergence described in Section 2 and should be treated as the primary confirmatory test, with a single preregistered significance threshold and no adjustment for the other five hypotheses. H2 and H3 are secondary confirmatory tests, each with a preregistered direction but tested at a threshold adjusted for the two comparisons; both depend on genotyped or breed-composition data of variable availability across herds and should be reported with the achieved sample size for the relevant subpopulation stated explicitly rather than implied by the overall study sample size. H4 and H5 are methodological rather than biological hypotheses and are appropriately evaluated descriptively, by direct comparison of calibration and transfer performance across architectures, without a formal significance threshold. H6 is exploratory by design, given the absence of bovine-specific dose-response evidence for contaminant exposure, and any effect estimate it produces should be reported with explicit language identifying it as hypothesis-generating, not confirmatory, regardless of the p-value obtained. This sequencing exists so that a null result on an exploratory hypothesis does not retrospectively discredit the confirmatory tests, and so that a positive result on H1 is not diluted by simultaneous testing of five other comparisons of lower prior probability. 7 From model to on-farm decision support A validated prediction model changes practice only if its output reaches a decision at the right time. The proposed pathway moves in four stages. Definition and readiness establish agreed variable definitions, data sources, and the ethical and data-governance approvals required before any heifer-level record leaves the farm management system. A pilot phase tests data capture and model scoring on a small number of herds without acting on the output, so that data quality problems are surfaced before they affect management decisions. Protected evaluation applies the held-out validation design in Section 4 on a larger multiherd sample, with the model used only in an advisory capacity alongside standard practice. Only after calibration and transfer performance meet prespecified thresholds does the model move to decision support, where its output informs, but does not replace, breeding and culling decisions made by farm management and attending veterinarians. Throughout this pathway, a heifer flagged as low-probability for pregnancy success should trigger a defined management response, such as targeted nutritional supplementation, delayed rather than accelerated culling pending a repeat assessment, or, where thermal load is the dominant driver, provision of shade or cooling rather than a genetic or nutritional intervention that would not address the operative mechanism. A model that reports risk without a corresponding action pathway converts a measurement exercise into an administrative burden rather than a management tool. The decision value of the prediction model depends on the asymmetry between the two error types it can make. A false-negative heifer, incorrectly predicted low-probability and consequently culled or withheld from further breeding investment, forfeits her full remaining productive value. A false-positive heifer, incorrectly predicted high-probability and consequently held in the breeding program without the corrective management a true low-probability animal would have received, extends the herd's non- productive interval and delays realization of the same information at the next diagnosis. These costs are not symmetric and are not constant across herds, since replacement heifer cost, feed cost, and the opportunity cost of a delayed diagnosis vary by operation and by season. Rather than reporting a single accuracy figure, the protected-evaluation and decision-support stages should report net benefit across a plausible range of decision thresholds, so that a given herd can select the threshold appropriate to its own cost structure instead of inheriting a default threshold optimized for the average development-herd cost ratio. None of the four stages is meaningful without capacity building among the staff who record the underlying data. Body condition scoring, respiration-rate spot checks, and water-trough audits are only as reliable as the training of the person performing them, and a multiherd study that trains scorers once at enrollment without periodic refresher calibration should expect measurement drift over the course of a multiyear study. The readiness stage should therefore include a documented training protocol for each manually recorded variable, with a scheduled recalibration interval, rather than treating training as a one-time event that precedes, rather than continues alongside, data collection. Governance of the decision-support stage should specify who is authorized to act on a model output, how a heifer's record can be corrected if a variable was misrecorded, and how the model's classification is documented alongside the eventual outcome so that miscalibration can be detected in ongoing use rather than only at the next formal validation cycle. Version control should tie each classification to the specific model version and the predictor values active at the time of scoring, so that a later audit can reconstruct why a given heifer received a given risk classification. 8 Discussion The physiological literature reviewed in Section 2 converges on a single structural point: nutritional status and thermal load act on the same upstream endocrine pathway, so a prediction model that enters them as additive covariates is not a simplification of the biology but a misspecification of it. The practical implication is that herds correcting only one limiting factor, for example investing in shade structures while allowing body condition to drift below the herd- specific optimum, should not expect the full pregnancy-rate benefit that either intervention would produce in isolation, because the two factors are expected to interact rather than to sum. The genetic literature carries a second, less intuitive implication. Genomic selection and breed- composition choices that raise average fertility merit or average thermotolerance do not remove the value of nutritional and thermal management; the comparative embryo-transfer literature indicates that the pregnancy-rate advantage associated with heat-adapted genetics is itself conditional on the thermal environment in which it is measured (Eberhardt et al., 2009), which is exactly the conditional structure proposed for hypothesis H3. A herd that adopts heat-tolerant genetics without also addressing body condition and forage quality should not expect the genetic gain to be fully expressed. Methodologically, the comparison across mixed-effects, tree-based, and genomic-enhanced architectures in Section 4 is included because no single architecture is expected to dominate across all use cases. A mixed-effects model with prespecified interactions is the more defensible choice where interpretability and physiological grounding matter, such as when the output must be explained to a farm manager or attending veterinarian, while a tree- based ensemble may recover thresholds, such as a temperature-humidity index breakpoint, that the linear model would need to be told to look for. The herd-held-out and season-held-out validation design exists specifically to prevent either architecture from being credited with accuracy that reflects development-herd management idiosyncrasy rather than a transferable predictor-outcome relationship. The inclusion of water-source contaminant burden as an exploratory covariate is deliberately conservative. The supporting evidence is drawn from human and environmental-health literatures rather than cattle-specific trials (Oyelade & Ohanebo, 2024; Oyelade & Ohanebo, 2025), and hypothesis H6 is framed accordingly as hypothesis-generating. Its inclusion nonetheless follows directly from the definition of environmental exposure used throughout this paper: an exposure that is present in a herd's water supply and independently linked to endocrine-relevant developmental outcomes in other species should not be excluded from a livestock reproduction model on the grounds that a bovine dose-response study does not yet exist. Absence of species-specific evidence is a reason to test the association, not a reason to omit the variable. The predictor taxonomy in Table 1 is also written to anticipate a shift already underway in commercial dairy management, in which continuous sensor streams, including rumination and activity collars, in-line milk and rumen sensors, and automated imaging for body condition, are replacing the periodic manual observations on which most existing fertility prediction work has relied. A framework specified only in terms of periodic manual measurement risks becoming obsolete as soon as continuous data become available, while a framework specified in terms of the underlying physiological construct, such as energy balance or integrated thermal load, remains valid regardless of whether the construct is measured by a monthly visual body condition score or a continuous automated proxy. The variable-level measurement matrix in Table 3 is deliberately written at the level of the construct and its currently practical measurement method rather than assuming any particular sensor technology, so that a herd with access to continuous monitoring and a herd relying on periodic manual scoring can both contribute data to the same underlying model, with measurement method itself retained as a covariate rather than treated as an unimportant implementation detail. Access to genomic testing is not uniform across the herds to which this framework would ideally apply. The genomic-enhanced architecture in Table 2 and hypothesis H4 both depend on genotyping, which remains cost-prohibitive for many smallholder and pastoral operations, including systems that rely most heavily on Bos indicus-influenced genetics for thermotolerance and that therefore stand to benefit most from a validated breed-composition-by-heat-load interaction. A framework that reports genomic-enhanced performance without also validating the pedigree-only architecture on the same held-out herds would produce a tool usable only by the subset of operations that can already afford genotyping, reproducing an access gradient rather than narrowing it. The stratified genotyped and non-genotyped validation specified in Section 4 exists specifically to prevent this outcome and to report honestly where the pedigree-only model falls short. The proposed protocol also differs from common current practice in a way worth stating plainly. Many on-farm decisions aids compute a single composite fertility score from a weighted sum of available variables, chosen for computational convenience rather than derived from the interaction structure argued for here. Such a composite is not wrong so much as underspecified: it can achieve reasonable average performance while systematically misclassifying the heifers at the tails of the condition, heat-load, or genomic-merit distributions, precisely because a weighted sum cannot represent a multiplicative interaction unless the interaction term is entered explicitly. The comparative architecture design in Section 4 is constructed to make this difference visible rather than assumed, by testing whether the mixed- effects model with explicit interaction terms outperforms both a simpler additive baseline and the more flexible tree-based learners on calibration in held-out herds, not only on raw discrimination in the development sample. The framework's limitations follow from its status as a design rather than a completed study. No accuracy, calibration, or effect-size estimate is reported here because none has been produced under this protocol; every quantitative claim in Sections 3 through 6 is stated as a hypothesis with a specified test, not as a result. The value of the paper lies in specifying that test precisely enough that a future study, using authorized herd data and appropriate institutional animal-care approval, can be judged against a preregistered plan rather than against outcomes chosen after the data were seen. 9 Conclusion Predicting pregnancy success in dairy heifers from body condition, forage dry matter intake, environmental exposure, and genetic background requires treating these four domains as a single, interacting predictor architecture rather than as four independent covariates entered additively into a regression. This paper has specified that architecture: a physiological rationale for two a priori interaction terms, a herd-clustered and temporally held-out validation design compared across four candidate modeling architectures, a variable-level measurement matrix with defined sources and timing, and six falsifiable hypotheses. The central claim is that no single domain, whether nutritional management, cooling infrastructure, or genomic selection, is sufficient on its own to maximize heifer pregnancy rates once any of the other three domains is limiting. Future empirical work should apply this protocol to authorized, multiherd data under appropriate institutional animal-care oversight, report calibration and transfer performance alongside discrimination, and treat a failed interaction hypothesis as informative rather than as a reason to abandon the interaction-based framing. If H1 and H2 are confirmed under held-out validation, the practical consequence for herd management is that nutritional, cooling, and genetic investments should be sequenced and co-targeted at the individual heifer level rather than applied as uniform herd-wide interventions, since the framework predicts that the marginal value of any one investment depends on the heifer's standing in the other domains. If they are not confirmed, the framework's more conservative but still useful contribution is the variable-level measurement matrix and validation design itself, which would allow the field to move toward a properly specified additive model with the same rigor this paper has applied to the interaction hypothesis, rather than defaulting back to an unvalidated composite score. Declarations Ethics approval and consent to participate: This manuscript is a predictive-modeling framework and protocol paper; no animal observations were collected for it. A future empirical study implementing this protocol must obtain institutional animal care and use committee approval for all animal handling, sampling, and management procedures, and must obtain the consent of participating herd owners for use of farm records, before any data are collected. Consent for publication: Not applicable. Availability of data and materials: No new animal, herd, or laboratory dataset was generated or analyzed for this paper. The predictor taxonomy, modeling comparison protocol, and measurement specifications required to reproduce the proposed study are reported in full within the manuscript. Competing interests: The authors declare that they have no competing interests.