References
class scale and system complexity Domain Feature Definition Source Timing Predictive rationale Baseline Facility topology Buildings halls power trains cooling trains Design basis Sanction Separates repeated structure and redundancy Baseline Delivery model Design-bid- build design- build EPCM or hybrid Contract records Sanction Captures governance and interface allocation Baseline Design maturity Percent packages approved at baseline Document control Baseline Measures readiness behind planned start Baseline Utility scope Owner utility or shared responsibilit y Interconnectio n plan Sanction Separates external gate exposure Schedule Logic density Relationship s per active activity Schedule snapshot Each update Flags underlinked networks Schedule Open-end rate Open-ended activities per 100 activities Schedule snapshot Each update Signals incomplete dependency logic Schedule Constraint density Hard constraints per 100 activities Schedule snapshot Each update Detects date forcing Schedule Near-critical mass Remaining activities under float threshold Schedule snapshot Each update Represents path convergence Schedule Forecast movement Days moved since prior snapshot Schedule archive Each update Measures deterioration or recovery Schedule Logic churn Added or deleted links per period Schedule archive Each update Captures plan instability Schedule Remaining- duration growth Net increase among in- progress tasks Schedule archive Each update Detects optimistic status correction Schedule Out-of-sequence Activities Schedule Each Measures Domain Feature Definition Source Timing Predictive rationale rate progressing against logic snapshot update execution-plan mismatch Schedule Critical path turnover Change in controlling path family Schedule archive Each update Signals instability Design Release reliability Packages released on or before plan Document control Each update Measures downstream readiness Design Revision intensity Revisions per active package Document control Each update Captures design volatility Design RFI aging Median and tail age of open RFIs RFI system Weekly Represents unresolved field information Design Submittal cycle time Days from submission to approved state Submittal log Weekly Measures approval queue Design First-pass approval Share approved without resubmissio n Submittal log Weekly Measures package quality Design Interface issue age Age of cross-system unresolved issues Issue log Weekly Signals integration risk Procurement PO release variance Actual versus planned award date Procurement system Event Measures start of lead-time chain Procurement Engineering approval Evidence state for vendor drawings Procurement documents Weekly Distinguishes promise from readiness Procurement Manufacturing stage Verified stage percentage or state Vendor report Weekly Tracks physical progress Procurement Factory-test readiness Prerequisite completenes s Quality system Weekly Predicts test date credibility Procurement Factory-test Pass fail and Quality system Event Updates Domain Feature Definition Source Timing Predictive rationale outcome open findings shipment risk Procurement Logistics status Booked departed customs and delivered Logistics system Event Tracks remaining transitions Procurement Promise volatility Count and magnitude of vendor date changes Procurement archive Each update Measures commitment instability Procurement Shared-vendor exposure Concurrent packages using supplier Portfolio data Monthly Captures common capacity risk Production Installed quantity rate Accepted quantity per crew-hour Field system Weekly Measures realized productivity Production Plan reliability Completed commitment s divided by planned Look-ahead plan Weekly Measures workflow control Production Constraint removal Constraints cleared before work date Constraint log Weekly Measures readiness discipline Production Crew stability Turnover and reassignmen t by work package Labor system Weekly Captures learning disruption Production Congestion Active crews or permits per zone Field records Daily Represents workspace interference Production Material readiness Work packages with full kit available Material system Weekly Separates labor from supply constraint Quality Inspection pass rate Accepted inspections divided by attempts Quality system Weekly Measures hidden completion Quality Nonconformanc e rate Open NCRs per active work Quality system Weekly Predicts rework demand Domain Feature Definition Source Timing Predictive rationale package Quality Rework hours Hours charged to corrective work Labor system Weekly Measures productivity loss Quality Closure age Median and 90th percentile defect age Quality system Weekly Captures unresolved backlog Quality Repeat defect rate Recurring defect families per system Quality system Weekly Signals systemic failure Commissionin g Prerequisite completeness Required documents and checks complete Cx platform Weekly Measures test readiness Commissionin g Script approval Approved scripts divided by due scripts Cx platform Weekly Captures preparation maturity Commissionin g Controls point verification Verified points divided by planned Controls platform Weekly Measures integration readiness Commissionin g First-test pass rate Tests passed on first execution Cx platform Weekly Predicts retest demand Commissionin g Severe open defects High- severity defects by system Cx platform Daily Directly threatens acceptance Commissionin g Retest cycle time Days from failure to passed retest Cx platform Weekly Measures defect recovery Commissionin g Witness capacity Available witness hours versus demand Resource plan Weekly Captures test queue constraint Commissionin g Turnover completeness Accepted records divided by required Document system Weekly Links physical work to acceptance External Utility evidence Completed Utility tracker Weekly Measures Domain Feature Definition Source Timing Predictive rationale state prerequisites for energization external milestone maturity External Permit comment age Age of unresolved authority comments Permit tracker Weekly Predicts approval slippage External Weather exposure Observed and forecast lost-work conditions Weather source Daily Represents exogenous productivity risk External Inspection availability Booked authority slots versus need Inspection plan Weekly Captures external queue Portfolio Shared specialist load Demand versus capacity by specialist role Portfolio resource plan Weekly Measures cross- project competition Portfolio Shared equipment exposure Projects dependent on common equipment family Portfolio procurement Monthly Captures common-mode supply risk Portfolio Utility concentration Capacity milestones under same provider Portfolio utility tracker Monthly Captures correlated external risk Outcome Committed milestone date Date active at index time Governance register Each index Defines contemporaneou s commitment Outcome Actual milestone date Verified accepted completion Acceptance record Outcom e Defines realized event Outcome Service readiness Operational acceptance state Owner acceptance Outcom e Primary business- relevant completion Governance Data freshness Age of newest source record Pipeline metadata Each run Flags stale predictions Governance Override Human Forecast Each Supports Domain Feature Definition Source Timing Predictive rationale magnitude adjustment to model forecast ledger decision human-model evaluation Governance Intervention completion Action executed after alert Action tracker Each alert Measures response pathway 5 Modeling strategy The model comparison begins with governance baselines: the current schedule forecast, no- change from prior forecast, historical median variance for a reference class, and earned-schedule projection. These baselines establish whether machine learning adds value. A complex model that fails to outperform the current forecast at the same horizon should not enter production. Statistical baselines include regularized logistic and linear regression. Their coefficients, shrinkage, and probability outputs support interpretation and calibration. Survival models estimate time to milestone while retaining incomplete projects. Mixed-effects models account for repeated snapshots and campus clustering. Generalized additive models allow nonlinear effects while retaining graphical interpretability. General guidance on building and tuning such baselines without overfitting follows standard applied predictive-modeling practice (Kuhn and Johnson, 2013). Interpretability requirements are not uniform across use cases and should not be treated as a single design choice. A forecast that informs internal prioritization can tolerate a less transparent model provided its calibration is monitored, but a forecast that influences vendor escalation, contractual notice, or executive commitment communication needs a model whose drivers can be explained to a party outside the data-science team without appeal to feature- importance scores alone. Regularized regression and mixed-effects models satisfy the second requirement directly through their coefficients; tree-based and temporal candidates satisfy it only through post hoc explanation methods that Section 7 requires to be validated against process- expert judgment before their output is treated as a driver rather than a correlate. Tree-based candidates include random forest and gradient boosting. They handle nonlinear interactions, mixed scales, and missingness patterns common in project data. Hyperparameters must be tuned within nested training folds. Class weighting or focal objectives should be evaluated for rare delay thresholds, but probability calibration must be checked after any imbalance treatment. Repeated-hall structure has a specific consequence for hyperparameter tuning that generic tabular-modeling guidance does not address. A gradient-boosted model tuned by ordinary cross-validation can inadvertently reward splits that key on hall identity within a single campus, since repeated halls sharing one design and one vendor set produce highly similar feature vectors with correlated outcomes. Nested tuning folds in this protocol must therefore be constructed at the project level, not the hall or snapshot level, so that a hyperparameter setting cannot be selected because it fits the idiosyncrasies of halls the final test set will later reuse under a different label. Temporal candidates use sequences of snapshot features. Recurrent networks, temporal convolution, or transformer-based models might capture acceleration and deterioration patterns. Their data demand is high. Sequence padding, irregular intervals, and changing project scope require explicit treatment. Simpler lagged features may perform similarly with lower governance cost. Graph-aware models use schedule and dependency structure. Message-passing models can represent propagation along precedence links or shared resources. Their advantage should be tested through ablation against flat models with critical-path and centrality features. Graph changes over time create versioning complexity, and explanations must identify which nodes and edges influence the forecast. An ensemble may combine models with distinct error patterns. Stacking weights should be learned only from validation predictions, not the test set. The final system should favor parsimony when performance differences are small. Model selection should consider calibration, lead time, stability, inference latency, maintainability, and actionability alongside discrimination. Table 2 summarizes the candidate models compared under this strategy and the controls each requires. A worked comparison illustrates how the candidate models would be expected to diverge. Consider a campus where a generator vendor's engineering-approval evidence stalls two weeks behind its promised date while the vendor's stated ship date remains unchanged. The current schedule forecast, anchored to the vendor's promise, would show no movement. The historical reference class would flag elevated risk only once enough comparable campuses existed to estimate a base rate for this equipment family. A regularized regression using procurement evidence features (Table 1) should detect the engineering-approval lag immediately, since it is an explicit predictor, while a gradient-boosted model with access to the full feature set should additionally weigh the lag against concurrent schedule-health and near-critical-mass signals to judge whether downstream float can absorb the delay. This scenario is the type of case the ablation and ensemble analyses in this section are designed to surface and score, not a claim about which model performs best in general. Table 2. Candidate model comparison Model Purpose Key controls Current schedule forecast Operational benchmark Same index date and milestone definition Historical reference class Outside-view benchmark Class frozen before testing Regularized regression Interpretable nonlinear baseline through transforms Nested tuning and calibrated probabilities Mixed-effects model Repeated snapshots and project clustering Random effects estimated without test leakage Survival model Incomplete projects and time to milestone Time-varying covariates frozen at index Random forest Nonlinear tabular benchmark Project-level folds and calibration Gradient boosting High-performance tabular candidate Early stopping inside training folds Temporal sequence model Trajectory learning Rolling-origin validation Graph-aware model Precedence and shared- resource propagation Time-versioned nodes and edges Stacked ensemble Combine distinct errors Weights learned from validation predictions only 6 Validation and statistical analysis The primary split holds out entire later projects. Training uses earlier projects, validation supports tuning and threshold selection, and a locked test set estimates final performance. A second leave-one-project-out analysis measures sensitivity to individual campuses. A geographic or delivery-partner holdout should test transfer where sample size permits. Random row splits may appear only as a demonstration of optimistic bias. Rolling-origin evaluation recreates the live forecasting process. At each origin, the pipeline trains on information previously available and predicts the next eligible projects or snapshots. Feature definitions, imputers, encoders, selection, and calibration must be fitted inside the origin. This prevents subtle leakage from future distributions. Regression evaluation includes mean absolute error, median absolute error, root mean squared error, signed bias, and error quantiles. Interval forecasts require empirical coverage, width, and conditional coverage across horizons. Classification evaluation includes precision-recall area, receiver-operating area, sensitivity, specificity, positive predictive value, Brier score, log loss, and calibration slope and intercept. Interval width should be interpreted against the decision it supports, not treated as a single universal target. A fourteen-day prediction interval around a design-freeze date eighteen months from service readiness may be entirely adequate for capital planning, while the same fourteen-day interval around an integrated-systems-testing date four weeks out would leave a response team with almost no usable lead time. This protocol therefore requires interval coverage and width to be reported by horizon, consistent with Table 3, so that a narrowing interval as service readiness approaches can be distinguished from a model that is simply miscalibrated at every horizon alike. Performance should be reported by forecast horizon, project phase, facility type, region, delivery model, and target prevalence. Bootstrap confidence intervals must resample projects rather than rows. Pairwise model differences should use the same held-out predictions. Statistical significance does not replace operational relevance; a predeclared minimum improvement should govern selection. Decision evaluation assigns costs to missed alerts, false alerts, and interventions. Net-benefit curves or an explicit cost matrix compare thresholds. Lead time measures days between the first sustained alert and the milestone. Alert burden measures alerts per project-month. A useful model should improve response opportunity without flooding teams. Cost assignment should be elicited from the same project- controls teams who would act on the alerts, not set analytically by the modeling team alone. The cost of a false alert on an energization-risk indicator, an unnecessary vendor executive call, a wasted contingency-planning cycle, is small relative to the cost of a missed alert that allows a campus to reach a scheduled integrated-systems-test window without adequate witness capacity or defect closure. Where such elicitation is not feasible before initial deployment, this protocol requires the evaluation to report net-benefit curves across a plausible range of cost ratios rather than committing to a single assumed ratio, so that a project-controls team can locate its own judgment on the curve after the fact. Ablation analysis removes each feature family to estimate incremental predictive contribution. Stability analysis repeats feature attribution across folds, horizons, and seeds. Error analysis reviews false negatives, false positives, large residuals, and unsupported cases. Reviewers should classify whether errors arise from missing data, definition failure, unprecedented events, model weakness, or human override. Table 3 summarizes the evaluation framework, mapping each dimension to an acceptance question a project-controls team can apply directly. Subgroup reporting should reflect the specific ways hyperscale delivery varies rather than generic demographic categories. Relevant strata include redundancy topology (N, N+1, 2N), cooling technology (air-cooled, direct liquid, hybrid), delivery model (design- build, design-bid-build, EPCM), utility ownership of the interconnection, and project generation (first-of-a-design versus repeated hall of an established design). A model that performs well on average but poorly for first-of-a-design halls or for projects where the owner rather than a utility controls interconnection would misinform the teams facing exactly the conditions where forecasting is hardest and most valuable. Table 3. Evaluation framework Dimension Measures Acceptance question Accuracy MAE median AE RMSE signed bias Does the model reduce date error against baselines? Uncertainty Interval coverage width and conditional coverage Do stated ranges contain outcomes at the promised rate? Discrimination Precision-recall ROC sensitivity specificity PPV Does the model separate milestone misses at useful thresholds? Calibration Brier score slope intercept reliability plot Do reported probabilities match observed frequencies? Decision value Net benefit intervention cost avoided loss Does use improve the decision under stated costs? Timeliness First alert lead time and sustained alert lead time Does warning arrive before response options close? Stability Fold horizon seed and attribution stability Does the signal persist across reasonable choices? Equity and transfer Performance by region partner type and project generation Where does the model fail to transfer? Operations Latency missingness alert burden override rate Does the service work inside reporting cadence? 7 Explainability governance and deployment Global explanation uses coefficient paths, permutation importance, grouped feature importance, partial dependence, and SHAP summaries as appropriate. Correlated features require grouped interpretation because individual rankings become unstable. Local explanations show the evidence behind a specific alert. Neither form proves causation. Process experts should assess whether the pattern is plausible and actionable. Model-agnostic explanation methods and their proper use follow established treatments of explanatory model analysis and interpretable machine learning (Biecek and Burzykowski, 2021; Molnar, 2022). A model card should document target, horizon, intended use, excluded use, training period, project population, algorithms, validation, calibration, subgroup results, limitations, and monitoring. A data sheet should document sources, definitions, timestamps, lineage, missingness, and access. Each prediction should link model version, feature snapshot, output, explanation, threshold, user response, and outcome. Deployment should start in silent mode, then advisory mode. In advisory mode, the model supplies probability, interval, driver groups, data freshness, and supported actions. It does not alter the contractual schedule automatically. Users who override the output should record new evidence, direction, magnitude, and rationale. Later analysis compares model-only, human-only, and combined forecasts. A concrete override case illustrates why structured adjustment, not silent human correction, is required. Suppose the model raises the probability of an energization miss because switchgear factory-test evidence has stalled, while a project executive knows, from a call the vendor has not yet logged in any system, that the test has actually passed and documentation is simply pending. Recording only a revised forecast would erase the evidence the model correctly used; recording the original forecast, the adjusted forecast, the stated reason, and the date the vendor's system evidence caught up allows later analysis to confirm whether the executive's early information was reliable in general or accurate only in this instance. Over enough cases, this record is what converts anecdote about human judgment into an evaluable claim. Monitoring covers data completeness, latency, feature drift, probability calibration, error, subgroup performance, alert rate, override patterns, and intervention outcomes. A stop rule should suppress predictions when required sources fail, the case lies outside training support, or calibration breaches a threshold. Recalibration may precede full retraining when ranking remains stable. Security and privacy controls should apply least privilege, encrypted transfer and storage, approved retention, and auditable access. Commercially sensitive vendor performance should have defined uses. Workforce data should remain aggregated when individual identity is unnecessary. Governance should prevent the system from becoming a punitive score that encourages data manipulation. Vendor-performance sensitivity deserves specific mention because it differs from typical construction-analytics privacy concerns, which focus mainly on workforce data. A manufacturing-stage or factory-test-outcome feature effectively scores a specific equipment supplier's delivery reliability. Sharing that score outside the immediate project-controls function, particularly with parties negotiating future contracts with the same vendor, raises commercial and legal considerations distinct from workforce privacy. Governance should therefore define, separately from workforce-data rules, which roles may see vendor- identified procurement features and which may see only de-identified, aggregated versions. 8 Expected contribution limitations and research agenda The protocol advances the field by treating prediction as a longitudinal, leakage-controlled decision service. Its contribution lies in the design, not unobserved results. Once executed, the study will show whether historical schedule evidence transfers across hyperscale projects, which feature families add value at each horizon, and whether calibrated alerts improve decisions beyond current controls. That framing has a practical consequence for how the protocol should be read by a reviewer or an adopting organization. It does not promise that any specific algorithm, gradient boosting, a graph neural network, or a survival model, will prove superior; it promises that whichever algorithm is chosen will have been evaluated under conditions that make the result trustworthy, on unseen projects, against honest baselines, with calibration and lead time reported alongside accuracy. A protocol that guaranteed a particular model's superiority in advance of execution would itself be an example of the presenting-conjecture-as- measured-performance problem this paper is designed to prevent. The main limitation is data availability. Hyperscale project histories are commercially sensitive and may use inconsistent coding. Small project counts constrain deep learning even when activity rows number in the millions. Organizational changes create concept drift. Actual service dates may be influenced by commercial decisions beyond construction. These limits require careful target definition and restrained generalization. A related limitation concerns the boundary between construction and commissioning records, which is drawn differently across organizations. Some owners treat commissioning as a subcontracted scope with its own schedule and data systems, entirely separate from the general contractor's integrated master schedule, while others embed commissioning activities directly inside the same schedule file used for construction sequencing. The feature dictionary in Table 1 assumes commissioning evidence is available at the same snapshot cadence as construction evidence; where it is not, an owner adopting this protocol should expect an early phase of aligning commissioning data governance with construction data governance before commissioning-family features can be populated reliably, and should treat any interim model that lacks them as provisional rather than final. Future work should examine causal interventions after predictive validation. Questions include whether earlier vendor escalation reduces delay, whether added commissioning capacity improves service readiness, and whether acceleration increases defect-driven variance. Predictive importance should nominate hypotheses, while causal designs evaluate actions. A further limitation deserves explicit statement: this protocol assumes that organizations already retain schedule snapshots, procurement events, and commissioning records in a form that supports as-of reconstruction. Many owners and contractors currently overwrite schedule updates or retain only the latest procurement status, which means the first deliverable of any execution of this protocol may be a data-retention change rather than a model. Organizations beginning from this position should treat Level 1 of the maturity progression, establishing controlled snapshots and objective progress measurement, as a prerequisite phase with its own timeline, not as a step that can be compressed to reach model development sooner. Federated analysis may support cross-owner learning without pooling raw data. Shared feature definitions, secure aggregation, and external validation would test transfer while protecting confidentiality. A public benchmark using de-identified or synthetic schedule networks would improve reproducibility. Any synthetic set must preserve temporal and dependency properties without being represented as actual performance. A federated design is attractive in this sector specifically because no single owner, however large, delivers enough campuses to satisfy the project-level holdout standard this protocol requires with high confidence. An owner with twenty completed campuses can estimate a reference class, but cannot reliably estimate rare, high- consequence events, a first-of-kind cooling technology failing integrated testing, for example, from twenty observations alone. Secure aggregation across owners, contractors, and commissioning firms would pool exactly these rare events without requiring any party to disclose commercially sensitive vendor performance or contract terms, provided the shared feature definitions in Table 1 are adopted consistently enough to make pooled records comparable. 9 Conclusion Historical schedule performance data support credible delay prediction only when records preserve what was known at each forecast date. The proposed protocol defines project-snapshot observations, operational outcomes, as-of feature construction, project-level temporal validation, strong baselines, uncertainty reporting, explainability, decision thresholds, and deployment monitoring. The decisive controls are methodological. Projects, not activity rows, must remain separated across training and test sets. Later information must not enter earlier forecasts. Performance must include calibration, intervals, horizon, subgroup behavior, and decision value. Advanced models must outperform current schedules and simple historical baselines. None of these controls is unique to machine learning. A purely statistical baseline, an earned- schedule projection, or a human forecaster is equally subject to project-level holdout logic, temporal integrity, and calibration reporting once the comparison is made explicit. The protocol's contribution is to make that comparison unavoidable: any candidate, however produced, enters the same evaluation pipeline and is judged against the same current-forecast and historical- median baselines under the same project-level split. A hyperscale owner adopting this protocol should therefore expect its first concrete output to be a disciplined benchmark of existing practice, not necessarily a new algorithm. The paper reports no synthetic performance claim. Its next step is execution on an authorized multi-project dataset, followed by silent prospective validation. Only measured results should support claims of accuracy, generalizability, or schedule improvement. This protocol is the second paper in a planned three-part program. The first paper in this series established why deterministic schedule assurance reaches its limit on hyperscale AI data-center programs and proposed a layered predictive architecture as a conceptual response; this paper operationalizes the data, feature, and validation requirements that architecture presupposes but did not itself specify in executable form. A subsequent paper is expected to develop and validate the resulting framework against the propositions and acceptance criteria defined here, closing the sequence from conceptual architecture through empirical protocol to tested framework.