Agents vs Wall Street / point-in-time run / 16 August 2026
Three estimates. One accountable number.
The system reconstructs what the Street believes, independently models the company’s operating reality, and asks prediction markets where real-money evidence disagrees. A source-aware meta-forecaster decides how much each deserves to move the final number.
1,139frozen documentsfilings · calls · slides
12forecast targetsequal-weight metrics
3explicit enginesestimate or abstention
69passing testsdocumented forecast suite
4/4valid workbookssubmission checker
01 / Architecture
Independence is a data contract, not a prompt.
Each top-level engine must return a value, predictive sigma, citations, evidence-family labels and reliability—or explicitly abstain. Driver nowcasts and validation-gated classical ML are nested inside fundamental research, so neither can masquerade as a fourth independent vote.
Engine A
Reconstructed Street
Exact company × metric calls are scored point in time. Historical errors are recency weighted, persistent source bias is corrected, and sparse samples shrink toward a conservative prior.
Inputs: public consensus, named analysts, company-published consensus, dated revisions.
Engine B
Fundamental research
Deterministic extractors turn repeating disclosure shapes into typed observations. Guidance calibration, seasonal models, regression bridges, post-guidance physical drivers and validation-gated classical ML produce the company forecast.
Inputs: filings, transcripts, guidance, peer and macro driver snapshots.
Engine C
Prediction markets
Dated beat markets contribute one quantile, never a fabricated distribution. No equivalent market means abstention. A low precision factor stops a thin binary proxy from overruling the operating model.
Inputs: snapshotted prices, strike, volume and parser self-tests.
Solid = forecast path · Dashed = outcomes resolved by the cutoff, fed back to score each Street source and to replay the whole path on history · Every engine returns value, sigma, citations and reliability, or abstains
Four evidence sources are selected only if dated at or before the cutoff: Street calls, the frozen 1,139-document corpus, physical driver snapshots and a Polymarket price snapshot.
The corpus passes through deterministic per-issuer extractors with unit and magnitude checks, producing 695 typed observations that each carry a period, units, source file and excerpt.
Engine A reconstructs the Street from dated calls, scored against outcomes that had already resolved at the cutoff.
Engine B builds the fundamental forecast from the typed store, with driver nowcasts and validation-gated machine learning nested inside it so neither becomes a second independent vote.
Engine C reads one quantile from a dated market price, and abstains on the nine of twelve targets with no defensible market.
The meta-forecaster weights each engine by reliability and a source-overlap penalty divided by predictive variance, and flags needs_review when engines disagree by more than two combined sigmas.
The run writes four workbooks, a machine-readable audit and a timestamped log.
02 / Controls
The system knows what it does not know.
The base weight is inverse predictive variance. Reliability scales that precision. If two engines reuse the same evidence family—such as a prediction-market strike copied from consensus—both receive an overlap penalty. The complete formula and every realized weight are written to submission/forecast-audit.json.
Corpus documents, Street calls, driver observations and market snapshots are each selected at or before --as-of. Historical source scores only use outcomes whose event date had resolved by the cutoff.
Disagreement boundary
Pairwise engine disagreement beyond two combined sigmas sets needs_review. Missing market coverage remains visible as an abstention warning instead of becoming a copied consensus value.
Unit boundary
Exact labels and units come from challenge/companies.json. Percentages are points, Hays EPS is pence, and non-submittable units hard-fail before a workbook is written.
One output path
The Street backtest is an input to forecast.run, not a second workbook builder. One Python orchestrator produces all four final templates and the machine-readable audit.
03 / Forecasts
Twelve final numbers, with the realized engine mix.
These are descriptive outputs from the 16 August run, not accuracy results; actuals were not available at forecast time. Weights are rounded for display. Full-precision values, sigmas, citations, abstentions and overlap penalties live in the JSON audit. To follow any single number back through the lens that produced it, the stock × method explorer shows each lens's validation history, gate, derivation and consumed sources with publication dates.
Company
Metric
Final
Units
Engine weights
HD
Net sales
47,105.31
USDm
Street 10% · Fundamental 90% · Market abstained
HD
Adjusted diluted EPS
4.64
USD / share
Street 19% · Fundamental 67% · Market 15%
HD
Comparable sales, total company
1.13
%
Street 63% · Fundamental 37% · Market abstained
ADI
Revenue
3,996.71
USDm
Street 4% · Fundamental 96% · Market abstained
ADI
Adjusted diluted EPS
3.43
USD / share
Street 19% · Fundamental 79% · Market 3%
ADI
Adjusted gross margin
73.15
%
Street 1% · Fundamental 99% · Market abstained
Hays
Net fees
903.49
GBPm
Street 1% · Fundamental 99% · Market abstained
Hays
Pre-exceptional basic EPS
1.19
GBp
Street 60% · Fundamental 40% · Market abstained
Hays
Pre-exceptional operating profit
45.39
GBPm
Street 56% · Fundamental 44% · Market abstained
DE
Worldwide net sales and revenues
12,486.77
USDm
Street 9% · Fundamental 91% · Market abstained
DE
Diluted EPS (GAAP)
5.03
USD / share
Street 24% · Fundamental 75% · Market 1%
DE
Production & Precision Ag operating profit
519.09
USDm
Street 32% · Fundamental 68% · Market abstained
04 / Evidence
Metric-specific models, not one generic earnings prompt.
Home Depot
Housing, category demand and ML
Full-year guidance is divided by measured quarterly seasonality. Census NAICS 444 updates comparable sales, while task-matched historical ML contributes only after beating a seasonal naive baseline; three-sample validation receives a wide uncertainty penalty.
Analog Devices
Guidance realization and mix
Revenue and EPS guidance are corrected by walk-forward realization history. Gross margin is bridged from guided operating margin by a fitted relationship, with TXN read-through as a physical driver.
Hays
Closed-year reconstruction
FY fees are H1 actual plus Q3/Q4 actual-basis growth on the prior H2 base, less disposed-country fees. Company-compiled consensus and the top-of-range management steer anchor profit.
Deere
Units → segments → EPS
AEM machinery units, price, FX and a calibrated shipment residual produce PPA sales and profit. Segment revenue and margin bridges then reconstruct worldwide revenue and diluted EPS.
05 / Validation
The system is scored on history, not asserted.
forecast.system_backtest replays the pipeline that writes the twelve submitted numbers. For every closed period it rebuilds the corpus as of the day before the actual was published, then runs the same extractors, calibration, estimators and meta-forecaster the final command runs. Look-ahead is structural: a leak would require a filing to travel backwards in time.
Scoring uses the competition's own formula — error divided by the benchmark's, floored and capped at 5.0 — against a seasonal median year-over-year replay standing in for the Wall Street consensus we are never shown. Below 1.00 beats a model you could write in twenty lines.
Solid = replay path · Dashed = the actual, revealed only after the cutoff · Amber = the feedback loop we measured but deliberately did not close
75 held-out cells · mean score 0.38 · median 0.15 · beat the benchmark on 68
Company · Metric
Held-out cells
Mean score
Beat benchmark
ADI · Revenue
19
0.16
18 / 19
ADI · Adjusted diluted EPS
19
0.16
19 / 19
ADI · Adjusted gross margin
19
0.48
18 / 19
HD · Net sales
5
0.91
3 / 5
HD · Adjusted diluted EPS
5
1.18
3 / 5
HD · Comparable sales, total company
5
0.54
4 / 5
Hays · Net fees
1
0.15
1 / 1
Hays · Pre-exceptional basic EPS
1
0.39
1 / 1
Hays · Pre-exceptional operating profit
1
0.27
1 / 1
Per-lens history, plotted actual against predicted with each gate, is in the forecast explorer. The same final command regenerates it.
Does aggregating actually help? Each engine, scored alone
The replay scores every engine separately, on the cells where it genuinely spoke. Getting the comparison at all took a repair: the Street engine keyed its panel on company and metric with no period, so it could only address the quarter being submitted and abstained on all 75 historical cells. The archive it needed was already in the repository — 39 consensus rows in forecasting/data/historical_forecasts.csv, each stamped a day before the company reported. Keying them by period gave Street 27 cells and made the blend measurable.
On the 27 cells where both engines spoke
Mean score
Median score
Beat benchmark
Fundamental engine alone
0.413
0.132
23 / 27
Street engine alone
0.480
0.153
24 / 27
Meta-forecaster, both blended
0.444
0.124
23 / 27
The reading is mixed. The blend has the best median and beats Street on 20 of 27, so it rescues the case where one engine is badly wrong. But its mean is worse than the fundamental engine alone, which it beats on only 11 of 27. Blending dilutes the stronger engine about as often as it protects against the weaker one.
The cause is identifiable. The fundamental engine's reliability is the fixed constant 0.85 rather than its measured skill, so the meta-forecaster cannot see which engine this evidence favours. Fitting reliability to measured skill is the change this backtest argues for; it is listed as a limitation rather than done quietly and claimed as design.
The prediction-market lens took two passes. A settled market prices the answer, not a forecast: the closed books in the live snapshot sit at 0.999 against 0.001, status resolved, so reading one at a later cutoff would have scored beautifully and been pure look-ahead. parse_beat_market now requires a market to still be open at the cutoff, making that refusal a rule rather than an accident of which snapshots exist. The legitimate measurement lives elsewhere: the CLOB price series, where the last trade at least 24 hours before the scheduled print is genuinely pre-print. forecast.polymarket_history scores all nine resolved beat markets that way — the whole HD, ADI and Deere series, which only begins in November 2025, and Hays has none.
The verdict is a clean negative: Brier 0.100 the day before the print, against 0.0988 for simply predicting the base rate. Across nine markets the crowd added nothing over knowing that companies usually beat. That is the measured justification for the market engine's 0.03 reliability, which until now rested only on reasoning about what a binary quantile can mean.
What the replay refuses to score
60 further cells are recorded as engine failures rather than dropped: each is the system declining to forecast on inputs it did not have. Deere abstains entirely, its AEM driver chain being anchored to a dated snapshot rather than a fiscal period, so no historical cutoff carries what it consumes. Hays scores one period because its engine needs a company-compiled consensus the corpus holds only for recent years.
The backtest also exposed the Home Depot guidance gap. HD does not always guide a range, and the extractor matched only the “X to Y” wording, so guidance existed for FY2026 alone. Reading single-point guidance as a fallback — never an override, so a published range always wins — took HD from 3 scored cells to 15 with all four workbooks byte-identical.
Two further backtests, both in the final command
forecast.backtest tests one claim: that correcting management guidance for its historical bias beats parroting it. Calibration is fitted only on the corpus as it stood the day before each guidance was issued. Across 54 held-out periods it improved 42, with error ratios of 0.65, 0.74 and 0.77.
forecasting.backtest ranks Street sources rather than the system, scoring 39 closed observations by recency-weighted error and shrinking sparse histories toward an 8% prior. This is what makes Street reliability measured rather than chosen.
The ML lens, and the gate it has to clear to speak
Classical models are an estimator inside the fundamental engine, never a fourth top-level vote. Validation is chronological only—a random split on a time series lets a model read its own future—and model choice is made by TimeSeriesSplit within the training window, never on the test period. Gates were fixed from what would be competitive against a Street benchmark before any model ran. A prediction is written only if it clears its gate and beats a seasonal naive baseline; sparse validation additionally widens sigma, so three good historical transitions cannot outvote current evidence. Seven of twelve metrics cleared. Five abstain.
Metric
Deployed model
Walk-forward
Gate
n
Verdict
HD · Net sales
Share-of-category nowcast, zero parameters, blended with a CV-weighted voting ensemble
0.40%
2%
6
pass
HD · Adjusted diluted EPS
Ridge on sequential ratios, plus the observed adjusted-to-GAAP wedge
4.17%
5%
3
pass
HD · Comparable sales
Ridge on lag 1 and lag 4
0.40pp
0.8pp
3
pass
ADI · Revenue
Guidance realization—guide × expanding mean of past realization ratios, zero parameters
1.38%
2%
14
pass
ADI · Adjusted diluted EPS
Guidance realization
3.44%
5%
14
pass
ADI · Adjusted gross margin
Guided operating margin × realization, plus a modelled opex spread
0.89pp
1.0pp
14
pass
Hays · Net fees
Quarterly composition from four disclosed growth rates, zero parameters
1.52%
2%
4
pass
Hays · Operating profit, EPS
None—seven annual observations
—
10%
7
abstain
DE · Revenue, EPS, PPA profit
Voting ensemble over Ridge, Random Forest, Gradient Boosting, SVR and KNN
7.3–24.3%
2–10%
4–9
abstain
Model inventory
Ridge, Random Forest, Gradient Boosting, SVR with an RBF kernel and KNN, plus equal-weight voting, CV-weighted voting and ridge stacking. Features are scaled; COVID-era target rows are excluded as a documented exogenous shock rather than silently winsorized.
Intrinsic floor
Before blaming a model we measure the dispersion of the target transition’s own history. Deere’s Q2→Q3 ratio disperses 8.1% on revenue and 37.2% on PPA profit against gates of 2% and 10%, so no history-only estimator could pass. Given four times the samples the model scored 7.26%—landing on its predicted floor. That is an information limit, not a tuning failure.
What we abandoned
A year-over-year framing scored 5.40% against a 5.24% naive baseline and was cut. Ridge stacking was the worst ensemble at 6.04%, overfitting a meta-learner on twenty rows. Ensembling bought 0.2pp; reframing the target around already-published data bought an order of magnitude.
A bug the gates did not catch
Hays tags every quarterly update with the fiscal year rather than the quarter, so keying on period collapsed a twenty-point series into six and we concluded Hays was unmodelable. The loader now raises on a colliding key; an audit found the same fault in neither ADI nor Deere. Hays net fees became a passing metric at 1.52%.
Every figure on this page is re-derived from the run artefacts by npm run check:architecture, which fails if the page and the audit disagree.
06 / Explorer
Follow any single number back to its evidence.
Pick a company and a lens. The explorer plots that lens's walk-forward history — reported actual against model prediction, period by period — with its pre-declared gate, the derivation of the submitted figure, and the source documents it consumed with publication dates. A lens that failed its gate still shows its history: that history is the evidence for the abstention.
Embedded in full below, so this page stays one self-contained file. npm run forecast rebuilds it from the run artefacts.
07 / Reproduce
One final command, then structural validation.
python3 -m venv .venv
.venv/bin/python -m pip install -r requirements.txt
npm ci
npm run test:forecast
npm run forecast
npm run check:forecasts
git rev-parse HEAD
npm run forecast regenerates the gated Home Depot ML artifact, materializes the Street backtest, runs the walk-forward skill test and the system replay, and invokes the final orchestrator—all at the frozen cutoff. The runner writes four workbooks, a timestamped log, submission/forecast-audit.json, and the backtest evidence under research/. Upload remains manual by challenge rule.
Known limits
The public exact-metric analyst panel is sparse for several operating metrics. Those Street reconstructions use a wider-sigma research fallback and lower reliability.
Direct prediction markets exist only for three EPS targets. A binary beat price supplies one quantile, so its precision is deliberately discounted; all other targets abstain.
The Home Depot ML gate is task matched but has only three historical Q2 transitions. Its uncertainty is widened before it enters the fundamental engine, and the all-transition stress view remains visible in the artifact.
Regex extractors depend on recurring issuer disclosure shapes. Magnitude checks and source excerpts turn format drift into an observable failure, but do not eliminate it.
Source-overlap labels are declared families, not learned causal dependence. The penalty is conservative and auditable, but still an approximation.
The system backtest is unbalanced. ADI supplies 57 of the 75 scored cells, so the headline mean is largely an ADI result; Hays contributes 3 and Deere none. The per-metric table is the honest reading, not the single average.
Engine reliabilities are not yet learned from the backtest. Street reliability is derived from measured source error, but the fundamental engine's 0.85 is a chosen constant even though it carries most of the weight. Feeding measured skill back into that number is the clearest next improvement and was out of scope for the day.
The replay benchmark is a seasonal baseline, not the Wall Street consensus the accuracy prize scores against. Beating it is necessary evidence, not sufficient.
Blending is measurable but not yet vindicated. On the 27 cells where two engines speak the meta-forecast has the best median and beats Street on 20, but its mean is worse than the fundamental engine alone and it beats that engine on only 11. On this evidence a judge would be right to ask whether the blend earns its place, and the answer today is that it earns it against a weak engine and costs against a strong one.
Engine reliability is chosen, not learned, and that is the direct cause of the result above. The fundamental engine sits at a fixed 0.85 while Street is derived from measured source error, so the meta-forecaster cannot see which engine the backtest says is better. Fitting reliability to measured skill is the clearest next change and was deliberately not attempted late in the day, because it would move submitted numbers with no time to validate the move.
The prediction-market record is nine markets long. The beat series only begins in November 2025 and Hays has none at all, so the Brier result is directionally clear but statistically thin.
The blinded LLM replay holds only two genuinely post-training quarters per metric. Three metrics that pass overall fail on that post-training subset, which is the honest place to look for memorisation, and percent metrics stay unscaled so recall on them is measured rather than prevented.