Agents vs Wall Street / point-in-time run / 16 August 2026

Three estimates.
One accountable number.

The system reconstructs what the Street believes, independently models the company’s operating reality, and asks prediction markets where real-money evidence disagrees. A source-aware meta-forecaster decides how much each deserves to move the final number.

1,139frozen documentsfilings · calls · slides
12forecast targetsequal-weight metrics
3explicit enginesestimate or abstention
69passing testsdocumented forecast suite
4/4valid workbookssubmission checker
01 / Architecture

Independence is a data contract, not a prompt.

Each top-level engine must return a value, predictive sigma, citations, evidence-family labels and reliability—or explicitly abstain. Driver nowcasts and validation-gated classical ML are nested inside fundamental research, so neither can masquerade as a fourth independent vote.

Engine A

Reconstructed Street

Exact company × metric calls are scored point in time. Historical errors are recency weighted, persistent source bias is corrected, and sparse samples shrink toward a conservative prior.

Inputs: public consensus, named analysts, company-published consensus, dated revisions.

Engine B

Fundamental research

Deterministic extractors turn repeating disclosure shapes into typed observations. Guidance calibration, seasonal models, regression bridges, post-guidance physical drivers and validation-gated classical ML produce the company forecast.

Inputs: filings, transcripts, guidance, peer and macro driver snapshots.

Engine C

Prediction markets

Dated beat markets contribute one quantile, never a fabricated distribution. No equivalent market means abstention. A low precision factor stops a thin binary proxy from overruling the operating model.

Inputs: snapshotted prices, strike, volume and parser self-tests.

Three-Lens Meta-Forecaster data flow Evidence selected at the point-in-time cutoff feeds deterministic extractors and three forecasting engines, which are combined by a source-overlap-aware meta-forecaster into four workbooks, a JSON audit and a run log. The same path is replayed at 75 historical cutoffs and scored against the figures those companies actually reported. 01 · EVIDENCE AT CUTOFF 02 · EXTRACTION 03 · ENGINES 04 · META 05 · OUTPUT Street calls Public and company-compiled consensus, each dated Frozen corpus 1,139 filings, transcript sections and slide decks Physical drivers Census NAICS 444, AEM units, FRED, peer read-through Market snapshot Polymarket prices, strike and volume, stored by date CUTOFF --AS-OF 2026-08-16 DETERMINISTIC Extractors Per-issuer table and regex shapes; unit and magnitude 695 TYPED OBSERVATIONS ENGINE A Reconstructed Street Point-in-time source scoring, bias correction, shrinkage ENGINE B Fundamental research Guidance calibration, seasonal shares, regression bridges DRIVER NOWCASTS Dated physical inputs VALIDATION- GATED ML Must beat the seasonal naive NESTED — NEITHER CASTS A SECOND VOTE Both draw on the same typed store. ENGINE C Prediction markets One quantile from a dated beat price — never a fabricated curve 9 OF 12 TARGETS: NO MARKET → ABSTAIN META-FORECASTER One accountable number w = reliability × overlap ÷ σ² Inverse-variance base weight, scaled by declared reliability and a source-overlap penalty. needs_review when engines differ by more than 2σ. Four workbooks SUBMISSION/*.XLSX Machine audit FORECAST-AUDIT.JSON Narrative trail LOGS/RUN-<TS>.LOG POINT-IN-TIME STREET CALLS DATED DRIVER SNAPSHOTS DATED MARKET PRICES 06 · VALIDATION — THE SAME PATH, REPLAYED ON QUARTERS THAT ALREADY REPORTED forecast.system_backtest Rewind the cutoff to the day before a closed quarter was reported, run everything above unchanged, and score the result against the figure the company actually published — competition formula, seasonal benchmark. 75 HELD-OUT CELLS · MEAN 0.38 · MEDIAN 0.15 · BEAT THE BENCHMARK ON 68 · 60 CELLS REFUSED · DEERE ABSTAINS MEASURED SKILL SHOULD SET ENGINE RELIABILITY NOT WIRED — RELIABILITY IS STILL A CONSTANT See 05 / Validation for the per-engine breakdown
Solid = forecast path · Dashed = outcomes resolved by the cutoff, fed back to score each Street source and to replay the whole path on history · Every engine returns value, sigma, citations and reliability, or abstains
  1. Four evidence sources are selected only if dated at or before the cutoff: Street calls, the frozen 1,139-document corpus, physical driver snapshots and a Polymarket price snapshot.
  2. The corpus passes through deterministic per-issuer extractors with unit and magnitude checks, producing 695 typed observations that each carry a period, units, source file and excerpt.
  3. Engine A reconstructs the Street from dated calls, scored against outcomes that had already resolved at the cutoff.
  4. Engine B builds the fundamental forecast from the typed store, with driver nowcasts and validation-gated machine learning nested inside it so neither becomes a second independent vote.
  5. Engine C reads one quantile from a dated market price, and abstains on the nine of twelve targets with no defensible market.
  6. The meta-forecaster weights each engine by reliability and a source-overlap penalty divided by predictive variance, and flags needs_review when engines disagree by more than two combined sigmas.
  7. The run writes four workbooks, a machine-readable audit and a timestamped log.
02 / Controls

The system knows what it does not know.

The base weight is inverse predictive variance. Reliability scales that precision. If two engines reuse the same evidence family—such as a prediction-market strike copied from consensus—both receive an overlap penalty. The complete formula and every realized weight are written to submission/forecast-audit.json.

raw_weight = reliability × overlap_penalty ÷ sigma²
final = Σ(raw_weight × estimate) ÷ Σ(raw_weight)
0silent missing engines
3direct EPS markets
9market abstentions
0run failures
Point-in-time boundary

Corpus documents, Street calls, driver observations and market snapshots are each selected at or before --as-of. Historical source scores only use outcomes whose event date had resolved by the cutoff.

Disagreement boundary

Pairwise engine disagreement beyond two combined sigmas sets needs_review. Missing market coverage remains visible as an abstention warning instead of becoming a copied consensus value.

Unit boundary

Exact labels and units come from challenge/companies.json. Percentages are points, Hays EPS is pence, and non-submittable units hard-fail before a workbook is written.

One output path

The Street backtest is an input to forecast.run, not a second workbook builder. One Python orchestrator produces all four final templates and the machine-readable audit.

03 / Forecasts

Twelve final numbers, with the realized engine mix.

These are descriptive outputs from the 16 August run, not accuracy results; actuals were not available at forecast time. Weights are rounded for display. Full-precision values, sigmas, citations, abstentions and overlap penalties live in the JSON audit. To follow any single number back through the lens that produced it, the stock × method explorer shows each lens's validation history, gate, derivation and consumed sources with publication dates.

CompanyMetricFinalUnitsEngine weights
HDNet sales47,105.31USDmStreet 10% · Fundamental 90% · Market abstained
HDAdjusted diluted EPS4.64USD / shareStreet 19% · Fundamental 67% · Market 15%
HDComparable sales, total company1.13%Street 63% · Fundamental 37% · Market abstained
ADIRevenue3,996.71USDmStreet 4% · Fundamental 96% · Market abstained
ADIAdjusted diluted EPS3.43USD / shareStreet 19% · Fundamental 79% · Market 3%
ADIAdjusted gross margin73.15%Street 1% · Fundamental 99% · Market abstained
HaysNet fees903.49GBPmStreet 1% · Fundamental 99% · Market abstained
HaysPre-exceptional basic EPS1.19GBpStreet 60% · Fundamental 40% · Market abstained
HaysPre-exceptional operating profit45.39GBPmStreet 56% · Fundamental 44% · Market abstained
DEWorldwide net sales and revenues12,486.77USDmStreet 9% · Fundamental 91% · Market abstained
DEDiluted EPS (GAAP)5.03USD / shareStreet 24% · Fundamental 75% · Market 1%
DEProduction & Precision Ag operating profit519.09USDmStreet 32% · Fundamental 68% · Market abstained
04 / Evidence

Metric-specific models, not one generic earnings prompt.

Home Depot

Housing, category demand and ML

Full-year guidance is divided by measured quarterly seasonality. Census NAICS 444 updates comparable sales, while task-matched historical ML contributes only after beating a seasonal naive baseline; three-sample validation receives a wide uncertainty penalty.

Analog Devices

Guidance realization and mix

Revenue and EPS guidance are corrected by walk-forward realization history. Gross margin is bridged from guided operating margin by a fitted relationship, with TXN read-through as a physical driver.

Hays

Closed-year reconstruction

FY fees are H1 actual plus Q3/Q4 actual-basis growth on the prior H2 base, less disposed-country fees. Company-compiled consensus and the top-of-range management steer anchor profit.

Deere

Units → segments → EPS

AEM machinery units, price, FX and a calibrated shipment residual produce PPA sales and profit. Segment revenue and margin bridges then reconstruct worldwide revenue and diluted EPS.

05 / Validation

The system is scored on history, not asserted.

forecast.system_backtest replays the pipeline that writes the twelve submitted numbers. For every closed period it rebuilds the corpus as of the day before the actual was published, then runs the same extractors, calibration, estimators and meta-forecaster the final command runs. Look-ahead is structural: a leak would require a filing to travel backwards in time.

Scoring uses the competition's own formula — error divided by the benchmark's, floored and capped at 5.0 — against a seasonal median year-over-year replay standing in for the Wall Street consensus we are never shown. Below 1.00 beats a model you could write in twenty lines.

How the system backtest works A closed period is chosen, the corpus is rewound to the day before its actual was published, the identical forecasting pipeline runs, and the result is scored against the reported actual using the competition formula. Measured skill is not yet fed back into engine reliability. 01 · A CLOSED PERIOD 02 · REWIND 03 · THE SAME PIPELINE 04 · SCORE 05 · WHAT IT SHOWED Pick a quarter that has already reported, so a real answer exists ADI FY2026Q2 Its actual published on the results 2026-05-20 · 3,623 USDm Cutoff publication minus one day. Corpus loads as_of=cutoff 2026-05-19 · 264 docs A LEAK WOULD NEED A FILING TO TRAVEL BACK IN TIME IDENTICAL CODE TO THE SUBMITTED RUN Extractors → engines → meta Same extractors, same calibration, same estimators, same meta-forecaster. The replay drives the real path, not a copy. forecast.system_backtest 75 cells · 4-way parallel · deterministic Against reality system 3,555 · actual 3,623 our miss 68 benchmark miss 736 Competition formula miss ÷ benchmark miss, floored, capped at 5.0 75 held-out cells mean 0.38 · median 0.15 beat benchmark on 68 under 1.00 is better Each engine alone fundamental 75 · street 27 market 0 cells the reported actual, revealed only after the cutoff MEASURED SKILL SHOULD SET ENGINE RELIABILITY — NOT WIRED, RELIABILITY IS STILL A CONSTANT WHAT THE REPLAY REFUSES TO SCORE 60 cells recorded as engine failures, each one the system declining to forecast on inputs it did not have. Deere abstains entirely: its driver chain is anchored to a dated snapshot, not a fiscal period. Settled market books price the answer, so they are rejected rather than scored.
Solid = replay path · Dashed = the actual, revealed only after the cutoff · Amber = the feedback loop we measured but deliberately did not close

75 held-out cells · mean score 0.38 · median 0.15 · beat the benchmark on 68

Company · MetricHeld-out cellsMean scoreBeat benchmark
ADI · Revenue190.1618 / 19
ADI · Adjusted diluted EPS190.1619 / 19
ADI · Adjusted gross margin190.4818 / 19
HD · Net sales50.913 / 5
HD · Adjusted diluted EPS51.183 / 5
HD · Comparable sales, total company50.544 / 5
Hays · Net fees10.151 / 1
Hays · Pre-exceptional basic EPS10.391 / 1
Hays · Pre-exceptional operating profit10.271 / 1

Per-lens history, plotted actual against predicted with each gate, is in the forecast explorer. The same final command regenerates it.

Does aggregating actually help? Each engine, scored alone

The replay scores every engine separately, on the cells where it genuinely spoke. Getting the comparison at all took a repair: the Street engine keyed its panel on company and metric with no period, so it could only address the quarter being submitted and abstained on all 75 historical cells. The archive it needed was already in the repository — 39 consensus rows in forecasting/data/historical_forecasts.csv, each stamped a day before the company reported. Keying them by period gave Street 27 cells and made the blend measurable.

On the 27 cells where both engines spokeMean scoreMedian scoreBeat benchmark
Fundamental engine alone0.4130.13223 / 27
Street engine alone0.4800.15324 / 27
Meta-forecaster, both blended0.4440.12423 / 27

The reading is mixed. The blend has the best median and beats Street on 20 of 27, so it rescues the case where one engine is badly wrong. But its mean is worse than the fundamental engine alone, which it beats on only 11 of 27. Blending dilutes the stronger engine about as often as it protects against the weaker one.

The cause is identifiable. The fundamental engine's reliability is the fixed constant 0.85 rather than its measured skill, so the meta-forecaster cannot see which engine this evidence favours. Fitting reliability to measured skill is the change this backtest argues for; it is listed as a limitation rather than done quietly and claimed as design.

The prediction-market lens took two passes. A settled market prices the answer, not a forecast: the closed books in the live snapshot sit at 0.999 against 0.001, status resolved, so reading one at a later cutoff would have scored beautifully and been pure look-ahead. parse_beat_market now requires a market to still be open at the cutoff, making that refusal a rule rather than an accident of which snapshots exist. The legitimate measurement lives elsewhere: the CLOB price series, where the last trade at least 24 hours before the scheduled print is genuinely pre-print. forecast.polymarket_history scores all nine resolved beat markets that way — the whole HD, ADI and Deere series, which only begins in November 2025, and Hays has none.

The verdict is a clean negative: Brier 0.100 the day before the print, against 0.0988 for simply predicting the base rate. Across nine markets the crowd added nothing over knowing that companies usually beat. That is the measured justification for the market engine's 0.03 reliability, which until now rested only on reasoning about what a binary quantile can mean.

What the replay refuses to score

60 further cells are recorded as engine failures rather than dropped: each is the system declining to forecast on inputs it did not have. Deere abstains entirely, its AEM driver chain being anchored to a dated snapshot rather than a fiscal period, so no historical cutoff carries what it consumes. Hays scores one period because its engine needs a company-compiled consensus the corpus holds only for recent years.

The backtest also exposed the Home Depot guidance gap. HD does not always guide a range, and the extractor matched only the “X to Y” wording, so guidance existed for FY2026 alone. Reading single-point guidance as a fallback — never an override, so a published range always wins — took HD from 3 scored cells to 15 with all four workbooks byte-identical.

Two further backtests, both in the final command

  • forecast.backtest tests one claim: that correcting management guidance for its historical bias beats parroting it. Calibration is fitted only on the corpus as it stood the day before each guidance was issued. Across 54 held-out periods it improved 42, with error ratios of 0.65, 0.74 and 0.77.
  • forecasting.backtest ranks Street sources rather than the system, scoring 39 closed observations by recency-weighted error and shrinking sparse histories toward an 8% prior. This is what makes Street reliability measured rather than chosen.

The ML lens, and the gate it has to clear to speak

Classical models are an estimator inside the fundamental engine, never a fourth top-level vote. Validation is chronological only—a random split on a time series lets a model read its own future—and model choice is made by TimeSeriesSplit within the training window, never on the test period. Gates were fixed from what would be competitive against a Street benchmark before any model ran. A prediction is written only if it clears its gate and beats a seasonal naive baseline; sparse validation additionally widens sigma, so three good historical transitions cannot outvote current evidence. Seven of twelve metrics cleared. Five abstain.

MetricDeployed modelWalk-forwardGatenVerdict
HD · Net salesShare-of-category nowcast, zero parameters, blended with a CV-weighted voting ensemble0.40%2%6pass
HD · Adjusted diluted EPSRidge on sequential ratios, plus the observed adjusted-to-GAAP wedge4.17%5%3pass
HD · Comparable salesRidge on lag 1 and lag 40.40pp0.8pp3pass
ADI · RevenueGuidance realization—guide × expanding mean of past realization ratios, zero parameters1.38%2%14pass
ADI · Adjusted diluted EPSGuidance realization3.44%5%14pass
ADI · Adjusted gross marginGuided operating margin × realization, plus a modelled opex spread0.89pp1.0pp14pass
Hays · Net feesQuarterly composition from four disclosed growth rates, zero parameters1.52%2%4pass
Hays · Operating profit, EPSNone—seven annual observations10%7abstain
DE · Revenue, EPS, PPA profitVoting ensemble over Ridge, Random Forest, Gradient Boosting, SVR and KNN7.3–24.3%2–10%4–9abstain
Model inventory

Ridge, Random Forest, Gradient Boosting, SVR with an RBF kernel and KNN, plus equal-weight voting, CV-weighted voting and ridge stacking. Features are scaled; COVID-era target rows are excluded as a documented exogenous shock rather than silently winsorized.

Intrinsic floor

Before blaming a model we measure the dispersion of the target transition’s own history. Deere’s Q2→Q3 ratio disperses 8.1% on revenue and 37.2% on PPA profit against gates of 2% and 10%, so no history-only estimator could pass. Given four times the samples the model scored 7.26%—landing on its predicted floor. That is an information limit, not a tuning failure.

What we abandoned

A year-over-year framing scored 5.40% against a 5.24% naive baseline and was cut. Ridge stacking was the worst ensemble at 6.04%, overfitting a meta-learner on twenty rows. Ensembling bought 0.2pp; reframing the target around already-published data bought an order of magnitude.

A bug the gates did not catch

Hays tags every quarterly update with the fiscal year rather than the quarter, so keying on period collapsed a twenty-point series into six and we concluded Hays was unmodelable. The loader now raises on a colliding key; an audit found the same fault in neither ADI nor Deere. Hays net fees became a passing metric at 1.52%.

Every figure on this page is re-derived from the run artefacts by npm run check:architecture, which fails if the page and the audit disagree.

06 / Explorer

Follow any single number back to its evidence.

Pick a company and a lens. The explorer plots that lens's walk-forward history — reported actual against model prediction, period by period — with its pre-declared gate, the derivation of the submitted figure, and the source documents it consumed with publication dates. A lens that failed its gate still shows its history: that history is the evidence for the abstention.

Embedded in full below, so this page stays one self-contained file. npm run forecast rebuilds it from the run artefacts.

07 / Reproduce

One final command, then structural validation.

python3 -m venv .venv
.venv/bin/python -m pip install -r requirements.txt
npm ci

npm run test:forecast
npm run forecast
npm run check:forecasts

git rev-parse HEAD

npm run forecast regenerates the gated Home Depot ML artifact, materializes the Street backtest, runs the walk-forward skill test and the system replay, and invokes the final orchestrator—all at the frozen cutoff. The runner writes four workbooks, a timestamped log, submission/forecast-audit.json, and the backtest evidence under research/. Upload remains manual by challenge rule.

Known limits

  • The public exact-metric analyst panel is sparse for several operating metrics. Those Street reconstructions use a wider-sigma research fallback and lower reliability.
  • Direct prediction markets exist only for three EPS targets. A binary beat price supplies one quantile, so its precision is deliberately discounted; all other targets abstain.
  • The Home Depot ML gate is task matched but has only three historical Q2 transitions. Its uncertainty is widened before it enters the fundamental engine, and the all-transition stress view remains visible in the artifact.
  • Regex extractors depend on recurring issuer disclosure shapes. Magnitude checks and source excerpts turn format drift into an observable failure, but do not eliminate it.
  • Source-overlap labels are declared families, not learned causal dependence. The penalty is conservative and auditable, but still an approximation.
  • The system backtest is unbalanced. ADI supplies 57 of the 75 scored cells, so the headline mean is largely an ADI result; Hays contributes 3 and Deere none. The per-metric table is the honest reading, not the single average.
  • Engine reliabilities are not yet learned from the backtest. Street reliability is derived from measured source error, but the fundamental engine's 0.85 is a chosen constant even though it carries most of the weight. Feeding measured skill back into that number is the clearest next improvement and was out of scope for the day.
  • The replay benchmark is a seasonal baseline, not the Wall Street consensus the accuracy prize scores against. Beating it is necessary evidence, not sufficient.
  • Blending is measurable but not yet vindicated. On the 27 cells where two engines speak the meta-forecast has the best median and beats Street on 20, but its mean is worse than the fundamental engine alone and it beats that engine on only 11. On this evidence a judge would be right to ask whether the blend earns its place, and the answer today is that it earns it against a weak engine and costs against a strong one.
  • Engine reliability is chosen, not learned, and that is the direct cause of the result above. The fundamental engine sits at a fixed 0.85 while Street is derived from measured source error, so the meta-forecaster cannot see which engine the backtest says is better. Fitting reliability to measured skill is the clearest next change and was deliberately not attempted late in the day, because it would move submitted numbers with no time to validate the move.
  • The prediction-market record is nine markets long. The beat series only begins in November 2025 and Hays has none at all, so the Brier result is directionally clear but statistically thin.
  • The blinded LLM replay holds only two genuinely post-training quarters per metric. Three metrics that pass overall fail on that post-training subset, which is the honest place to look for memorisation, and percent metrics stay unscaled so recall on them is measured rather than prevented.