Research lab

ZeonLab

A research lab working on causal reasoning about investments: how a machine can hold a structural account of why an asset moves, keep that account calibrated against thin evidence, and then be graded on it under a rule it did not write.

The programme

The premise runs against the grain of most machine learning applied to markets. The standard move learns a mapping from a large sample — many instruments, many periods, a function fitted to whatever regularity survives — and it works where the sample is large and the world sits still. The questions that decide an investment are neither. Whether a regulator approves a filing, whether a capacity cycle has turned, whether a central bank has changed its reaction function: the informative precedents number in the dozens, and the episodes nearest in time are often the least like the present.

So we begin from the other end. The structure is assumed up front — a causal topology of drivers, entities, themes and archetypes, with edges stating what propagates to what — and observation is given a narrower job: not to discover that structure but to calibrate the latent quantities inside it, how strong a driver is now, how much weight an edge carries, what stage a theme has reached. The familiar instance of the manoeuvre is implied volatility. Assume a diffusion, observe a handful of quoted contracts, back out the one parameter that reconciles the assumed structure with the market, and the whole surface follows at strikes nobody ever quoted. None of that is a pattern fitted to history; the leverage comes from having committed to a structure before looking. The same shape recurs in the world-model line of machine learning research, where an encoder and a predictor are posited and sparse pairs infer the latents that hold the structure together. Our subject is that manoeuvre applied to causal reasoning about investments, with a language model doing the inference.

The hard part is that a topology assumed once does not stay true, and a structure calibrated straight through a regime break is worse than none, because it is confidently wrong. Half the work is therefore about the validity envelope rather than the estimate: deciding whether a contradiction is noise, an overshoot, a transition or a genuinely new regime; storing every edge with the horizon and the regime over which it is claimed to hold; and treating the lab's own published conclusions as part of the dynamics they describe, since a view that is acted on has stopped being an observation of an independent world.

A system that reasons in prose can be argued with but not scored, so an unusual share of the effort goes into grading rather than into performance. Claims are stored as dated commitments whose resolution criterion was fixed before the fact. A backtest driven by a language model over any window inside that model's own training period is contaminated by construction — the weights may already hold the outcome — so every decision carries a measure of how much it knew that it should not have known, and the result is discounted by it. The scoring code has a verdict meaning not evaluable, and it is not decorative: where the discount is itself uncalibrated, or the adjusted interval straddles zero skill, a report must say that no skill claim is defensible rather than publish the number it happens to hold. That refusal is why no performance figure appears anywhere on this page, and we would sooner defend the refusal than a figure.

Currently

The line under way is forecasting. A milestone — a filing approved, a plant at volume production, a policy rate moved at a named meeting — is stored not as a probability attached to a date but as a distribution over time to occurrence with an explicit mass on never, because the collapsed scalar cannot be expanded back into the law it came from, and a book that re-marks weekly asks every week a question the scalar cannot answer. Base rates are retrieved from a reference class of historical episodes and never fitted from the episode being forecast; the present situation enters only through a multiplier; and any surface showing a hazard must print beside it how much of that number is prior rather than evidence. Running alongside it is the contamination work, written up with its identification strategy and the out-of-sample gate that downgrades the discount when the discount cannot itself be validated. Both are working papers, and they sit in a reading room that stays invitation-only for now. What they describe runs on paper positions: it manages no capital and routes no orders.

Contact

There is nothing here to sign up for. If this sits close to your own work — small-sample causal inference, scoring judgemental forecasts, contamination in model-driven backtests — write to admin@zeonlab.ai.