The research behind the hardest question.

Kapnova uses econometrics, forecasting, simulation and optimization to model business decisions. Its core research advantage is causal identification: determining whether a lever actually changed an outcome or simply moved alongside it.

That capability is rooted in published MIT doctoral research and adapted here from survival analysis to real business decisions.

Six methods. One decision engine.

Each method answers a different part of the question. Causal identification determines when the evidence supports saying a decision actually caused the outcome.

1
Descriptive
What happened, and where?
2
Forecasting
What is likely to happen if nothing changes?
3
Econometric estimation
How strongly is a lever related to an outcome, and with what uncertainty?
4
Causal identification
Did the lever actually change the outcome, or did it simply move alongside it?
5
Optimization
What is the best move inside real budget, margin and inventory constraints?
6
Experiments
What should we run when the data cannot support the claim on its own?

The methods work together. Identification is what lets a forecast or an optimization be stated as a claim about what a decision will cause, rather than what has tended to move alongside it, which is the difference between a number that describes the past and one you can act on.

MIT / doctoral thesis / 2025
Causal Inference with Survival Outcomes via Orthogonal Statistical Learning
Shenbo Xu / Massachusetts Institute of Technology

It develops causal methods for the highest-stakes setting there is, whether a drug causes a change in cancer or survival outcomes, estimated from observational data. Kapnova is built on years of this doctoral research.

Read the thesis dspace.mit.edu/handle/1721.1/158798
SURVIVAL TIME 1.0 0.5 0 causaleffect treated control

The published methods behind that step.

These are the methods Kapnova uses to build the causal models behind every answer. They are how the engine works out what actually drove an outcome, and what would happen if you changed it. Each one below names the chapter it comes from and the decision it answers.

Target trial emulationthe core promise Ch. 4 / the metformin study

The thesis reconstructs a randomized trial it cannot run, whether a diabetes drug prevents cancer, from observational data alone. Your company cannot run a controlled experiment on an 8% price move either, so the engine emulates that trial from your own history. This is the named, published version of “what will this cause,” and it is why the whole approach is more than a forecast.

Double & debiased estimation, Neyman orthogonalitythe estimation engine Ch. 2

Estimators built to stay consistent even when the underlying models are imperfect. This is why an effect the engine reports is defensible rather than fragile, and why the number survives a finance review.

Overlap weighting under poor covariate balancethe loyalty case Ch. 2

A real effect even when the treated and untreated groups barely compare. That is the loyalty self-selection problem exactly, where members and non-members are not alike to begin with.

Competing risks & separable direct effectsthe decomposition engine Ch. 2 & 3

Modeling mutually exclusive outcomes so one cannot mask another, and splitting an effect into its direct and mediated paths. These are the flat-number-hides-a-loss case and the price-to-sentiment-to-demand case.

Heterogeneous treatment effectsaudience-level answers Ch. 3

The effect of a decision, and for whom, not just on average.

Causally-informed AIthe agents Ch. 5

The thesis uses foundation models to screen many drug-disease pairs for causal signal. The agents do the same for decisions, generating and ranking hypotheses by estimated causal effect, which grounds the agentic argument in published work, not aspiration.

Peer reviewed / American Journal of Epidemiology / 2025
Can metformin prevent cancer relative to sulfonylureas?
Shenbo Xu, first author, with Zheng, Su, Finkelstein, Welsch, Ng and Shahn

The target trial emulation the causal engine runs on. It compares two diabetes therapies across 93,353 patient records from a UK primary-care database, and it faces the two problems every business decision faces: poor overlap between the groups being compared, and a competing risk that can mask the outcome. The answer arrives as an interval, not a point.

Read the paper doi.org/10.1093/aje/kwae217
Accepted / ICML 2026 / Seoul
Finding the minimal parameter budget for implicit reasoning
Wang, Tan, Shenbo Xu, Jin, Wang, Panda and Shen

A scaling law for how much model capacity implicit reasoning actually needs, accepted at one of the three top-tier machine learning conferences alongside NeurIPS and ICLR. This is the machine learning research behind the agents that build the causal maps, rather than the causal methods themselves.

Read on OpenReview openreview.net/forum?id=iuIPAhZpxz
Publications

The work behind the engine, in full.

Shenbo Xu’s peer-reviewed and forthcoming work in causal inference and machine learning. The methods Kapnova runs are drawn from these papers and from the dissertation above.

Peer-reviewed journals2
  • Can metformin prevent cancer relative to sulfonylureas? A target trial emulation accounting for competing risks and poor overlap via double/debiased machine learning estimators
    Xu, Zheng, Su, Finkelstein, Welsch, Ng, Shahn
    American Journal of Epidemiology, 2025, 194(2), 512–523First author
    The target trial emulation the causal engine runs on.
  • Systematically exploring repurposing effects of anti-hypertensives
    Shahn, Spear, Lu, Jiang, Zhang, Deshmukh, Xu, Ng, Welsch, Finkelstein
    Pharmacoepidemiology and Drug Safety, 2022, 31(9), 944–952
Conference proceedings4
  • Finding the minimal parameter budget for implicit reasoning: a data complexity driven scaling law for language models
    Wang, Tan, Xu, Jin, Wang, Panda, Shen
    Accepted, ICML 2026, SeoulOpenReview
    One of the three top-tier machine learning conferences, alongside NeurIPS and ICLR.
  • Foundational model-aided automated high-throughput drug screening using self-controlled cohort study
    Xu, Cobzaru, Finkelstein, Welsch, Ng
    AI for New Drug Modalities, 38th NeurIPS, Vancouver, 2024First author
  • Anti-diabetic drug repurposing using electronic health records: design, emulation and analysis of a synthetic in-silico clinical trial for Alzheimer’s disease
    Finkelstein, Xu, Su, Zheng, Charpignon, Tzoulaki, Middleton, Welsch
    Machine Learning for Healthcare, Ann Arbor, 2019
  • Repurpose anti-diabetic drugs for cancer based on causal evidence
    Xu, Finkelstein, Welsch, Su, Zheng, Charpignon, Tzoulaki
    CFE-CMStatistics, London, 2019First author
Under review3
  • Estimating heterogeneous treatment effects on survival outcomes using counterfactual censoring unbiased transformations
    Xu, Cobzaru, Finkelstein, Welsch, Ng, Shahn
    Journal of Machine Learning ResearchFirst authorarXiv:2401.11263
  • Double/debiased machine learning for time-to-event outcomes under poor overlap
    Xu, Finkelstein, Welsch, Ng, Tzoulaki, Shahn
    ICLR 2026, causal reasoningFirst authorarXiv:2305.02373
  • Beyond pre-training memorization: reinforcement learning for recovering reasoning capabilities in over-parameterized models
    Xu, Wang, Tan, Shen
    Working paperFirst author

The same estimators that weigh a cancer-drug question weigh your revenue decisions.

The thesis is built on the counterfactual, what would have happened under the other choice, the exact quantity a dashboard never contains and the one the engine estimates. The techniques that make that estimate trustworthy in medicine are the ones running under every Kapnova decision.

In the interest of precision

The author's own research page frames these methods as cross-domain, listing sales and marketing among the applications.

The platform applies methods from this research and the broader causal-inference field. It is not itself a peer-reviewed artifact, and the thesis validated the methods in medicine, not business.

Start with your URL. Go deeper with your data.

Start with your URL and Kapnova will surface what it can find from public data alone. Then connect your own data to measure what past decisions actually caused and where the next revenue or profit opportunity may be.