Journal
PerspectiveJuly 2026

Why More Data Does Not Always Solve the Problem

More observations can reduce sampling uncertainty. They cannot repair a wrong measurement, remove hidden confounding, recover a missing causal quantity or guarantee transfer to a new regime.

More data is one of the most reliable ways to improve a model—when insufficient sample size is the problem.

It can reduce sampling variability, make rare events easier to estimate, support more precise subgroup analysis and reveal patterns that a small dataset could not distinguish from noise. Those gains are real.

But not every modelling failure is statistical. Some are structural. The system may measure the wrong quantity, combine observations generated by different processes, omit a common cause, lack the counterfactual needed to answer the question or rely on a correlation that exists only in the development environment.

In those cases, adding more rows of the same kind can make an answer more precise without making it more correct.

The central question is therefore not simply, “How much data do we have?” It is, “Does this data contain the information required to identify the quantity we need?”

More data solves a particular class of problem

Suppose the target is well defined, the measurement is valid, the observations represent the intended population and the relationship being estimated remains stable. Under those conditions, additional independent data can reduce uncertainty in the estimate. A wider sample may also improve coverage of uncommon but relevant cases.

That is a statistical limitation: the information exists in the data-generating process, but the current sample is too small or noisy to estimate it reliably.

A structural limitation is different. The available observations do not identify the requested answer under the assumptions that can be defended. The missing information may be a variable, an intervention, a regime label, a valid measurement or evidence from the target environment. Repeating the existing observation process does not supply it.

This distinction can be stated plainly:

  • Statistical uncertainty asks how precisely a defined quantity can be estimated from appropriate data.
  • Structural uncertainty asks whether the observed data and assumptions determine the right quantity at all.

A large dataset can narrow a confidence interval around a systematically biased value. Scale does not convert validity into a sampling problem.

Wrong measurements produce precise answers to the wrong question

A dataset does not contain the real world. It contains readings produced by instruments, labels, proxies, coding practices and collection decisions.

Some measurement error is approximately random. Repeated independent readings can reduce its effect. Systematic error behaves differently. If an instrument is biased under a particular condition, collecting more readings under that condition estimates the biased output more precisely. If a label is an imperfect proxy for the decision target, more labelled examples improve prediction of the proxy—not necessarily the quantity that matters.

The distinction is visible in pulse oximetry. Pulse oximeters estimate arterial oxygen saturation non-invasively; arterial blood-gas measurement provides a different, direct reference. In a retrospective analysis of two hospital cohorts, Sjoding and colleagues found that occult hypoxaemia—low arterial oxygen saturation not detected by pulse oximetry—occurred nearly three times as frequently among self-identified Black patients as among White patients in the analysed range. The issue was not a shortage of pulse-oximeter readings. The measurement relationship itself differed across the compared groups. Sjoding et al., New England Journal of Medicine (2020)

This example does not imply that every device or measurement error has the same cause, nor that a group label is a substitute for directly measuring the relevant physical factors. It demonstrates the structural point: an indirect measurement must be validated against an appropriate reference across the conditions and populations in which it will be used.

The remedy depends on the failure. It may require a better instrument, direct measurement of the target, redesigned labels, calibration across the operating range, adjudication by qualified experts or explicit modelling of measurement uncertainty. More data helps only if the new collection changes the information available rather than repeating the same bias at greater scale.

Confounding survives an arbitrarily large observational sample

Confounding occurs when a variable influences both the candidate cause and the outcome, creating an association that does not equal the causal effect of changing the candidate cause.

For example, in observational treatment data, disease severity can influence both which treatment is selected and the eventual outcome. A model may learn that one treatment is associated with worse outcomes because it is given to sicker patients. With enough records, that association can be estimated very precisely. It still does not reveal what would have happened had otherwise comparable patients received a different treatment.

This is not a weakness that model capacity can remove. It is an identification problem.

Pearl's causal-diagram framework formalised this distinction. Subject-matter assumptions about which variables cause others determine whether a causal effect can be recovered from non-experimental observations. When the assumptions and observed variables are insufficient, the framework indicates that additional observations or auxiliary experiments are required. The observational distribution alone does not decide the causal structure. Pearl, Biometrika (1995)

Statistical adjustment is therefore not a ritual in which every available variable is added to a model. Adjusting for the wrong variable can introduce bias, while an unmeasured common cause cannot be removed by a more flexible fit to the measured columns.

Defensible causal work begins by stating the intervention, outcome, time horizon and causal assumptions. Depending on the problem, evidence may then come from randomisation, a justified natural experiment, measurement of the missing common causes, negative controls, sensitivity analysis or another design that changes what can be identified. These methods have different assumptions; none allows data volume to substitute for design.

Hidden operating regimes turn one dataset into several problems

Many systems do not follow one stable input-output relationship. They occupy regimes with different dynamics, constraints or failure modes.

A battery may behave differently across temperature and degradation states. A manufacturing process may change after maintenance or a supplier substitution. A patient population may contain clinically distinct subtypes under the same coarse label. An autonomous platform may move between environmental conditions that change which sensors remain informative.

If the regime is unobserved or omitted, a model trained on pooled data estimates an average across processes. That average can perform well on the dominant regime while failing a small but consequential one. Adding more common-regime observations can improve the aggregate score without materially improving the weak regime.

Oakden-Rayner and colleagues called this hidden stratification in medical imaging: a labelled class contains meaningful but unidentified subclasses with different performance. Across multiple medical-imaging datasets, they found relative performance differences greater than 20% on clinically important subsets. In one hip-fracture dataset, a model with an overall reported AUC of 0.994 had sensitivity of 0.981 across fractures at the evaluated operating point, but sensitivity fell to 0.900 for subtle fractures and 0.911 for cervical fractures. Oakden-Rayner et al., ACM CHIL (2020)

The specific numbers belong to those imaging tasks. The general lesson is about aggregation: average performance cannot reveal a regime that the labels and evaluation do not distinguish.

Resolving hidden regimes may require more data, but it must be different data: measurements that expose regime membership, deliberate coverage of transitions and rare states, finer labels, time-ordered records or experiments that vary the operating conditions. Evaluation should report regime-specific and worst-relevant-group performance rather than relying only on the pooled mean.

Missing causal quantities are not hidden in the rows

Prediction asks what outcome is likely given what was observed. Many decisions ask a different question: what outcome would occur if an action were changed?

For any individual case, only the outcome under the action actually taken is observed. The outcome under the alternative action is not another column waiting to be collected from the same event. Holland described this absence of simultaneous potential outcomes as the fundamental problem of causal inference. Holland, Journal of the American Statistical Association (1986)

At population level, causal effects can still be estimated under suitable designs and assumptions. Randomisation creates comparable groups in expectation. Observational methods require conditions that make the comparison defensible. But if historical policy always assigns one action in a particular region of the data, there is no empirical overlap there: the alternative outcome for comparable cases is unsupported.

More observations generated by the same policy preserve the gap. They can make the observed branch extremely well characterised while leaving the unobserved branch unidentified.

Missing causal quantities appear beyond treatment effects. An organisation may have records of equipment that was operated, but no evidence about what would have happened under a different control policy. It may observe completed inspections but not failures prevented by inspections that were never scheduled. It may record an alarm and a response without independent evidence of whether the underlying event was real.

The response is to define the required quantity before data collection. If the decision concerns intervention, the evidence programme must create or justify the comparisons needed to estimate an intervention—not merely predict the history produced by the existing policy.

Correlations do not automatically transfer

A correlation can be stable inside one institution, device fleet, laboratory protocol or market and disappear elsewhere because it depended on local practice rather than the underlying phenomenon.

Zech and colleagues tested pneumonia-detection models using 158,323 chest radiographs from three hospital systems. In three of five natural comparisons, performance was significantly lower on external-hospital data than on held-out data from the training hospital. The models could also identify hospital system and, in one analysis, perfectly distinguished inpatient from emergency-department radiographs. Those site and workflow signals were correlated with pneumonia prevalence and could support prediction internally without representing the disease finding the model was intended to detect. Zech et al., PLOS Medicine (2018)

Pooling more data from the same sites improved performance on new pooled data from those sites but did not establish external performance. The missing evidence was not another internal image. It was a test of the relationship in a genuinely different hospital environment.

Transport is a separate scientific claim. Bareinboim and Pearl formalised conditions under which causal information from one environment can be combined with observational or experimental evidence from another. Their work makes the dependency explicit: transfer requires assumptions about what differs and what remains invariant between source and target populations. It is not guaranteed by source-sample size. Bareinboim and Pearl, Proceedings of the National Academy of Sciences (2016)

For predictive systems, the practical requirements are similarly concrete. Define the target population and environment, test on independent sites or future periods, document changes in measurement and policy, and identify which relationships have a reason to remain stable. A random test split from the source population cannot support an unrestricted claim about another population.

More data can reveal structure—but only if collection changes

The argument is not that structural problems are immune to evidence. It is that the useful unit is new information, not raw row count.

Additional data can resolve a structural uncertainty when it does at least one of the following:

  • measures the quantity previously represented only by a proxy;
  • captures a missing common cause or effect modifier;
  • includes a justified intervention or comparison group;
  • labels or spans previously hidden operating regimes;
  • covers the ranges and combinations required for deployment;
  • comes from the target site, population, device or future period; or
  • tests an assumption against an independent reference.

In contrast, duplicated measurements, denser sampling of the dominant regime or additional records created by the same selection policy may add volume without adding the information required by the decision.

This is why a smaller, deliberately designed dataset can be more valuable than a much larger convenience dataset. The comparison is not “small versus large.” It is “identifying evidence versus non-identifying evidence.” Once the design is valid, scale becomes valuable again.

A structural audit before a collection programme

Before funding more collection, an organisation should be able to answer eight questions.

  1. Decision quantity: Is the required output a prediction, a causal effect, a physical state, a risk, a counterfactual or a control action?
  2. Measurement: Is that quantity observed directly? If not, what proxy or instrument produces it, and against which reference has it been validated?
  3. Selection: What determines which cases enter the dataset and which outcomes become visible?
  4. Confounding: Which variables can influence both the proposed action or exposure and the outcome?
  5. Regimes: Which physical, biological, operational or policy conditions can change the governing relationship?
  6. Support: Are all deployment ranges, combinations and action alternatives represented with meaningful coverage?
  7. Transfer: What differs between the development environment and the target environment, and why should the learned relationship persist?
  8. Falsification: Which observation or experiment could show that the assumed structure is wrong?

These questions change the data plan. The next investment may be more observations, but it may instead be a better sensor, a reference measurement, a new site, an intervention, a regime label or an experiment designed around the missing comparison.

Structure is not a substitute for validation

Domain knowledge can also be wrong. A mechanistic assumption may omit an interaction, use an invalid boundary condition or simplify away a regime that matters. An interpretable causal story can be more persuasive than the evidence warrants.

Structural modelling therefore requires the same discipline it demands of statistical learning: explicit assumptions, held-out tests, sensitivity analysis, uncertainty and failure criteria. Known constraints should be treated as testable commitments, not as immunity from data.

The strongest systems use data and structure to challenge one another. Data estimate quantities and expose mismatches. Structure determines which quantities matter, which comparisons are valid and where extrapolation requires justification.

The Mondren perspective

Mondren does not treat data scale as unimportant. We treat it as one part of the evidence architecture.

For problems governed by physical behaviour, biological history, operating constraints or changing regimes, the first task is to determine what must be represented and what the available data actually measure. Only then is it meaningful to decide whether the limitation is sample size, missing information or an invalid problem formulation.

Our public position is simple: collect data that changes what can be known. More observations are valuable when they improve coverage or precision inside a defensible design. When the missing element is a causal comparison, regime indicator, reference measurement or target-environment test, that is the evidence the programme should seek.

This article reports no Mondren benchmark and discloses no proprietary model architecture, state representation, feature construction, parameter values or implementation logic. The problem classes and validation principles are public. Construction remains proprietary.

The commercial implication

Data programmes are often priced by volume: more sensors, more records, more storage and more labelled examples. That can be expensive without addressing the decision risk.

A better investment case asks what uncertainty the next unit of data will reduce. If the answer is sampling variance in a valid design, scale may be exactly the right investment. If the answer is unclear because the target, measurement or operating regime is unresolved, collection should pause until the information gap is specified.

This changes procurement and research planning. Instrument validation, experimental design, external-site testing and regime coverage become first-class deliverables rather than secondary checks after a model has been trained.

More data is not the opposite of better modelling. It is most valuable when the structure of the problem tells us which data will be informative.

The failure to distinguish statistical scarcity from structural absence creates false confidence. A billion observations of the wrong quantity remain observations of the wrong quantity. A perfectly estimated correlation remains non-causal if the required assumptions are absent. A source-domain relationship remains unproven elsewhere until transfer is tested or justified.

The objective is not to collect less. It is to ensure that every collection programme is capable of changing the answer to the question that matters.

The healthcare examples in this article illustrate measurement and validation principles. They are not medical advice or claims about the performance of all devices or clinical models.