Why We Need a Method at All
Start with one uncomfortable fact: you can never study every patient. The whole of clinical research exists because we are forced to learn about a large group we cannot reach — every woman who will ever have a caesarean, every pregnancy that will ever be complicated by pre-eclampsia — by looking at a small group we can actually observe. The large, often loosely defined group we care about is the population. The small group we measure is the sample. Everything else in this chapter follows from the gap between those two words.
That gap has a consequence worth sitting with: if you repeated your study tomorrow with a fresh sample, you would get a slightly different number. The true effect in the population does not move; your estimate of it wobbles from sample to sample. The scientific method is simply the set of disciplines that keep that wobble — and the systematic distortions that can hide inside it — from fooling you and harming a patient.
So the scientific method is not a school-level slogan. It is disciplined clinical doubt: the route by which a registrar moves from "this seems to be happening in our labour ward" to a claim that can be tested, criticised, repeated, and used to change practice. It protects patients from confident anecdotes. In obstetrics and gynaecology this is not abstract philosophy: magnesium sulphate for pre-eclampsia, tranexamic acid for postpartum haemorrhage, HPV testing for cervical screening, and the national confidential enquiry into maternal deaths all exist because somebody converted observation into a reproducible method.
The Primary FCOG examiner is not asking you to recite "observe, hypothesise, experiment, conclude." The examiner wants to know whether you understand why a claim may be wrong. Was there a denominator? Was the sample representative? Was the comparator fair? Was the outcome defined before the data were seen? Could a confounder explain the association? Did the measurement actually measure the clinical construct? Could bias, chance, or selective reporting create the apparent effect? The scientific method is the grammar behind those questions, and the rest of this chapter builds them up one step at a time: first the sample-to-population leap, then the cycle that turns observation into a testable claim, then how that claim can be defeated by chance, bias, or confounding, and finally how transparent reporting lets others check your work.
The Classical Cycle
Science is usually drawn as a sequence, but in real clinical work it is a loop. A result rarely closes a question permanently; it refines the question, exposes a measurement problem, or identifies a subgroup needing a better study.
| Step | What it demands | O&G example | Failure mode |
|---|---|---|---|
| Observation | A pattern noticed systematically enough to be worth testing | More post-caesarean wound sepsis after a theatre move | Anecdote mistaken for incidence |
| Question | A population, exposure/intervention, comparator, and outcome | Among women having caesarean birth, did the theatre move change 30-day surgical-site infection? | "Did infection get worse?" is too vague |
| Hypothesis | A falsifiable prediction made before analysis | Infection risk is higher after the move than before | Hypothesis invented after seeing the data |
| Protocol | Eligibility, variables, outcome definitions, data source, and analysis plan | Include all caesareans over two 6-month periods; define SSI by standard criteria | Changing definitions between periods |
| Measurement | Valid, reliable, complete data collection | Same wound-sepsis definition, same follow-up window, same denominator | Ascertainment improves after the move and mimics harm |
| Analysis | Effect size, precision, chance, bias, and confounding | Risk ratio, risk difference, confidence interval, case-mix comparison | Reporting only p < 0.05 |
| Interpretation | Clinical judgement against biology and limitations | Was the move causal, or did referral case-mix change? | Association called causation |
| Revision | Practice, theory, or method adjusted cautiously | Audit prophylactic antibiotic timing and aseptic workflow | Overreacting to a biased result |
The cycle matters because clinical systems change constantly. A district hospital may introduce a new induction protocol, a province may change cervical screening pathways, or a unit may open a high-care bed. The scientific method is how one decides whether the apparent improvement or deterioration is real.
Sample, Population and Inference
Every step of that cycle rests on the leap introduced at the start: we measure a sample and reason about a population. The word "statistics" misleads many doctors into thinking the subject is arithmetic. It is not. Statistics is the discipline of inference — of stating, honestly and with quantified uncertainty, what a sample can and cannot tell us about the population it was drawn from. A paper can be drowning in numbers and contain no real statistics; a thoughtful critique of how a sample was selected, or whether a reference standard was valid, is itself a statistical act even without an equation.
Two ideas anchor everything that follows.
Representativeness. The sample's value is entirely conditional on how well it stands in for the population. A larger sample generally drifts closer to the truth, but size alone does not guarantee it: a systematically skewed recruitment method produces a precisely wrong answer no matter how large the sample grows. This is the single most common reason a confident-looking O&G study fails to apply to your patients — a tertiary fetal-medicine cohort is not a representative sample of all pregnancies, so its abnormal-Doppler risk does not transfer to a district antenatal clinic.
Sampling variability. Because each sample is a different draw from the same population, repeating a study gives a different number even when nothing has changed. On average the estimates would agree with the population value, but any single study sits somewhere on that scatter. This is why we report not just a point estimate but its precision (the confidence interval) — it expresses how far the truth might plausibly sit from the number we happened to measure. Two errors flow directly from this wobble, and the examiner expects you to separate them.
| Error | What happens | Plain-language meaning | Mainly driven by |
|---|---|---|---|
| Type I (false positive) | The sample shows a "significant" effect that does not exist in the population | A chance finding mistaken for a real one | The significance threshold (conventionally a 5% risk) and multiple testing |
| Type II (false negative) | The sample misses a real effect that does exist in the population | A true effect overlooked | Too small a sample — inadequate power |
The conventional 5% significance threshold and 80–90% power are arbitrary conventions, not laws of nature, and the mechanics of how they are set belong to the Primary hypothesis-testing chapter. The scientific-method point is narrower and permanent: chance can manufacture a false signal, and a study too small to detect a real effect cannot rule it out. A non-significant result from an underpowered trial does not mean the treatments are equivalent — absence of evidence is not evidence of absence.