Category Zero says: stop guessing, fuse everything, and let the structure nominate the questions. Here is the actual machinery, the real papers behind it, and the one rule that stops the whole thing from being elaborate nonsense. Cheeky, but every method below is real and cited.
It is called data-intensive, or data-driven, or hypothesis-free discovery. Jim Gray called it the fourth paradigm of science, a mode as distinct from theory and experiment as those are from each other: you start from the data, at scale, and let patterns nominate the theory rather than the other way around. The ocean is a near-perfect candidate, because it has been generating far more data than anyone has had the working memory to combine.
→ Hey, Tansley & Tolle (eds.), The Fourth Paradigm, 2009.
A machine can hunt all three at once, across more variables than a person can hold, and with no opinion about which fields are allowed to talk to each other. Each one has real, off-the-shelf tooling.
Here is the trap. Point a tireless machine at every possible correlation and it will find millions, and nearly all of them are noise dressed as signal. The number of ways to slice data is astronomical, and if you go looking you will always find something. Statisticians call it the garden of forking paths: even honest, unplanned choices inflate false positives until a chance pattern looks profound.
→ Gelman & Loken, The Garden of Forking Paths, 2013.
So the discovery is never the correlation. The discovery is the correlation that survives. The rule, non-negotiable:
→ The held-out principle is old and boring and load-bearing: Kohavi, A Study of Cross-Validation, IJCAI 1995.
This whole Guide was built collaboratively by AI agents, which is exactly the kind of thing that fools itself if you let it. So the same discipline that guards Category Zero guarded the build:
None of this makes the machine right. It makes it catchable when it is wrong, which is the most you can honestly ask of a discovery engine, silicon or otherwise.