Measurement First

Measure first, explain later.

NucleoScope's job isn't to hand you a theoretical answer — it's to turn SVS images into inspectable, exportable, reproducible statistical measurements of nuclear populations.

Tumor versus non-tumor status has so far been observed to correspond to a single statistical feature of the RSi distribution (Tail Runaway / Tail Closed) — without prior specification of species, tumor type, or tissue site.

Why Is This Unexpected

Why is this observation unexpected?

For decades, pathology has treated nuclear atypia as one of the most fundamental indicators of malignancy. However, pathology has generally evaluated nuclear atypia within task-specific frameworks — different species, different organs, different tumor types — each with their own grading systems or diagnostic criteria.

Our observation asks a different question. Rather than asking whether nuclear atypia exists, we ask: can one predefined quantitative morphology variable exhibit the same statistical correspondence with tumor versus non-tumor status, without prior specification of species, tumor type, or tissue site? So far, we have repeatedly observed such a correspondence. Whether this represents a genuine biological regularity remains an open question requiring independent validation.

One predefined statistical measurement ↓ No task-specific calibration ↓ Cross-task statistical correspondence
The novelty is not that nuclear morphology relates to cancer. The novelty is the possibility that one predefined statistical measurement may generalize across pathological contexts without task-specific calibration.
Instrument Logic

From image to statistical spectrum

SVS
Nucleus segmentation
Per-nucleus RSi
Population distribution
Tail readout
SVS → ≥10,000 nuclei → one RSi per nucleus → RSi distribution → Tail Closed / Tail Runaway
How does one nucleus become a number?

R → S → RSi → Nuclear-State Spectrum

The computer first sees a nucleus, not "cancer." It then computes two structural quantities: R (Structural Field Fluctuation) and S (Geometric Bending Deviation), combined as RSi = R / S into a single value per nucleus.

R: Structure Field Fluctuation

  • Looks atThe 2D structural field formed by the whole nucleus
  • In plain termsHow pronounced the internal structural fluctuation is
  • Is notSimple edge-roughness statistics

S: Geometric Bending Deviation

  • Looks atThe nucleus's outline
  • In plain termsHow far the boundary deviates from an ideal circle
  • Is notA measure of real material stiffness
R looks inside the nucleus; S looks at its boundary. RSi = R / S combines the inside and outside structural information into one number.

Each nucleus also gets a second, independent measurement: C (Concavity — Dent Count). The Tail Closed / Tail Runaway readout described on this page is based on the RSi distribution alone; C is tracked separately as an additional per-nucleus measurement, not currently part of the Tail classification.

C: Concavity (Dent Count)

  • Looks atHow many inward-bending regions ("dents") appear along the nucleus boundary
  • In plain termsHow many "dents" the nucleus has — like counting dents on the surface of a ball
  • Is notDerived from R or S — it's measured independently, directly from the boundary shape

RSi vs. C

  • RSiInside + outside structure, combined into one number (R / S) — this is what today's Tail Closed / Tail Runaway readout is based on
  • CHow many dents the boundary has, counted independently — tracked, but not currently used in the Tail classification
  • BothMeasured for every nucleus, and both belong to the fuller Nuclear-State Spectrum this instrument can, in principle, read out
C counts how many "dents" the nucleus has — a smooth nucleus has a low C, a heavily deformed one has a high C. Today's Tail readout doesn't use it yet.
What NucleoScope Actually Measures

It measures the Nuclear-State Spectrum — Tail is only one readable feature of it

NucleoScope does not measure "Tail Status." It measures the Nuclear-State Spectrum of a nuclear population. Tail Closed / Tail Runaway is one readout drawn from that spectrum — not the spectrum itself, and not the definition of the instrument.

SVS whole-slide image ↓ Per-nucleus state (RSi, C) — RSi = R / S, C = concavity (dent count) ↓ Distribution across a large nuclear population ↓ Nuclear-State Spectrum ↓ Multiple spectral features / readouts

Core shape

Peak position, spectral width

Tail behavior

Tail width, tail ratio, Tail Closed / Tail Runaway, spectral cliff

Population structure

Bi/multimodal structure, subpopulations, spatial heterogeneity, transition bands between states

Cancer is the first spectral association found so far, not the instrument's definition. In the data personally analyzed to date, the tail region of the spectrum shows a recurring association: tumor/cancer samples tend to present as Tail Runaway, non-tumor samples as Tail Closed. This association still requires independent validation. The instrument is not a "Tail detector" or a "cancer detector" — it is a Nuclear-State Spectrometer. Other spectral features (peak position, width, multimodality, subpopulations, spatial organization) may in the future be informative for other biological questions — precancerous change, inflammation, aging, drug response, development, tissue injury and repair, cross-species nuclear states, intratumoral heterogeneity, or variation among normal tissues — none of which have been tested yet.
Understanding RSi through a "temperature-like variable"

RSi itself isn't temperature — but the RSi distribution functions like temperature, a macroscopic quantity.

A single nucleus's RSi is like a single molecule's kinetic energy — a microscopic quantity, not temperature. Temperature is macroscopic: it isn't a property of one molecule but emerges statistically from the kinetic-energy distribution of a huge number of molecules (this is also how "effective temperature" is used in non-equilibrium systems). Likewise, one nucleus's RSi is just a microscopic state value; what actually functions like temperature is the RSi distribution across at least 10,000 nuclei — the Nuclear-State Spectrum.

Microscopic quantity

One nucleus → one RSi, like one molecule's kinetic energy.

Macroscopic quantity

The RSi distribution across ≥10,000 nuclei → the Nuclear-State Spectrum, the true temperature analogue.

Scientific boundary: "Functions like temperature" applies at the level of the RSi distribution (Nuclear-State Spectrum), not any single nucleus's RSi value. This distribution is not Celsius temperature, nor an established thermodynamic temperature or entropy measurement — it's closer to the generalized "effective temperature" concept used in non-equilibrium systems.

A thermometer outputs one scalar number (39.1°C). NucleoScope outputs a full distribution — closer, structurally, to a spectrometer, mass spectrometer, or flow cytometer: an instrument that produces a whole spectrum, from which a researcher can read peak position, width, tail behavior, and subpopulation structure — not just one number.

A Proposed Framework, Not a Claimed Discipline

Nuclear-State Spectroscopy — a proposal, in three layers

Established spectroscopies (optical, mass, NMR, Raman) each separate three things: a methodology (the field itself), a measurement object (the spectrum — frequency spectrum, mass spectrum), and the instruments that implement it, of which there are usually many, made by different people, all producing the same kind of object. We propose applying that same separation here — as a proposal, not an established discipline.

Methodology

Nuclear-State Spectroscopy — a proposed framework for representing large nuclear populations as reproducible state spectra.

Measurement object

Nuclear-State Spectrum — the population-level distribution itself, independent of which software computes it.

Instrument

NucleoScope — one implementation of this framework, a first-generation instrument, not the framework itself.

Why we introduce a term, not just a method: the existing vocabulary for what we compute — histogram, distribution, density, kernel density estimate — describes the mathematical shape of the data, but not the idea that this shape represents a real state of the population being measured. That is exactly what "spectrum" has meant in established science.

Two different questions matter here, and they are easy to conflate: do we need a concept to describe our own work, and will that concept later become an established academic discipline. The second is not ours to answer. "Nuclear-State Spectrum" follows the naming precedent of terms like genome, proteome, and connectome — introduced to name a newly definable object of study. "Nuclear-State Spectroscopy" follows the naming precedent of terms like information theory, chaos theory, and systems biology — introduced as a way of organizing knowledge, whose originators did not get to decide whether it became a textbook field; that was decided later, by adoption or the lack of it.

The term is introduced because existing vocabulary does not adequately describe the measurement of reproducible state spectra from large nuclear populations. It is intended as a methodological framework for organizing measurement, phenomenology, and explanation — not a claim that a discipline already exists. Whether this framework eventually develops into an established research field is a question for the broader scientific community, not something we seek to determine ourselves.
Stage 1 — Measurement Does a reproducible Nuclear-State Spectrum exist at all? (stable statistical structure, not noise — independent of any disease label) ↓ Stage 2 — Phenomenology Does a specific spectral feature (e.g. Tail Runaway) associate with a specific biological state (e.g. malignancy)? (testable, falsifiable — this is the hypothesis under open validation) ↓ Stage 3 — Mechanism What explains that association? (future work, contingent on Stage 2 surviving independent replication)

This mirrors how other sciences developed: astronomers established that spiral arms exist and mapped their features long before density-wave theory explained why they form; X-ray diffraction established the double-helix structure years before base-pairing chemistry explained why it forms that way. Existence precedes explanation. This report works on Stages 1 and 2 only. Stage 3 is not addressed.

Stage 2 is not one finding — it's a space of possible entries, and cancer is only the first one filled in. Once Stage 1 (the spectrum) is established, many independent Stage 2 questions become askable, each with its own future Stage 3:

Filled in (this report)

  • Tail Runaway ↔ Cancer

Open — unexplored

  • Peak shift ↔ Aging?
  • Spectral broadening ↔ Drug response?
  • Subpopulation structure ↔ Development?
  • Spatial heterogeneity ↔ Tissue injury/repair?
A Note on Language

Measurement language, not definition language

Throughout this project we try to keep a specific linguistic discipline: describing what has been measured, not asserting what something is. This matters most at Stage 2 (Phenomenology) above, where it is easy to accidentally overstate an observed association as a definition.

What we use

  • corresponds to
  • is associated with
  • repeatedly appears together with
  • is observed together with
  • shows a consistent statistical relationship with

What we avoid at this stage

  • is
  • defines
  • proves
  • causes
  • determines
Tumor versus non-tumor status has so far been observed to correspond to a single statistical feature of the RSi distribution (Tail Runaway / Tail Closed) — without prior specification of species, tumor type, or tissue site.
Head-to-head · Pathology Foundation Models

Not "AI vs. no AI" — a difference in where generality comes from.

Pathology foundation models (e.g., UNI, Virchow, CONCH, Prov-GigaPath) are trained on very large, multi-organ whole-slide datasets to learn a representation that generalizes across tasks without per-organ retraining. The comparison worth making here is not whether NucleoScope uses AI, but where the generality comes from.

Foundation Model

Model Training ↓ Learned Representation ↓ Cross-task Generalization

NucleoScope

Fixed Measurement ↓ Statistical Distribution ↓ Measured Statistical Correspondence

Foundation models

  • Source of generalityLearned representation — a model trained on large-scale data to generalize across tasks
  • RequiresLarge training corpora, GPU compute, model weights
  • Generality isAn emergent property of training

NucleoScope's observation

  • Source of generality, if confirmedMeasured statistical correspondence — a fixed, predefined statistic applied without retraining
  • RequiresNo training, no labels, no per-organ calibration
  • Generality isAn empirical property of the nuclear population itself, if the observation replicates
We are not proposing a new classifier. We are asking whether pathology contains an intrinsic statistical measurement that has simply never been measured before. Foundation models ask a model to learn what is common across tasks from data; the observation described here asks whether some cross-task commonality can instead be directly measured, without training, using a single predefined statistic — a statistical correspondence, not a statistical convergence, since "convergence" would imply a process this report does not claim to observe. If this observation continues to hold under independent replication, it would represent an empirical phenomenon worth explaining — not a new classifier, and not a replacement for foundation-model approaches, which address a much broader range of tasks than the single readout (Tail Runaway/Closed) discussed in this report. If widely replicated, the question that would matter is no longer "why is this AI so accurate," but: why does a fixed statistical measurement, without knowing species, cancer type, or tissue site, still exhibit a consistent statistical correspondence?
Head-to-head · QuPath

The real difference isn't "who computes statistics" — it's "what is defined as the object of measurement."

Almost any image-analysis software can compute statistics on segmented objects — that alone is not a meaningful point of comparison. QuPath, Python, or even a spreadsheet can all compute distributions, histograms, or an FFT. The question that actually matters is different: does the field treat a particular output — like a frequency spectrum in signal processing, or a mass spectrum in mass spectrometry — as a defined, reproducible object of measurement in its own right, independent of which software computed it?

QuPath

  • Primary taskGeneral pathology image analysis, annotation, segmentation, object measurement
  • Primary objectIndividual cells, individual nuclei, tissue regions
  • StatisticsLeft to the researcher to design after export — QuPath does not define a specific population-level measurement object
  • OutputObject-level measurement tables, annotations, CSV

NucleoScope

  • Primary taskProposes a specific population-level measurement object: the Nuclear-State Spectrum
  • Primary objectA population of at least 10,000 nuclei, represented as a spectrum
  • StatisticsDefined by the framework itself — peak position, width, tail behavior, subpopulation structure
  • OutputA Nuclear-State Spectrum, with Tail Closed / Tail Runaway as one of its readable features
QuPath is a tool for computing statistics. NucleoScope proposes a measurement object — the Nuclear-State Spectrum — that other tools could, in principle, compute too.
Head-to-head · Hydrogen spectroscopy

Not the same physical mechanism — the same measurement logic.

Hydrogen spectroscopy

  • MeasuresA large number of hydrogen atoms
  • State variableWavelength, frequency, or energy
  • Population representationSpectrum
  • Final readoutCharacteristic spectral lines and pattern
  • Scientific useInfer atomic state from the spectrum

NucleoScope's Nuclear-State Spectrum

  • MeasuresA large number of nuclei
  • State variableImage-derived structural variable RSi
  • Population representationRSi statistical distribution
  • Final readoutTail Closed / Tail Runaway
  • Scientific useInfer population state from the distribution
A single object can't form a spectrum; only the state distribution of many objects produces a reproducible population readout.

What it is

  • A quantitative measurement tool for research use
  • Focused on population-level nuclear structural statistics
  • Outputs auditable values and distributions

What it is not

  • Not a clinical diagnostic system
  • Not a black-box cancer-probability model
  • Not a theory of cancer mechanism