The problem with one unexplained score

Technical software often compresses a complex process into a single number: health 82, efficiency 71, anomaly 0.84, risk high. A summary can be useful, but without context the user cannot tell whether the result comes from one sensor, ten signals, a heuristic threshold or an AI model.

Trustworthy software should make the path inspectable even if the user normally sees only the summary.

1. Preserve the raw or source data

Whenever practical, keep the original input: measurement file, image, protocol response, time series or configuration. Derived values should be reproducible from that source rather than replacing it.

2. Record version and configuration

A result produced by algorithm v1.4 with calibration A is not necessarily comparable with the same file processed by v2.0. Versioning matters for research software, building automation and AI-assisted workflows alike. Store software version, algorithm version, relevant settings and timestamp.

3. Separate measurement from interpretation

This is one of the strongest general rules. A BACnet Present_Value is a measured or reported value; the conclusion “valve is probably stuck” is diagnostic reasoning. A GDV image feature is measured; an organ-sector statement is interpretation. A World State correlation is calculated; a causal claim is a hypothesis.

4. Make uncertainty visible

Not every result deserves the same confidence. Image quality may be poor, a sensor may be stale, a statistical sample may be small or an AI model may be uncertain. Hiding these limitations makes software look more confident while making it less trustworthy.

Useful outputs can include confidence, sample size, quality score, repeatability, data age and whether the finding is direct or inferred.

5. Prefer repeated evidence over one dramatic event

A single anomalous point can be noise. Repetition under similar conditions is more informative. In building automation, a valve command and temperature response can be observed across cycles. In GDV research, the same morphology can be checked across repeated captures. In time-series research, a discovery relationship should be tested on unseen data.

6. AI should explain and assist, not erase the trail

AI is valuable for summarising logs, comparing sessions, finding inconsistencies and translating technical results into readable language. It becomes risky when it replaces the data trail with a fluent conclusion. The original measurement and deterministic calculations should remain visible where possible.

7. Reproducibility is a product feature

Reproducibility is often treated as an academic concern, but it is useful in ordinary engineering. A commissioning engineer wants to know which register format produced a value. A researcher wants to know which algorithm created a metric. A support technician wants to recreate the customer's state.

Examples across PICALLW projects

NODVIA: preserve device identity, address, protocol context, write permissions and trends before diagnostic interpretation.

GDV Studio: preserve original captures, quality checks, morphology, repeatability and personal baseline before interpretation.

World State Explorer: keep discovery and validation periods separate, control false discoveries and reject relationships that fail on unseen data.

Digital Radionics: treat symbolic workflows as experimental practice, preserve session settings and avoid presenting subjective interpretation as measured physical evidence.

A practical checklist

  • Can the user see or recover the source data?
  • Can the result be reproduced with the same version and settings?
  • Are measured values clearly separated from interpretation?
  • Are uncertainty and data quality visible?
  • Can a later reviewer understand why the software produced the conclusion?
  • Does AI preserve rather than hide the evidence trail?

Sources and frameworks

Related reading