Skip to content

ILLUSTRATIVE · NASA / SVS · NOT ADITYA-L1 DATA

Scientific conclusions

A threshold beat the models where it counts.

Machine learning ranks individual minutes better — but it detects no more flares, and it raises about five times as many false alarms. A one-line detector on the SoLEXS count rate remains the operational recommendation.

What is this?

The scientific conclusions of the project, stated at the scale of a claim — including the result that says machine learning did not help.

Why should I trust it?

Every figure on this page is read from the same committed benchmark artifact that the full method tables render from, so the summary cannot drift from the evidence it summarises.

Where can I verify it?

Every model, against the simple detector

M/X nowcast · higher is better

  • Randomtrivial0.4970.483–0.509
  • Majoritytrivial0.5000.500–0.500
  • Climatologytrivial0.5000.500–0.500
  • Persistencetrivial0.9820.978–0.986
  • Threshold (rate)0.9540.940–0.966
  • Logistic regression0.9640.953–0.974
  • Random forest0.9660.956–0.976
  • LightGBM0.9610.949–0.972
0.450.731.00

ROC-AUC · bars show the 95% confidence interval

On ROC-AUC the intervals overlap: no model separates itself from the threshold. That is the basis for the verdict — but ROC-AUC is not the whole story at a 1.2% base rate, which is why the validation page reports precision, recall and the false-alarm cost alongside it.

Why this result holds

The comparison is decided on the metric that matters operationally — how many real flares are caught, and at what alarm cost — not on the metric that flatters a model.

Decided in advance

seed 20260718

The protocol was frozen before any model was fit.

Extra spectral bands

+0.0033

Δ ROC-AUC from added spectral resolution. A confirmed null.

Held-out test

192,541

Minutes from 2026-01-01 00:00:00+00:00 onward, never seen in training.